Method for identifying and describing whole-process activity patterns of personnel based on mobile phone signaling data

Through the segmentation, matching and clustering methods based on mobile phone signaling data, the entire process activity chain of urban residents is identified and optimized, and the problem of insufficient accuracy and timeliness of traditional methods is solved, high-precision activity pattern recognition is achieved, and decision-making and planning of smart cities are supported.

CN120075745BActive Publication Date: 2025-07-11BEIHANG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510547321.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-11
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

Traditional data collection and analysis methods are difficult to meet the needs of modern urban management for high precision and timeliness. Mobile phone signaling data has challenges in data redundancy, noise interference and privacy protection. How to extract valuable information from these complex data to build a method system for urban residents' full-process activity mode is difficult.

Method used

By identifying the user's stay status, the base station sequence matching is used to identify the mobile mode, and the travel method is identified by combining clustering and speed and distance rules, description tags are added and merged optimized to finally generate the entire process activity chain.

Benefits of technology

It has achieved rapid, comprehensive and accurate portrayal of personnel activity patterns within the target city, provided high-precision travel pattern recognition, provided scientific basis for decision-making such as urban transportation organization, emergency response and commercial layout, and promoted the construction of smart cities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075745B_ABST
    Figure CN120075745B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for recognizing and describing the whole-process activity patterns of personnel based on mobile phone signaling data, belonging to the field of data mining, including: S1: Based on the signaling data of the user throughout the day, by identifying the user's staying state, making a preliminary activity semantic segmentation of the user's all-day activity chain, and forming the staying segments and moving segments of the activity chain; S2: Using the method of base station sequence matching, identifying the travel patterns in the moving segments, further segmenting part of the moving segments, and filling them into the activity chain to obtain a basic version of the activity chain; S3: Identifying the travel patterns on the basis of the basic version of the activity chain and filling them into the basic version of the activity chain to obtain a synthesized version of the activity chain; S4: Checking and analyzing the staying segments and adding description tags; for the synthesized version of the activity chain, optimizing and merging it to generate the final version of the whole-process activity chain. The method of the present invention comprehensively uses technical means such as data mining, machine learning, and spatio-temporal analysis to achieve a comprehensive and accurate characterization of the activity patterns of urban residents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of data mining, and in particular relates to a method for identifying and describing a personnel full-process activity pattern based on mobile phone signaling data. Background Art

[0002] With the acceleration of urbanization and the rapid development of information technology, the concept of smart cities has become an important way to improve urban management efficiency and optimize the quality of life of residents. In the construction of smart cities, how to comprehensively and accurately grasp the activity patterns of urban residents has become a key issue that needs to be solved urgently. Traditional data collection and analysis methods, such as questionnaire surveys and traffic flow statistics, often have shortcomings such as long data collection cycles, limited coverage, and poor real-time performance, which are difficult to meet the needs of modern urban management for high precision and high timeliness.

[0003] In recent years, with the popularization of mobile communication technology and the widespread use of smart phones, mobile signaling data, as a new data source, has gradually attracted attention in the fields of urban planning and traffic management. Mobile signaling data refers to the records of location updates, calls, text messages, data connections, etc. generated by mobile phone users interacting with base stations during the use of mobile communication services. These data are not only massive, but also can reflect the movement trajectory and activity behavior of mobile phone users in real time, providing unprecedented possibilities for analyzing the activity patterns of urban residents.

[0004] However, mobile phone signaling data also faces challenges such as data redundancy, noise interference, and privacy protection. How to extract valuable information from these complex data and build a method system that can accurately describe the activity patterns of urban residents throughout the entire process is a hot topic and difficulty in current research. Summary of the invention

[0005] In order to solve the above technical problems, the present invention provides a method for identifying and describing the whole-process activity pattern of personnel based on mobile phone signaling data, comprising the following steps:

[0006] Step S1: Based on the user's signaling data throughout the day, by identifying the user's stay status, the user's activity chain throughout the day is preliminarily segmented into stay segments and movement segments of the activity chain;

[0007] Step S2: using base station sequence matching to identify the movement mode in the movement segment, including subway travel mode, high-speed rail travel mode, and bus travel mode, further segmenting the movement segment, filling the activity chain, and obtaining a basic version of the activity chain;

[0008] Step S3: On the basic version of the activity chain, identify travel modes of driving, walking, and cycling using clustering and inference rules based on speed and travel distance, and fill them into the basic version of the activity chain to obtain the synthesized version of the activity chain;

[0009] Step S4: Check and analyze the stay segments and add descriptive labels; optimize and merge the synthesized version of the activity chain to generate the final version of the full-process activity chain.

[0010] Beneficial effects:

[0011] The present invention provides a method for identifying and describing the full-process activity patterns of personnel based on mobile phone signaling data. By extracting key spatio-temporal information from a large amount of rough signaling data and applying a big data processing framework that integrates multiple machine learning means, it realizes a fast, comprehensive, and accurate description of the activity patterns of all personnel within the target city scope. This method solves the difficulties of large data volume, low accuracy, and roughness of signaling data itself, and achieves a high-precision travel mode recognition effect with a big data processing framework that supports real-time calculation, providing technical support for downstream tasks in human behavior understanding and pattern extraction, providing a scientific basis for decision-making in urban traffic organization and optimization, emergency response, commercial layout, etc., and having important significance for promoting the construction and development of smart cities. Description of the Drawings

[0012] Figure 1 It is a schematic flowchart of a method for identifying and describing the full-process activity patterns of personnel based on mobile phone signaling data according to the present invention;

[0013] Figure 2 It is a schematic diagram of the line GID grid sequence and the trajectory GID grid sequence;

[0014] Figure 3 It is a schematic diagram of the generation process of the basic version of the activity chain. Detailed Embodiments

[0015] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0016] Embodiment 1

[0017] As Figure 1 shown, a method for identifying and describing the full-process activity patterns of personnel based on mobile phone signaling data provided by an embodiment of the present invention includes the following steps:

[0018] Step S1: Based on the user's signaling data throughout the day, by identifying the user's staying status, perform a preliminary activity semantic segmentation on the user's activity chain throughout the day to form the staying segments and moving segments of the activity chain;

[0019] Step S2: Adopt the method of base station sequence matching to identify the moving patterns in the moving segments, including subway travel pattern, high - speed rail travel pattern, and bus travel pattern, further segment the parts of the moving segments, and fill them into the activity chain to obtain the basic version of the activity chain;

[0020] Step S3: On the basic version of the activity chain, adopt clustering and inference rules based on speed and travel distance to identify the travel patterns of driving, walking, and cycling, and fill them into the basic version of the activity chain to obtain the synthetic version of the activity chain;

[0021] Step S4: Check and analyze the staying segments, and add description tags; for the synthetic version of the activity chain, perform optimization and merging to generate the final version of the full - process activity chain.

[0022] In one embodiment, the above - mentioned Step S1: Based on the user's signaling data throughout the day, by identifying the user's staying status, perform a preliminary activity semantic segmentation on the user's activity chain throughout the day to form the staying segments and moving segments of the activity chain, specifically including:

[0023] Step S11: Put the signaling data into a sliding window, determine whether the time period is a staying status according to the repetition degree of the longitude and latitude data in the sliding window, calculate the staying points, and the calculation formula is as follows:

[0024] (1)

[0025] (2)

[0026] Among them, represents the repetition degree of the coordinate of the signaling data t in the window, represents the size of the window, is an indicator function, and are the coordinates of the signaling data numbered i and t respectively; represents the activity status of the signaling data t, including staying and moving;

[0027] Step S12: For each signaling data with a staying activity status, retain its longitude and latitude coordinates and staying time, weight them according to the staying time and perform clustering on the map, and take the center coordinate of each cluster after clustering as the staying point, specifically including:

[0028] Step S121: For each screened staying signaling data, calculate its staying time in the sliding window ( ). The residence time can be obtained from the timestamps of the signaling data, i.e., the end timestamp ( ) minus the start timestamp ( ). This time difference represents the duration that the user stays at this location. According to the length of the residence time, a weight is assigned to each stay signaling. The longer the residence time, the greater the weight. The purpose of weighting is to make the points with longer stay times have a greater impact on the clustering result in the subsequent clustering process. The calculation formula for the weight is as follows:

[0029] (3)

[0030] Where, represents the weight of the currently calculated signaling, represents the residence time of this signaling.

[0031] Step S122: Perform clustering using the DBSCAN clustering algorithm. Input all the weighted stay signaling data (including: longitude and latitude and weight weight) into the selected clustering algorithm for clustering analysis.

[0032] Step S123: For each cluster obtained after clustering ( ), calculate its average value to obtain the center coordinates ( ). The cluster center coordinates will serve as the stationary points, representing the typical stay locations of the user in this area. Output the center coordinates (i.e., the stationary points) of each cluster obtained after clustering, and record relevant information, including the ID of the cluster and the stay time period.

[0033] Step S13: Using the stationary points as the segmentation points, divide the activity chain into multiple moving segments and stay segments to complete the preliminary semantic segmentation, providing a basis for subsequent activity chain segmentation and semantic annotation.

[0034] In one embodiment, the above Step S2: Adopt the method of base station sequence matching to identify the movement patterns in the moving segments, including subway travel patterns, high-speed rail travel patterns, and bus travel patterns, further divide the parts of the moving segments, and fill them into the activity chain to obtain the basic version of the activity chain, specifically including:

[0035] Step S21: Obtain subway, bus, and high-speed rail operation data, including: line description information data, station description information data, and line topology structure data, specifically including:

[0036] Step S211: Obtain subway line description information data, including: line ID, line name, line direction, whether the line is a loop line, starting and ending stations, and passing stations;

[0037] In the line description information data, each of the two running directions of a physical line corresponds to a line ID, that is, a line ID uniquely determines a line traveling in a fixed direction. For example, the line IDs corresponding to Beijing Subway Line 10 are "Line 10_Inner Loop" and "Line 10_Outer Loop".

[0038] Step S212: Obtain high-speed rail line description information data, including line ID, line name, line direction, whether it is a double-track line, starting and ending stations, and passing stations.

[0039] Step S213: Obtain high-speed rail line description information data, including line ID, line name, line direction, starting and ending stations, and passing stations.

[0040] Step S214: According to the poi information, obtain the latitude and longitude position coordinates of all stations.

[0041] Step S22: Generate running line data according to the line topologies of subways, buses, and high-speed rails. The form of each running line data is an ordered sequence of latitude and longitude coordinates; then map the latitude and longitude coordinates of the running line data to the grids named in the form of GID to form a line GID grid sequence;

[0042] Step S23: Scan the latitude and longitude coordinates of all base stations, establish a mapping table between the base station IDs that may interact when traveling into a certain GID grid and the GID grid ID, and convert the signaling data into a trajectory GID grid sequence through this mapping table;

[0043] As Figure 2 shows a schematic diagram of the line GID grid sequence (left) and the trajectory GID grid sequence (right).

[0044] Step S24: Calculate the edit distance between the trajectory GID grid sequence and the line GID grid sequence to measure whether the trajectory of the person matches the line; if the edit distance does not exceed the matching threshold, it is determined that the travel segment is the corresponding subway, bus, or high-speed rail travel mode, and fill it into the activity chain to obtain the basic version of the activity chain.

[0045] After multiple experiments, the embodiment of the present invention sets the matching threshold to 80%.

[0046] The definition and calculation pseudo-code of the edit distance are as follows:

[0047] The Edit Distance, also known as the Levenshtein distance, is a way to measure the difference between two strings. It is defined as the minimum number of edit operations required to transform one string into another, where the edit operations include inserting, deleting, and replacing characters. Here, the elements for the Edit Distance operation are the sequences of GID grids, that is, it calculates the number of edits required to transform one grid sequence into another.

[0048] When calculating the Edit Distance, the dynamic programming algorithm is generally used, and its pseudocode is as follows:

[0049] FUNCTION EditDistance(gridSequence1, gridSequence2)

[0050] LET n = LENGTH(gridSequence1)

[0051] LET m = LENGTH(gridSequence2)

[0052] LET dp = ARRAY[0..n, 0..m] / / Create a two-dimensional array, where dp[i][j] represents the Edit Distance between the first i grid sequences of gridSequence1 and the first j grid sequences of gridSequence2

[0053] / / Initialize the first row and the first column

[0054] FOR i FROM 0 TO n

[0055] dp[i][0] = i / / It takes i deletions to change the first i grid sequences of gridSequence1 into an empty grid sequence string

[0056] END FOR

[0057] FOR j FROM 0 TO m

[0058] dp[0][j] = j / / It takes j insertions to change an empty grid sequence string into the first j grid sequences of gridSequence2

[0059] END FOR

[0060] / / Fill the dp array

[0061] FOR i FROM 1 TO n

[0062] FOR j FROM 1 TO m

[0063] If gridSequence1[i - 1] == gridSequence2[j - 1]

[0064] cost = 0 / / The grid sequences are the same, no editing is required

[0065] Else

[0066] cost = 1 / / The grid sequences are different, 1 edit is required

[0067] End If

[0068] dp[i][j] = MIN(dp[i - 1][j] + 1, / / Delete operation

[0069] dp[i][j - 1] + 1, / / Insert operation

[0070] dp[i - 1][j - 1] + cost) / / Replace operation

[0071] End For

[0072] End For

[0073] / / dp[n][m] contains the edit distance between the entire gridSequence1 and gridSequence2

[0074] Return dp[n][m]

[0075] End Function

[0076] The pseudocode for determining the travel mode of a travel segment is as follows:

[0077] Input: Sequence of line coordinates , sequence of user travel base stations .

[0078] Output: The corresponding travel mode

[0079] for :

[0080] = ( )

[0081] for :

[0082] if is in s or is absent from , where then

[0083] add to

[0084] for :

[0085] if is in then

[0086] add to

[0087] =

[0088] if then

[0089] belongs to a certain mode of transportation

[0090] Figure 3 Shows the generation process of the basic version of the activity chain.

[0091] In one embodiment, the above step S3: On the basic version of the activity chain, using clustering, inference rules based on speed and travel distance, identify travel modes of driving, walking, and cycling, and fill them into the basic version of the activity chain to obtain the synthesized version of the activity chain, specifically including:

[0092] Step S31: Based on the basic version of the activity chain, select the start, end nodes and intermediate nodes of each section of travel, and calculate the average speed, travel time and travel distance of each section of travel;

[0093] Step S32: Use the average speed, travel time and travel distance as travel characteristics to cluster the travel segments;

[0094] Step S33: Select the clusters with low speed, short time and short distance as walking segments, the clusters with medium speed, medium time and medium distance as cycling segments, and the segments with high speed, long time and long distance as driving segments to obtain the synthesized version of the activity chain.

[0095] In one embodiment, the above-mentioned step S4: Check and analyze the stay segments, and add descriptive tags; optimize and merge the synthetic version of the activity chain to generate the final version of the whole-process activity chain, specifically including:

[0096] Step S41: Add tags for the stay mode to the stay segments according to the stay duration and stay moment, including: work, home, short stop outside, specifically including:

[0097] Step S411: Obtain the stay time of each point, and divide the stay duration into long stays and short stays according to a preset time threshold;

[0098] After multiple experiments, the value of this time threshold in the embodiment of the present invention is 1 hour.

[0099] Step S412: Analyze whether the start and end time periods of the stay are within the working hours on weekdays, and combine the time pattern (such as morning and evening rush hours) with whether it is a long stay to further infer whether the stay point is a work point;

[0100] Step S413: Add detailed mode tags to the stay points, such as "working on weekdays", "staying at home on weekends", "short stop outside on weekdays", etc.;

[0101] Step S42: Generate a travel sequence according to the synthetic version of the activity chain, remove and replace unreasonable travel mode sequences, specifically including:

[0102] Step S421: Detect public transportation modes according to the travel period, including subways, buses, and high-speed rails. If it is in the non-operating time period (from zero to five o'clock every day), then replace it with driving, cycling, and walking modes according to the speed;

[0103] Step S422: Mark driving, a non-public travel mode, and buses and subways as mutually exclusive options. If it is detected that the activity chain contains driving adjacent to buses and subways, then replace it with adjacent activity modes;

[0104] Step S43: Merge travel modes in the same state to generate the final version of the whole-process activity chain.

Claims

1. A method for recognizing and describing the whole-process activity pattern of personnel based on mobile phone signaling data, characterized in that Including: Step S1: Based on the user's signaling data throughout the day, by identifying the user's stay status, perform a preliminary activity semantic segmentation on the user's activity chain throughout the day to form stay segments and movement segments of the activity chain, specifically including: Step S11: Put the signaling data into a sliding window, determine whether the corresponding time period is a stay status according to the repetition degree of the longitude and latitude data in the sliding window, and calculate the stay points. The calculation formula is as follows: (1) (2) Among them, represents the repetition degree of the coordinates of signaling data t within the window, represents the size of the window, is an indicator function, and are the coordinates of signaling data with numbers i and t respectively; represents the activity state of signaling data t, including staying and moving; Step S12: For each signaling data with a stay activity status, retain its longitude and latitude coordinates and retention time, weight them according to the retention time and perform clustering on the map, and use the center coordinates of each cluster after clustering as the stay points; Step S13: Use the stay points as segmentation points to segment the activity chain into multiple movement segments and stay segments to complete the preliminary semantic segmentation; Step S2: Adopt the method of base station sequence matching to identify the movement patterns in the movement segments, including subway travel mode, high-speed rail travel mode, and bus travel mode, further segment the part of the movement segment, and fill it into the activity chain to obtain a basic version of the activity chain, specifically including: Step S21: Obtain subway, bus, and high-speed rail operation data, including: line description information data, station description information data, and line topology structure data; Step S22: According to the line topology structures of the subway, bus, and high-speed rail, generate operation line data, and the form of each operation line data is a set of ordered longitude and latitude coordinate sequences; then map the longitude and latitude coordinates of the operation line data to a grid named in the form of GID to form a line GID grid sequence; Step S23: Scan the longitude and latitude coordinates of all base stations, establish a mapping table between the base station IDs that may interact when traveling to a certain GID grid and the GID grid ID, and convert the signaling data into a trajectory GID grid sequence through this mapping table; Step S24: Calculate the edit distance between the trajectory GID grid sequence and the line GID grid sequence to measure whether the trajectory of the person matches the line; if the edit distance does not exceed the matching threshold, determine that the travel segment is the corresponding subway, bus, or high-speed rail travel mode, and fill it into the activity chain to obtain a basic version of the activity chain; Step S3: On the basic version of the activity chain, adopt clustering and inference rules based on speed and travel distance to identify travel modes of driving, walking, and cycling, and fill them into the basic version of the activity chain to obtain a synthesized version of the activity chain, specifically including: Step S31: Based on the basic version of the activity chain, select the start, end nodes and intermediate nodes of each travel segment, and calculate the average speed, travel time and travel distance of each travel segment; Step S32: Use the average speed, travel time and travel distance as travel characteristics to perform clustering on the travel segment; Step S33: Select the clusters with low speed, short time, and short distance as walking segments, the clusters with medium speed, medium time, and medium distance as cycling segments, and the segments with high speed, long time, and long distance as driving segments to obtain a synthesized version of the activity chain; Step S4: Check and analyze the stay segments and add description tags; for the synthesized version of the activity chain, perform optimization and merging to generate a final version of the whole-process activity chain, specifically including: Step S41: Add tags of stay modes to the stay segment according to the stay duration and stay moment, including: working, staying at home, short stop while going out; Step S42: Generate a travel sequence according to the synthesized activity chain, and remove and replace unreasonable travel mode sequences; Step S43: Merge travel modes in the same state to generate the final whole-process activity chain.

Citation Information

Patent Citations

  • Mobile phone user trip mode identification method based on mobile phone signaling data and navigation route data

    CN106197458A