A method and system for identifying a career type of a shared bicycle commuter

By filtering shared bicycle data and combining gravity models and Bayesian rules, the occupational types of shared bicycle users can be identified, solving the problem of insufficient accuracy in existing technologies and providing technical support for urban management and resource optimization.

CN117056823BActive Publication Date: 2026-06-02SOUTHEAST UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SOUTHEAST UNIV
Filing Date
2023-07-24
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify the occupational types of shared bicycle users, especially in large-sample data identification where accuracy is insufficient. They neglect the importance of users' social attributes and fail to consider the randomness of shared bicycle parking and network latency errors.

Method used

By receiving POI data, urban road network data, and shared bicycle data, valid shared bicycle data is filtered out. Morning rush hour commuting data and cycling destination clustering are used, combined with gravity models and Bayesian rules to identify user occupation types and establish a correspondence between POIs and occupation types.

Benefits of technology

It has enabled relatively accurate identification of the occupational types of shared bicycle commuters, providing technical support for urban space utilization and resource allocation, and optimizing shared bicycle dispatch management and facility planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117056823B_ABST
    Figure CN117056823B_ABST
Patent Text Reader

Abstract

The application discloses a method and system for identifying the professional type of a shared bicycle commuting user, relates to the technical field of traffic management and urban planning, and comprises the following steps: receiving POI data, urban road network data and shared bicycle data, and processing the shared bicycle data to obtain effective shared bicycle data; calculating the normal parameter of the user using the shared bicycle in the morning peak of a weekday, performing primary screening on the effective shared bicycle data, and extracting primary effective shared bicycle data; performing secondary screening on the primary effective shared bicycle data in the time dimension to obtain secondary effective shared bicycle data; performing tertiary screening on the secondary effective shared bicycle data in the space dimension to obtain tertiary effective shared bicycle data; determining the decision area of the professional type of the user according to the maximum walking radius and in combination with the urban road network data, determining the corresponding relationship between the POI data and the professional type, and identifying the professional type of the user by using the POI data, the tertiary effective shared bicycle data and in combination with the decision area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of traffic management technology and urban planning, specifically a method and system for identifying the occupational types of shared bicycle commuter users. Background Technology

[0002] Commuting is a fundamental activity in residents' daily lives. Exploring the spatial and temporal patterns of commuting among different groups, in conjunction with residents' social attributes, helps to strengthen the refined research on "population-spatial behavior patterns," tap into the differentiated mechanisms of "human-land interaction," and specifically improve the supply of corresponding urban systems to build high-quality livable cities. With the popularization of the sharing economy and the improvement and application of related technologies such as internet communication technology, GNSS (Global Navigation Satellite System) positioning technology, and satellite remote sensing, shared bicycles have become the optimal choice for short-distance travel and connecting to public transportation to complete the "last mile" of travel due to their high convenience and cost-effectiveness. According to the 2017 Aurora Data "China Small and Medium-Sized Cities Shared Bicycle Development Report," the purposes and proportions of shared bicycle use are: commuting to and from get off work (26.4%), shopping (25.4%), exercise (23.3%), scenic tourism (11.4%), going to school (10.5%), and other (2.9%). It is evident that shared bicycles have become one of the important public transportation tools for residents to complete their commutes. Therefore, exploring the spatiotemporal patterns of shared bicycle use from the perspective of social attribute differentiation is also of great value and significance: it involves disciplines such as geography, urban planning, and traffic management, which is conducive to a deeper understanding of behavioral geography and temporal geography from a social perspective, promoting the research of related theoretical models, optimizing and adjusting urban spatial structure and functional zoning, improving the supply of urban transportation systems, and building intelligent transportation management systems, etc.

[0003] The social attributes of shared bike users typically include age, gender, education level, occupation, and income level. Among these, occupation is most closely related to commuting. However, this type of data is not only difficult to obtain directly through web scraping, but it is also easily overlooked when filling gaps in the data through questionnaire sampling. Regarding the feasibility of identifying the occupation types of shared bike users, on the one hand, the occupation types of most users largely match their work locations; on the other hand, the spatial information and nature of work locations can be reflected through the industry field of POI (Point of Interest); furthermore, work locations are closely related to the land use and functional structure of cities, and the deployment of shared bikes is often closely linked to urban land use functions (e.g., more shared bikes are usually deployed in high-tech industrial parks and office areas to match supply and demand). This provides a realistic basis and practical possibility for identifying the occupation types of shared bike commuting users from a large data sample. In short, as one of the few social attributes that can be derived from data, the identification of occupational types still has research gaps and is closely related to commuting, which has a significant impact on urban space. More importantly, previous studies have often remained at a rough description of a general state, rather than being able to finely explore deeper patterns.

[0004] Current research on shared bicycles, both domestically and internationally, mainly focuses on riding routes and hotspot area identification, while research on methods for identifying user occupation types is relatively limited. From a content perspective, existing research on user occupation type identification primarily focuses on identifying travel purposes, and then exploring user activity types and spatiotemporal patterns through methods such as POI matching. Technically, this mainly relies on vehicle spatiotemporal trajectory data, combined with mobile phone signaling data, POI data, AOI (Area of ​​Interest) data, and TUD (Technical Data Interpretation). Data from the University of Dortmund, urban road network data, etc., are used to predict shared bicycle destinations using Bayesian rules, gravity models, DMR models, k-nearest neighbors, and other models. Some scholars have also used K-means++ clustering based on shared bicycle data and POI data to study shared bicycle travel patterns and travel purposes. In general, the above methods have the following problems: (1) In terms of research content, they mainly focus on the spatiotemporal aspects and travel purposes, and lack attention to the social attributes of users such as age, gender, education level, occupation type, and income level. In particular, occupation type, which is the social attribute most closely related to commuting, is difficult to obtain directly using web crawlers and is also easily ignored in sampling questionnaires; (2) In terms of technical methods, there is still insufficient consideration of the complex real-world situation, and the spatiotemporal differences in the commuting behavior characteristics of residents of different occupation types are ignored. At the same time, using a single riding destination to search for the corresponding destination does not take into account the randomness of shared bicycle parking and the errors caused by network latency, so it is difficult to guarantee high accuracy in the identification of large sample data. Summary of the Invention

[0005] To address the shortcomings mentioned in the background section, the present invention aims to provide a method and system for identifying the occupational types of shared bicycle commuter users.

[0006] The objective of this invention can be achieved through the following technical solution: a method for identifying the occupational type of shared bicycle commuter users, the method comprising the following steps:

[0007] Receive POI data, urban road network data, and shared bicycle data; clean and filter the shared bicycle data to obtain valid shared bicycle data.

[0008] Calculate the normalized parameters of users' use of shared bicycles during the morning rush hour on weekdays. Use the normalized parameters of users' use of shared bicycles during the morning rush hour on weekdays to filter and extract the effective data of shared bicycles, and obtain one set of effective data of shared bicycles.

[0009] The system counts the cumulative number of days users use shared bikes during a week, and then performs a secondary filtering and extraction of the valid shared bike data based on the time dimension to obtain secondary valid shared bike data.

[0010] The system analyzes the destinations of users using shared bikes on weekdays throughout the week, and then performs a third round of filtering and extraction on the secondary shared bike data in terms of spatial dimensions to obtain the third round of valid shared bike data.

[0011] By using the maximum walking radius and urban road network data, the decision area of ​​the user's occupation type is determined, the correspondence between POI data and occupation type is determined, and the user's occupation type is identified by using POI data and three valid shared bicycle data and combining them with the decision area of ​​the user's occupation type.

[0012] Preferably, the information in the shared bicycle data includes order number, user number, vehicle number, start and end time of the ride, and geographical coordinates of the riding trajectory.

[0013] Preferably, the shared bicycle data is cleaned and filtered to select shared bicycle data with a riding time of no more than 1 hour and an average riding speed in the range of 1 km / h-30 km / h as valid shared bicycle data.

[0014] Preferably, the normalized parameters for users using shared bicycles during the morning rush hour on weekdays include calculating the maximum tolerable walking distance threshold from the starting point to the pick-up point / riding start point and identifying the morning rush hour.

[0015] The basis and method are as follows:

[0016] Since commuting typically ends at the residence and workplace, and considering the daily commuting usage patterns of shared bikes, the starting points generally include three possible locations: residence (for the entire journey), bus or subway (for transfers). According to relevant research reports, shared bikes primarily replace walking, with limited substitution for public transportation, and the use of bikes to connect to the subway is more widespread than using buses. Therefore, the order of shared bike data extraction is: subway station entrance / exit → bus stop → residence. Let r represent these locations respectively. M r B and r H Using subway station entrances / exits, bus stops, and residential buildings in open communities (or entrances / exits of gated communities) as basic units, multi-ring buffer zones are created. The frequency of shared bicycle riding start points within each buffer ring is then statistically analyzed. The distance corresponding to the inflection point of the frequency change of shared bicycle riding start points within a buffer ring is used as a threshold (R). M R B R H ).

[0017] Because morning commutes typically involve time constraints for clocking in and out, while users have more free time after get off work, the purposes for using shared bikes in the evening are more diverse and complex than in the morning. Therefore, shared bike data during the morning rush hour is a more accurate reflection of commuting usage. Based on everyday experience, the typical weekday morning rush hour is determined to be 7:00-10:00 AM. Then, valid data on shared bike order start times within this timeframe is extracted.

[0018] Preferably, the process of performing secondary filtering and extraction of valid data from a single shared bicycle transaction along the time dimension is as follows:

[0019] Users who use shared bikes for 3 or more days out of a 5-day work week are considered valid users. Each valid shared bike usage session is then deemed valid. Let t be the cumulative number of days (5 workdays) of shared bike usage within a 5-day work week, and then the following criteria are used for determination:

[0020] If t≤2, then the valid data for the corresponding shared bicycle will be removed;

[0021] If t>2, then extract the valid data corresponding to the first shared bicycle and mark it as valid data for the second shared bicycle.

[0022] The basis for this is that users generally work at least 3 days a week (i.e., production workers working 3 shifts, excluding freelancers). Accordingly, there should be at least 3 or more days out of a 5-day work week where shared bike usage data is recorded. Therefore, users who have used shared bikes for 3 days or more are considered valid users, and their corresponding shared bike usage data is considered valid data.

[0023] Preferably, the effective data of secondary shared bicycles is filtered and extracted three times in the spatial dimension, based on the following: First, considering the randomness of users parking shared bicycles in real life, it is difficult to guarantee that users will accurately park their shared bicycles in the same location (same geographical coordinates) every day on their way to work. Therefore, the riding destinations that recur within a week should be clustered within a certain spatial range, forming a cluster of all possible riding destinations for a certain destination. Second, commuters have fixed workplaces, so the spatial location (geographical coordinates) of shared bicycle riding destinations should recur on weekdays with different date markings. Third, users have a stronger rigid demand for using shared bicycles for commuting during weekdays compared to other travel purposes, so the number of riding destinations in the effective cluster should be the largest. The specific steps are as follows:

[0024] For any valid user, identify and classify all possible cycling destinations for a given destination. For any user, use spatial coordinate data to calculate the distance between each pair of all cycling destinations for that user, and perform cluster analysis with distance as a parameter to output all possible cycling destination location data for each user's destination.

[0025] If, for any valid user, the number of different date labels in all possible cycling destinations for a given destination exceeds half of the user's cumulative usage days, then that destination will be considered as one of the user's possible candidate work locations.

[0026] The candidate work location with the most possible cycling destination data is identified as the most likely work location. For the user's possible work locations and all possible cycling destination data, the number of cycling destination data contained in them is counted. The destination with the most cycling destination data is extracted as the user's most likely work location. The corresponding secondary shared bike valid data is extracted and marked as tertiary shared bike valid data.

[0027] Preferably, if the number of possible work locations and all possible cycling destinations of a user has multiple identical maximum values, then the sum of the distances between each pair of possible cycling destinations of the candidate work location is compared, and the candidate work location with the smallest sum of distances is taken as the most likely work location for the user.

[0028] Preferably, the process for defining the decision area for the user's occupation type is as follows:

[0029] Calculate the maximum tolerable walking distance r (maximum walking radius) from the parking point / riding endpoint D to the destination. This is done by calculating the distance between the shared bike's riding endpoint and its nearest neighbor POI data, and plotting a histogram of the maximum walking radius versus the proportion of shared bike data. The maximum tolerable walking distance from the parking point / riding endpoint to the destination is adaptively determined, and the following two conditions should be met simultaneously:

[0030] Within a circular area centered on the destination of the shared bicycle ride and with radius r, at least one POI can be found.

[0031] Shared bike data that meets the above criteria accounts for 90% of the total shared bike data.

[0032] For any most probable work location and all possible cycling endpoints, using one of these endpoints as the center and the maximum tolerable walking distance *r* from the parking point / endpoint D to the destination as the radius, calculate the walkable range based on the actual road network. Then, merge the walkable ranges of all cycling endpoints to form the decision region A for the shared bicycle user's occupation type. d .

[0033] Preferably, the method for establishing the correspondence between POI data and occupation types is as follows:

[0034] The industry classification of Gaode Map's POI data was compared with the People's Republic of China National Standard "Occupational Classification and Codes (GB / T6565-2009)" to accurately identify users' occupational types. The identified occupational types include five categories: professional and technical personnel, clerical and related personnel, business service personnel, production and transportation equipment operators and related personnel, and agricultural, forestry, animal husbandry, fishery, and aquatic product production personnel.

[0035] Preferably, the process of identifying the user's occupation type is as follows:

[0036] Based on the fundamental formulas of the gravity model and considering the factors influencing the user's occupation type, the variables involved in the calculation are determined to be distance, type weight, and environment. Distance is calculated using the decision region A. d The average distance between a given POI and each cycling endpoint in any valid cluster; the type weight is calculated based on the decision region A. d The ratio of the number of POIs mapped to a specific occupation type in the study area to the number of POIs mapped to that occupation type in the entire study area; the environment is calculated as the decision area A. d The ratio of the proportion of any one occupational type to the sum of the proportions of all five occupational types is given by the following formula:

[0037]

[0038]

[0039]

[0040] Among them, t ij For, P i and P j The populations of neighborhoods i and j are respectively, dij ρ is the distance between cell i and cell j, k is the model parameter; ρ is the type weight, number(occupation) i ) represents the number of POIs mapped to a specific occupation type within a user's decision region, sum(occupation i ) represents the number of POIs mapping to this occupation type across the entire decision region; C represents the environmental parameter, ρoccupation i The proportion of a certain occupation type in a user's decision region. The sum of the proportions of the five occupational types in a user's decision area;

[0041] Using valid shared bicycle riding destination data of a user combined with POI data, a gravity model is applied to calculate the probability that the user belongs to each of the following occupational types: professional technician, clerk and related personnel, business service personnel, production and transportation equipment operator and related personnel, and agricultural, forestry, animal husbandry, fishery, and aquatic product production personnel. Then, Bayesian rules are used to calculate the normalized probability of the user belonging to the above occupational types. The occupational type corresponding to the highest normalized probability is taken as the user's occupational type, and the identification result is output. The formula is as follows:

[0042]

[0043]

[0044] Among them, GD,P i The probability that a user belongs to a certain occupation type. P represents the average distance between a given point of interest (POI) within a user's decision region and all possible cycling destinations most likely to be their workplace; r P i D is the normalized probability that a user belongs to a certain occupation type. Let be the sum of the probabilities of a user being classified into 5 different occupational types;

[0045] If a user has two or more occupation types with the same normalized probability, the absolute values ​​of the single parameters are compared in the order of distance, type weight, and environment. The occupation type with the largest absolute value is the user's final occupation type. That is, when two or more occupation types have the same maximum value of the distance parameter, the absolute value of the type weight parameter is compared, and so on.

[0046] Secondly, in order to achieve the above objectives, this invention discloses a system for identifying the occupational type of shared bicycle commuter users, comprising:

[0047] Data processing module: Used to receive POI data, urban road network data and shared bicycle data, clean and filter the shared bicycle data to obtain valid shared bicycle data;

[0048] The first-stage filtering module is used to calculate the normalized parameters of users' use of shared bicycles during the morning rush hour on weekdays. It uses these normalized parameters to filter and extract the effective data of shared bicycles, thus obtaining the first set of effective data.

[0049] Secondary filtering module: Used to count the cumulative number of days users use shared bikes in a week, and to perform secondary filtering and extraction on the effective data of shared bikes in the time dimension to obtain secondary effective data of shared bikes;

[0050] The three-stage filtering module is used to statistically analyze the destination locations of users using shared bikes on weekdays. It performs three-stage filtering and extraction on the valid shared bike data in the spatial dimension to obtain the valid shared bike data in the three stages.

[0051] The identification module is used to determine the decision area of ​​a user's occupation type based on the maximum walking radius and urban road network data, determine the correspondence between POI data and occupation type, and identify the user's occupation type by using POI data and three valid shared bicycle data and the decision area of ​​the user's occupation type.

[0052] The beneficial effects of this invention are:

[0053] This invention provides a method for relatively accurately identifying the occupational types of users commuting using shared bicycles by employing multi-source LBS data. First, based on initial data screening, the maximum tolerable walking distance threshold from the starting point to the pick-up point / riding origin (O) is determined, and peak morning hours are identified to initially screen shared bicycle data for commuting trips. Second, effective work locations and all possible riding destinations are extracted through date labeling, database categorization, and other operations to further screen shared bicycle data for commuting trips. Based on this, all service areas centered on effective riding destinations and with a radius equal to the maximum tolerable walking distance from the parking point / riding destination (D) to the destination are merged to generate a decision area for determining the user's occupational type. Then, a mapping relationship between POI industry fields and occupational types is established. Finally, the occupational types of users commuting using shared bicycles are identified based on a gravity model and Bayesian rules. This invention can conveniently and accurately determine the occupational type of users who commute using shared bicycles, providing a foundation for the study of the spatiotemporal patterns of different shared bicycle travel activities. This will further optimize the scheduling and management of shared bicycles in specific spaces and time periods, as well as the planning and layout of related supporting facilities, providing strong technical support for the efficient use of urban space and the rational allocation of urban resources. Attached Figure Description

[0054] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0055] Figure 1 This is a schematic diagram of the method flow of the present invention;

[0056] Figure 2 This is a schematic diagram of the workflow of the present invention;

[0057] Figure 3 This invention is a line graph showing the "maximum tolerable walking distance - shared bicycle data frequency" from subway station entrances / exits, bus stops, and residences to bike pick-up points in MYXC Street, NJ City.

[0058] Figure 4 This invention is a histogram of "maximum walking radius - percentage of shared bicycle data" for MYXC Street in NJ City.

[0059] Figure 5 This is a schematic diagram of the decision area for identifying the occupational type of a valid user of a shared bicycle in MYXC Street, NJ City, according to the present invention.

[0060] Figure 6 This is a schematic diagram of the system structure of the present invention. Detailed Implementation

[0061] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0062] like Figure 1 As shown, a method for identifying the occupation type of shared bicycle commuter users includes the following steps:

[0063] Receive POI data, urban road network data, and shared bicycle data; clean and filter the shared bicycle data to obtain valid shared bicycle data.

[0064] It should be noted that the information in the shared bike data includes the order number, user number, vehicle number, start and end time of the ride, and geographical coordinates of the ride route. Then, the valid data of the shared bikes is cleaned and filtered. Specifically, the data with a ride time of no more than 1 hour and an average ride speed in the range of 1km / h-30km / h is considered valid data. Abnormal data caused by GNSS positioning offset of shared bikes, users forgetting to lock the bikes, etc. are removed.

[0065] In this embodiment, POI data and urban road network data of MYXC Street in NJ City in 2020 are obtained from the Gaode Map API, and usage data of shared bicycles on February 10 (Monday) and February 14 (Friday) in 2020 are obtained from a third-party company (including order number, vehicle number, user number, order start and end time, and shared bicycle riding trajectory data). The effective data of shared bicycles are initially cleaned and filtered, that is, data with riding time not exceeding 1 hour and average riding speed in the range of 1km / h-30km / h are extracted and retained for subsequent calculations.

[0066] Calculate the normalized parameters of users' use of shared bicycles during the morning rush hour on weekdays. Use the normalized parameters of users' use of shared bicycles during the morning rush hour on weekdays to filter and extract the effective data of shared bicycles, and obtain one set of effective data of shared bicycles.

[0067] Calculate the maximum tolerable walking distance threshold from the starting point to the pick-up point / cycling origin (O). Since commuting endpoints are generally residences and workplaces, and considering the daily commuting usage patterns of shared bikes, the cycling origin typically includes three possible locations: residence (for the entire journey), bus or subway (for transfers). According to relevant research reports, shared bikes primarily replace walking, with limited substitution for public transportation, and the use of bikes to connect to the subway is more widespread than bus use. Therefore, the order of shared bike data extraction is subway station entrance / exit → bus stop → residence. Let r be the threshold value for each location. M r B and r H Using subway station entrances / exits, bus stops, and residential buildings in open communities (or entrances / exits of gated communities) as basic units, multi-ring buffer zones are created. The frequency of shared bicycle riding start points within each buffer ring is then statistically analyzed. The distance corresponding to the inflection point of the frequency change of shared bicycle riding start points within a buffer ring is used as a threshold (R). M R B R H ), and extract the valid data of shared bicycles in sequence (taking MYXC Street in NJ City as an example, respectively using r M =200m, r B =100m and r HUsing 50m as the basic unit, the maximum tolerable walking distance thresholds from subway station entrances / exits, bus stops, and residences to car pick-up points are identified as 600m, 200m, and 100m, respectively. (See appendix) Figure 2 ).

[0068] Identify the morning rush hour. Since morning commutes typically involve time constraints for clocking in and out, while users have more free time after get off work, the purposes for using shared bikes in the evening are more diverse and complex than in the morning. Therefore, shared bike data during the morning rush hour is a more accurate reflection of shared bike usage during commutes. Based on everyday experience, the typical weekday morning rush hour is determined to be 7:00-10:00 AM. Then, extract valid data on shared bike order start times within this period.

[0069] The system counts the cumulative number of days users use shared bikes during a week, and then performs a secondary filtering and extraction of the valid shared bike data based on the time dimension to obtain secondary valid shared bike data.

[0070] Select any user and label the usage data of each shared bike for each of their 5 working days of the week with the date (e.g., Mon, Tue, Wed, Thu, Fri).

[0071] Generally, users work at least 3 days a week (i.e., production workers working 3 shifts, excluding freelancers). Accordingly, shared bike usage data should be recorded for at least 3 out of 5 working days a week. Therefore, users who have used shared bikes for 3 days or more are considered valid users, and their corresponding shared bike usage data is considered valid data.

[0072] Specifically, for any user, the cumulative number of days (t) of their use of shared bikes across 5 working days per week is counted, where the days of shared bike use can be non-consecutive and the shared bike numbers can be different; then, all users are iterated through and categorized into databases.

[0073] If t≤2, then remove the user and their shared bicycle usage data;

[0074] If t=3 or t=4, then the user and their shared bicycle usage data will be included in database 1 (DS1);

[0075] If t = 5, then the user and their shared bicycle usage data will be included in database 2 (DS2);

[0076] The system analyzes the destinations of users using shared bikes on weekdays throughout the week, and then performs a third round of filtering and extraction on the secondary shared bike data in terms of spatial dimensions to obtain the third round of valid shared bike data.

[0077] It should be noted that distance-based clustering identifies and categorizes all possible cycling destinations for any valid user at a given destination. For any user in DS1 and DS2, the distances (projected distances) between all their cycling destinations are calculated using spatial coordinate data. Clustering analysis is then performed using these distances as parameters, outputting a dataset {Cluster}, which represents the location data of all possible cycling destinations for each user at every destination. The output result formed from the shared bicycle usage data of all users in DS1 is {Cluster1} = {{C 11},{C 12},…{C 1i}}, and so on, the output results formed from the data of DS2.

[0078] If, for any valid user, the number of different date markers among all possible cycling destinations for a given destination exceeds half of the user's cumulative usage days, then that destination is considered one of the user's potential candidate work locations. This step aims to filter and extract the user's potential work locations and related cycling data. Specifically:

[0079] For any user in DS1, all possible cycling destinations {C 1i}, count the number of different date tags (n 1i ), keep n 1i Data with ≥2 elements and their corresponding user IDs form the dataset {Cluster1'}={{C 11 '},{C 12 '},…{C 1i '}};

[0080] For any user in DS2, all possible cycling destinations {C 2i}, count the number of different date tags (n 2i ), keep n 2i Clusters with a density of ≥3 and their corresponding user IDs form a dataset {Cluster2'} = {{C 21 '},{C 22 '},…{C 2i '}}.

[0081] The candidate work location with the most possible cycling destination data is identified as the most likely work location. For any set {C} in {Cluster1'}... 1i '} represents the user's possible work locations and all possible cycling destinations. The number N of cycling destinations included in this data is counted, forming {N1} = {N}. 11 N 12 ,…N 1i}. Perform pairwise comparisons on all elements in {N1}, search for the maximum value in {N1} (i.e., the element with the most cycling destinations), and set its corresponding {C}. 1i {} is identified as valid data for all cycling destinations most likely to be the workplace. Similarly, the same operation is performed on {Cluster2}.

[0082] Specifically, if a user's possible work locations and the number of all possible cycling destinations are the same (i.e., the maximum value of elements in {N1} or {N2} is two or more), then the sum of the distances between each pair of possible cycling destinations corresponding to the candidate work location is compared, and the candidate work location with the smallest sum of distances is taken as the most likely work location for that user. Then, relevant data (user ID, all associated cycling destination data, etc.) are extracted.

[0083] The decision area for determining a user's occupation type is determined using the maximum walking radius and urban road network data. It should be noted that in this embodiment, the maximum tolerable walking distance from the parking point / cycling endpoint (D) to the destination is calculated. Since dockless shared bicycles offer relatively flexible parking, users tend to park them near the geofence close to their destination. Therefore, the maximum tolerable walking distance from the shared bicycle cycling endpoint (parking point) to the destination is used as the maximum walking radius (r). By calculating the distance between the destination of a shared bicycle ride and its nearest neighboring Point of Interest (POI) data, and plotting a histogram of "maximum walking radius - percentage of shared bicycle data" (the change in the proportion of shared bicycle data where at least one POI can be found within a certain range to the total shared bicycle data), the specific value of the maximum walking radius is adaptively determined. This value should simultaneously meet the following two conditions: First, at least one POI data point can be found within a circular area centered on the shared bicycle ride destination and with radius r; second, the proportion of shared bicycle data meeting the above conditions to the total shared bicycle data is 90%. (Note that the calculation result of the maximum walking radius varies by location. Taking MYXC Street in NJ City as an example, the calculated maximum walking radius r = 100m. See...) Figure 3 ).

[0084] For any most probable work location, among all possible cycling endpoints, using one of these endpoints as the center and the maximum tolerable walking distance (r) from the parking point / cycling endpoint (D) to the destination as the radius, calculate the walking reachability based on the actual road network. Then, merge the walking reachability of all cycling endpoints to form the decision region A for the shared bicycle user's occupation type. d (See Figure 4 ).

[0085] The correspondence between POI data and occupation types is shown in Table 1:

[0086] Table 1. Correspondence between POI industry sector classification and occupation type

[0087]

[0088]

[0089] By using POI data, valid data from three shared bicycle trials, and the decision region of the user's occupation type, the user's occupation type can be identified.

[0090] S1: Determine the variables involved in the calculation and their calculation methods. Based on the fundamental formula of the gravity model (Eq.1), and considering factors that may influence the user's occupation type, the variables involved in the calculation are determined to include distance (d), type weight (ρ), and environment (C). The distance is calculated using the decision region (A). d The average distance between a given point of interest (POI) and every cycling destination in any valid cluster. The method for calculating the type weight is based on the decision region (A). d The number of POIs mapped to a specific occupation type in the data. i The total number of POIs (occupation points) mapped to this occupational type across the entire study area. i The ratio of (Eq.2) to the decision region; the environment is calculated as the decision area (A) d The ratio of the proportion of any one occupational type to the sum of the proportions of all five occupational types (Eq.3). The relevant formula is as follows:

[0091]

[0092]

[0093]

[0094] Among them, t ij For, P i and P j The populations of neighborhoods i and j are respectively, d ij ρ is the distance between cell i and cell j, k is the model parameter; ρ is the type weight, number(occupation) i ) represents the number of POIs mapped to a specific occupation type within a user's decision region, sum(occupation i ) represents the number of POIs mapping to this occupational type across the entire study area; C represents the environmental parameter, ρoccupation i The proportion of a certain occupation type in a user's decision region. This represents the sum of the weights of the five occupational types in a user's decision area.

[0095] S2: Using valid shared bicycle riding destination data for a user, combined with POI data, a gravity model is applied to calculate the probability that the user is a professional technician, clerk or related personnel, business service personnel, production and transportation equipment operator or related personnel, or agricultural, forestry, animal husbandry, fishery, or aquatic product production personnel (Eq.4). Then, Bayesian rules are used to calculate the normalized probability of the user belonging to the above occupational types (Eq.5). The occupational type corresponding to the highest normalized probability is taken as the user's occupational type, and the identification result is output. The relevant formulas are as follows:

[0096]

[0097]

[0098] Among them, GD,P i The probability that a user belongs to a certain occupation type. P represents the average distance between a given point of interest (POI) within a user's decision region and all possible cycling destinations most likely to be their workplace; r P i D is the normalized probability that a user belongs to a certain occupation type. Let be the sum of the probabilities of a user being one of the five occupational types.

[0099] Specifically, if a user has the same normalized probability for two or more occupation types, then the probability will be determined by distance. The absolute values ​​of individual parameters are compared in the order of type weight (ρ) and environment (C), and the occupation type with the largest absolute value is taken as the user's final occupation type. That is, when two or more occupation types have the same distance from the maximum value of the parameter, the absolute value of the type weight parameter is compared, and so on.

[0100] In the vicinity of an office area in MYXC Street, NJ City, the results were verified through on-site observation and interviews with residents. The accuracy rate was found to be approximately 67.9%, proving the necessity and effectiveness of this method.

[0101] Secondly, in order to achieve the above objectives, such as Figure 6 As shown, this invention discloses a system for identifying the occupational type of shared bicycle commuter users, comprising:

[0102] Data processing module: Used to receive POI data, urban road network data and shared bicycle data, clean and filter the shared bicycle data to obtain valid shared bicycle data;

[0103] The first-stage filtering module is used to calculate the normalized parameters of users' use of shared bicycles during the morning rush hour on weekdays. It uses these normalized parameters to filter and extract the effective data of shared bicycles, thus obtaining the first set of effective data.

[0104] Secondary filtering module: Used to count the cumulative number of days users use shared bikes in a week, and to perform secondary filtering and extraction on the effective data of shared bikes in the time dimension to obtain secondary effective data of shared bikes;

[0105] The three-stage filtering module is used to statistically analyze the destination locations of users using shared bikes on weekdays. It performs three-stage filtering and extraction on the valid shared bike data in the spatial dimension to obtain the valid shared bike data in the three stages.

[0106] The identification module is used to determine the decision area of ​​a user's occupation type based on the maximum walking radius and urban road network data, determine the correspondence between POI data and occupation type, and identify the user's occupation type by using POI data and three valid shared bicycle data and the decision area of ​​the user's occupation type.

[0107] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0108] The foregoing has shown and described the basic principles, main features, and advantages of this disclosure. Those skilled in the art should understand that this disclosure is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of this disclosure. Various changes and modifications can be made to this disclosure without departing from its spirit and scope, and all such changes and modifications fall within the scope of this disclosure as claimed.

Claims

1. A method for identifying the occupational type of shared bicycle commuter users, characterized in that, The method includes the following steps: Receive POI data, urban road network data, and shared bicycle data; clean and filter the shared bicycle data to obtain valid shared bicycle data. Calculate the normalized parameters of users' use of shared bicycles during the morning rush hour on weekdays. Use the normalized parameters of users' use of shared bicycles during the morning rush hour on weekdays to filter and extract the effective data of shared bicycles, and obtain one set of effective data of shared bicycles. The system counts the cumulative number of days users use shared bikes during a week, and then performs a secondary filtering and extraction of the valid shared bike data based on the time dimension to obtain secondary valid shared bike data. The system analyzes the destinations of users using shared bikes on weekdays throughout the week, and then performs a third round of filtering and extraction on the secondary shared bike data in terms of spatial dimensions to obtain the third round of valid shared bike data. Based on the maximum walking radius and combined with urban road network data, the decision area of ​​the user's occupation type is determined, the correspondence between POI data and occupation type is determined, and the user's occupation type is identified by using POI data and three valid shared bicycle data and combined with the decision area of ​​the user's occupation type. The process of defining the decision area for the user's occupation type is as follows: Calculate the maximum tolerable walking distance from the parking point / cycling endpoint D to the destination. This is the maximum walking radius. By calculating the distance between the shared bike's destination and its nearest neighboring Points of Interest (POIs), and plotting a histogram of the maximum walking radius versus the proportion of shared bike data, the maximum tolerable walking distance from the parking point / destination to the destination is adaptively determined. This should simultaneously meet the following two conditions: With the destination of the shared bicycle ride as the center, and... Within a circular area with a radius of , at least one POI can be found; Shared bicycle data that meets both of the above conditions accounts for 90% of the total shared bicycle data; For any most probable work location, among all possible cycling endpoints, take one of these endpoints as the center and calculate the maximum tolerable walking distance from the parking point / cycling endpoint D to the destination. Using the radius as the basis, calculate the walking reachability based on the actual road network, and then merge the walking reachability of all cycling destinations to form the decision area for the occupation type of shared bicycle users. ; The process for identifying a user's occupation type is as follows: Based on the fundamental formulas of the gravity model and considering the factors that influence the user's occupation type, the variables involved in the calculation are determined to include distance, type weight, and environment. Distance is calculated using the decision region. The average distance between a given POI and each cycling endpoint in any valid cluster; the type weight is calculated based on the decision region. The ratio of the number of POIs mapped to a specific occupation type in the study area to the number of POIs mapped to the same occupation type across the entire study area; the environment is calculated based on the decision-making region. The ratio of the proportion of any one occupational type to the sum of the proportions of all five occupational types is given by the following formula: in, and Each of the following is a residential community and population, For the community and The distance between them These are model parameters; For type proportion, The number of POIs mapped to a specific occupation type within a user's decision region. The number of POIs mapped to occupation types for the entire decision region; For environmental parameters, The proportion of a certain occupation type in a user's decision region. The sum of the proportions of the five occupational types in a user's decision area; Using valid shared bicycle riding destination data of a user combined with POI data, a gravity model is applied to calculate the probability that the user belongs to each of the following occupational types: professional and technical personnel, clerks and related personnel, business service personnel, production and transportation equipment operators and related personnel, and agricultural, forestry, animal husbandry, fishery, and aquatic product production personnel. Then, Bayesian rules are used to calculate the normalized probability of the user belonging to the above occupational types. The occupational type corresponding to the highest normalized probability is taken as the user's occupational type, and the identification result is output. The formula is as follows: in, The probability that a user belongs to a certain occupation type. The average distance between a user’s POI within their decision area and each of all possible cycling destinations most likely to be their workplace. Let be the normalized probability of a user belonging to a certain occupation type. Let be the sum of the probabilities of a user being classified into 5 different occupational types; If a user has two or more occupation types with the same normalized probability, the absolute values ​​of the single parameters are compared in the order of distance, type weight, and environment. The occupation type with the largest absolute value is the user's final occupation type. When the maximum value of the distance parameter is the same for two or more occupation types, the absolute value of the type weight parameter is compared, and so on.

2. The method for identifying the occupational type of shared bicycle commuter users according to claim 1, characterized in that, The shared bicycle data includes the order number, user number, vehicle number, start and end times of the ride, and the geographical coordinates of the ride route.

3. The method for identifying the occupational type of shared bicycle commuter users according to claim 1, characterized in that, The shared bicycle data is cleaned and filtered to select shared bicycle data with a riding time of no more than 1 hour and an average riding speed in the range of 1km / h-30km / h as valid shared bicycle data.

4. The method for identifying the occupational type of shared bicycle commuter users according to claim 1, characterized in that, The normal parameters for users using shared bikes during weekday morning rush hours include calculating the maximum tolerable walking distance threshold from the starting point to the pick-up point / riding start point and identifying the morning rush hour.

5. The method for identifying the occupational type of shared bicycle commuter users according to claim 1, characterized in that, The process of secondary filtering and extraction of effective data from a single shared bicycle transaction along the time dimension is as follows: Users who use shared bikes for 3 or more days out of a 5-day work week are considered valid users. Each valid shared bike usage session is deemed valid. The cumulative number of days of shared bike usage within a 5-day work week is defined as t. The usage dates are allowed to be non-consecutive, and different shared bike numbers are allowed. The following criteria are used for evaluation: like If the value is ≤2, then the corresponding valid data for one shared bicycle will be removed. like If the value is >2, then extract the corresponding valid data for one shared bicycle and mark it as valid data for a second shared bicycle.

6. The method for identifying the occupational type of shared bicycle commuter users according to claim 1, characterized in that, The process of performing three rounds of filtering and extraction on the effective data of secondary shared bicycles in the spatial dimension is as follows: For any valid user, identify and classify all possible cycling destinations for a given destination. For any user, use spatial coordinate data to calculate the distance between each pair of all cycling destinations for that user, and perform cluster analysis with distance as a parameter to output all possible cycling destination location data for each user's destination. If, for any valid user, the number of different date labels in all possible cycling destinations for a given destination exceeds half of the user's cumulative usage days, then the destination will be considered as one of the user's possible candidate work locations. The candidate work location with the most possible cycling destination data is identified as the most likely work location. For the user's possible work locations and all possible cycling destination data, the number of cycling destination data contained in them is counted. The destination with the most cycling destination data is extracted as the user's most likely work location. The corresponding secondary shared bike valid data is extracted and marked as tertiary shared bike valid data.

7. The method for identifying the occupational type of shared bicycle commuter users according to claim 6, characterized in that, If a user's possible work locations and the number of all possible cycling destinations have multiple maximum values, then the sum of the distances between each pair of possible cycling destinations for each candidate work location is compared, and the candidate work location with the smallest sum of distances is selected as the user's most likely work location.

8. A system for identifying the occupational type of shared bicycle commuter users, employing the method for identifying the occupational type of shared bicycle commuter users as described in any one of claims 1 to 7, characterized in that, include: Data processing module: Used to receive POI data, urban road network data and shared bicycle data, clean and filter the shared bicycle data to obtain valid shared bicycle data; The first-stage filtering module is used to calculate the normalized parameters of users' use of shared bicycles during the morning rush hour on weekdays. It uses these normalized parameters to filter and extract the effective data of shared bicycles, thus obtaining the first set of effective data. Secondary filtering module: Used to count the cumulative number of days users use shared bikes in a week, and to perform secondary filtering and extraction on the effective data of shared bikes in the time dimension to obtain secondary effective data of shared bikes; The three-stage filtering module is used to statistically analyze the destination locations of users using shared bikes on weekdays. It performs three-stage filtering and extraction on the valid shared bike data in the spatial dimension to obtain the valid shared bike data in the three stages. The identification module is used to determine the decision area of ​​a user's occupation type based on the maximum walking radius and urban road network data, determine the correspondence between POI data and occupation type, and identify the user's occupation type by using POI data and three valid shared bicycle data and the decision area of ​​the user's occupation type.