Process for grouping journeys traveled
The method efficiently groups journeys by combining spatial and temporal clustering to differentiate speeds and trajectories, addressing inefficiencies in existing technologies and enabling precise analysis of large datasets for transport mode, emissions, and speed determination.
Patent Information
- Application Number
- FR2023009371
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-09-06
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2043-09-06
AI Technical Summary
Existing methods for grouping journeys using GPS and mobile phone data are inefficient in distinguishing between journeys based on different speeds and modes of transport, require significant computing resources, and are not suitable for large datasets.
A method involving time-stamped positioning data acquisition, filtering out aberrant journeys, and grouping using an agglomerative spatiotemporal clustering technique that combines spatial and temporal distances to differentiate journey speeds and trajectories.
Effectively groups journeys to determine mode of transport, pollutant emissions, and vehicle speed with reduced computing time and resources, enabling precise analysis of large datasets.
Smart Images

Figure 00000024_0000 
Figure 00000025_0000 
Figure 00000025_0001
Abstract
Description
Title of the invention: Method for grouping journeys traveled Technical field
[0001] The present invention relates to the field of grouping journeys traveled, in particular with a view to determining the mode of transport, a quantity of pollutants emitted, a vehicle speed and / or a vehicle flow rate.
[0002] The field of transport is an essential aspect of daily life on which the global economy depends heavily. However, it also plays a major role in the degradation of air quality and the increase in CO2 emissions, particularly in urban areas. According to the World Health Organization (WHO), transport is one of the main sources of air pollution, and it is directly linked to multiple respiratory and cardiovascular diseases. In addition, global transport is responsible for approximately 24% of direct CO2 emissions from fuel combustion, of which almost three-quarters correspond to cars, trucks, buses and motorcycles, further aggravating the phenomenon of global warming. Other types of harmful emissions include nitrogen oxides (NOx), ground-level ozone (O3) and particulate matter (PM), which generally exceed the recommended limits.The importance of transport's role in air quality is also reflected in the growing number of policy actions taken in Europe in recent years, making it a central issue. Strategies to address this issue include: vehicle efficiency and low-carbon fuel standards, the development of low-emission vehicles, fiscal policies and tax regimes.
[0003] New mobility data ("Floating Car Data", which can be translated as tracking vehicle data, data from GPS-type geolocation sensors of mobile phones, telephone boundary data, etc.) present a potential that is still underexploited in the characterization of mobility behaviors in a territory. Indeed, the mobile phone has become one of the essential objects of everyday life. This device is, on the one hand, connected to the network, but more recent technologies allow, on the other hand, to be connected and detected by other devices, such as satellites via a geolocation system such as the Global Positioning System (GPS) protocol, wireless network terminals via the WIFI protocol or other mobile phones via the Bluetooth protocol.These connection data can be used to determine more or less precisely the position of the mobile device, and therefore a priori that of its owner, and can therefore be useful for characterizing mobility.
[0004] Telephone operators collect connection information to their mobile antennas for all users during data transfers via the mobile telephone network, such as calls, messages and internet connections, which we will call events hereinafter. These data, called Call Detail Records (CDR), contain for each recorded event, the user ID, the start and end dates of the event and the position of the antenna through which the information was transmitted. This type of data offers a very interesting population penetration rate, because today most of the population has a mobile phone. However, the generation of this data is often not regular since it depends on the activity of users on their mobile phone.It is thus observed that a good part of this data appears in bursts and during static phases. It is for this reason that this type of data is widely used for the estimation of origin-destination matrices of movements, or to determine the attraction zones of a territory in a dynamic way. Nevertheless, mobility being repeatable, it is possible to find recurring patterns in the mobility of users. Some works even aim to complete data in the dynamic phases by using the observed patterns. Few solutions aim to determine the trajectory taken by a user to go from a point of origin to a destination. Prior art
[0005] Regarding trajectory aggregation in general, there is a lot of work on GPS data that consists of finding in real time which groups of individuals are following the same trajectory to group them together, in a more Lagrangian point of view. Most solutions do not simply use GPS data, but a lot of other information that phones can provide them, such as the speed and instantaneous acceleration of individuals.
[0006] Regarding trajectory aggregation with mobile phone network data, Network Signal Data (NSD), which is generated as soon as a phone connects to an antenna to find network, is often preferred over CDR data. Since CDR data is a subset of NSD data, anything that can be done with CDR data can be done with NSD data. Origin-destination matrix estimation is allowed with NSD data.
[0007] Several solutions have been developed with the aim of categorizing journeys made from data measured using a mobile phone.
[0008] Patent application FR 3112398 (US 2022 / 0011123) relates to a grouping of trajectories into clusters (groupings) from GPS data, to enable to associate trajectories from the same cluster with the same mode of transport. To implement this grouping, a spatial distance is used, which does not allow discrimination between modes of transport that could use the same path, but at a different speed. In addition, the modified Fréchet distance described in this patent application is not suitable for processing a large set of journeys.
[0009] Patent application FR 3130488 relates to a method for determining a mode of transport using a mobile telephone network. This method, which does not involve grouping the journeys taken, requires significant computing time and substantial computing resources (memory and processor) for processing a large set of journeys taken.
[0010] The document: "Loïc Bonnetain, Angelo Furno, Nour-Eddin El Faouzi, Marco Fiore, Razvan Stanica, Zbigniew Smoreda, Cezary Ziemlicki, TRANSIT: Fine-grained human mobility trajectory inference at scale with mobile network signaling data, Transportation Research Part C: Emerging Technologies, 2021" describes a method based on a DBSCAN clustering strategy, which makes it possible to refine trajectories of the same user from NSD data, thanks to the repeatability of the trajectories. The goal is to have trajectories close to GPS data. At the end of the article, the authors superimpose trajectories and realize that the public transport and road networks stand out through this implementation. They use these operations to observe the attraction areas of certain zones.This method does not take into account the temporal aspect, which does not allow to differentiate between journeys using the same path at a different speed (for example due to a different mode of transport, or due to a congested road).
[0011] Patent application US10073908 allows spatio-temporal clustering of trajectories using a combination of functional analysis of the trajectories (temporal clustering), then spatial clustering of the identified functions. The main problem with this method lies in the computation time, the necessary computing resources (memory and processor) and in the parameterization of the Functional Data Analytic (FDA) methods for estimating a parametric function for each trajectory and then Principal Component Analysis (PCA) to arrive at a subset of parametric functions explaining the majority of the variance of each group of functions. This approach seems poorly suited to processing a large set of paths. Summary of the invention
[0012] The present invention aims to group together journeys traveled making it possible to dissociate the speeds of different journeys, with a large number of journeys, with a limited computing time and computing resources, in particular to determine the mode of transport, a quantity of pollutants emitted, a vehicle speed and / or a vehicle flow rate. For this, the invention relates to a method for grouping journeys, with a step of acquiring time-stamped positioning data of the journeys, a step of filtering out aberrant journeys, then a step of grouping (clustering) by an agglomerative grouping method based on a spatiotemporal distance. The agglomerative grouping method based on a spatiotemporal distance makes it possible, in a single step, to differentiate the speeds and trajectories of the journeys in the groupings, which in particular facilitates the determination of a mode of transport, the quantity of pollutants emitted, a speed or a flow rate.
[0013] The invention further relates to a method for determining a mode of transport, a method for determining a quantity of pollutants emitted, and a method for determining a vehicle speed or flow rate implementing such a method for grouping journeys.
[0014] The invention relates to a method for grouping journeys taken by at least one user by any means of transport, in which the following steps are implemented: a. Time-stamped positioning data of a plurality of journeys traveled are acquired; b. Eliminating aberrant paths from said plurality of acquired paths, by a filtering method applied to said time-stamped positioning data; and c. Grouping said plurality of acquired and filtered paths into at least two groupings using an agglomerative grouping method based on a spatiotemporal distance, said spatiotemporal distance being a weighted sum of a spatial distance between two paths and a temporal distance between two paths depending on the difference in duration between two paths, said spatial distance and said temporal distance being determined from said time-stamped positioning data, the groupings differentiating speeds and trajectories.
[0015] According to one embodiment, the method comprises a step of preprocessing said time-stamped positioning data to determine a vector having a predetermined number of path points.
[0016] According to one implementation, said spatial distance between two paths i and j is determined by means of the following formula: > / vz / / / ï. », \ / s. \ \ with path i, Xj the aspat\Ai> Aj ) ~ N ^^^haversine^ J' J jjc ) j path j, dspat the spatial distance, N the number of path points, dhaversine a distance geodesic between two points, ^X^, y.^ ) the preprocessed coordinates of the k-th path point of path i, Xj y the preprocessed coordinates of the k-th path point of path j.
[0017] Advantageously, said time distance between two paths i and j is determined by means of the following formula: rrv with path i, X j path j, dt time distance, ( Xp Xj ) = — 2' N the number of path points, the time of the last path point of path i, t,j the time of the first path point of path j, tj the time of the last path point of path j, tjj the time of the first path point of path j.
[0018] According to one aspect, said spatiotemporal distance is determined using the following formula: dspatjemp = dspat + X^with dSpatje}np the spatiotemporal distance, dt the temporal distance, h a weighting scalar.
[0019] According to one configuration, said time-stamped positioning data of said plurality of journeys are acquired from measurements using a geolocation device or from measurements of connection to a telephone network.
[0020] According to one embodiment option, the method comprises a step prior to step b) of filtering in which journeys having the same geographical area of origin and the same geographical area of destination are selected, and the filtering and grouping steps are applied to the selected journeys.
[0021] Advantageously, the filtering and grouping steps are repeated for several selections of journeys with different geographical areas of origin and / or destination.
[0022] According to one embodiment, said aberrant paths are filtered by means of a DBSCAN type data partitioning method, said aberrant paths being the paths not grouped by said DBSCAN type data partitioning method.
[0023] The invention also relates to a method for determining a mode of transport of a plurality of journeys, in which the following steps are implemented: a. The grouping of each path is determined by means of the method according to one of the preceding characteristics; and b. A mode of transport is determined for each grouping based on the speed of the journeys of each grouping and / or by interpolation of a known mode of transport of at least one journey belonging to said grouping and / or based on the trajectory of said grouping; and c. Each journey of said group is assigned the determined mode of transport of said group.
[0024] Furthermore, the invention relates to a method for determining, for a plurality of journeys, an average speed and / or a median speed and / or a flow rate of vehicles, in which the following steps are implemented: a. The grouping of each path is determined by means of the method according to one of the preceding characteristics; and b. The said average and / or median speed and / or the flow of vehicles of each grouping is determined by means of the said time-stamped positioning data of the journeys belonging to the said grouping.
[0025] Furthermore, the invention relates to a method for determining a quantity of at least one pollutant emitted by a plurality of paths, in which the following steps are implemented: a. The grouping of each path is determined by means of the method according to one of the preceding characteristics; and b. For each grouping, a pollutant emissions model is applied, which links the speed and trajectory of said grouping to a quantity of at least one pollutant emitted; and c. The quantity of at least one pollutant emitted by said plurality of paths is determined by means of a weighted sum of said quantities of pollutant emitted for each grouping.
[0026] Other characteristics and advantages of the method according to the invention will appear on reading the following description of non-limiting examples of embodiments, with reference to the appended figures described below. List of figures
[0027] [Fig.l]
[0028] [Fig.l] illustrates the steps of the method for grouping paths according to a first embodiment of the invention.
[0029] [Fig.2]
[0030] [Fig.2] illustrates the steps of the method for grouping paths according to a second embodiment of the invention.
[0031] [Fig.3]
[0032] [Fig.3] illustrates the steps of the method for grouping paths according to a third embodiment of the invention.
[0033] [Fig.4]
[0034] [Fig.4] illustrates the steps of the method for determining a mode of transport according to an implementation of the invention.
[0035] [Fig.5]
[0036] [Fig.5] illustrates the steps of the method for determining a quantity of pollutants emitted according to an implementation of the invention.
[0037] [Fig.6]
[0038] [Fig.6] illustrates the steps of the method for determining a speed or a flow rate according to an implementation of the invention.
[0039] [Fig.7]
[0040] [Fig.7] illustrates a dendrogram for an example implementation of the agglomerative clustering step.
[0041] [Fig.8]
[0042] [Fig.8] illustrates, for an example, a set of paths between an origin zone and a destination zone.
[0043] [Fig.9]
[0044] [Fig.9] illustrates for the example of [Fig.8] the aberrant paths identified by the filtering step according to an embodiment of the invention.
[0045] [Fig. 10]
[0046] [Fig. 10] illustrates for the example of figures 8 and 9 the trajectories of the groupings formed by the grouping method according to an embodiment of the invention.
[0047] [Fig. 11]
[0048] [Fig. 11] illustrates for the example of figures 8 to 10 the curves of the distance as a function of time for the groupings formed by the grouping method according to an embodiment of the invention. Description of the embodiments
[0049] The present invention relates to a method for grouping (clustering) journeys traveled. A journey traveled is a trip made by a user between an origin and a destination with any means of locomotion (vehicle, bicycle, public transport: bus, tram, train, metro, walking, etc.). The grouping concerns the formation of groups (clusters) of journeys having similarities (similar attributes) among the plurality of journeys, in this case similarities in terms of speed of movement and in terms of trajectories (paths taken). According to one embodiment of the invention, the grouping can be used to determine the mode(s) of transport (i.e. the means(s) of locomotion). Alternatively or additionally, the grouping can be used to determine a quantity of pollutants emitted.Alternatively or cumulatively, the grouping can also be used to determine a vehicle speed and / or flow rate. In addition, the grouping can be used to determine users' preferred routes and roads, particularly for the purpose of assisting navigation (determining a route to be taken for a user).
[0050] According to a first embodiment, the method can implement the following steps:
[0051] 1. Acquisition of time-stamped positioning data
[0052] 4. Filtering of aberrant paths
[0053] 5. Grouping of journeys
[0054] These steps will be detailed in the remainder of the description. These steps 4 and 5 can be implemented by computer means, in particular a computer or a server, comprising at least one processor and a computer memory.
[0055] [Fig.l] illustrates, schematically and in a non-limiting manner, the steps of the path grouping method according to the first embodiment. Firstly, time-stamped positioning data of the plurality of paths are acquired ACQ. Then, the aberrant paths are removed by a FIL filtering method. Finally, the paths are grouped by an agglomerative grouping method based on a spatiotemporal distance CST.
[0056] According to one embodiment of the invention, the method may comprise a step of preprocessing the time-stamped positioning data, for a homogeneous representation of all the paths, which facilitates the calculation of the spatiotemporal distance. Advantageously, this preprocessing may consist of the formation of a vector having a predetermined number of path points and identical for all the paths.
[0057] This second embodiment then comprises the following steps:
[0058] 1. Acquisition of time-stamped positioning data
[0059] 3. Preprocessing of time-stamped positioning data
[0060] 4. Filtering of aberrant paths
[0061] 5. Grouping of journeys
[0062] These steps will be detailed in the remainder of the description. These steps 3 to 5 can be implemented by computer means, in particular a computer or a server, comprising at least one processor and a computer memory.
[0063] [Fig. 2] illustrates, schematically and in a non-limiting manner, the steps of the path grouping method according to the second embodiment of the invention. The steps identical to [Fig. 1] are not detailed again. The method further comprises a pre-processing step PRE before the step of removing aberrant paths by FIL filtering.
[0064] According to one embodiment of the invention, the method may further comprise a step of selecting routes between a departure zone (also called origin zone) and an arrival zone (also called destination zone). In this way, the filtering and grouping steps are implemented for the selection of routes, which reduces the volume of routes processed, and consequently reduces the calculation time and the requirements for computer resources (memory and processor). For this embodiment, the different steps can be repeated for several route selections, for example for several departure zones and / or for several arrival zones. Thus, groupings are constructed for each route selection.
[0065] This third embodiment then comprises the following steps:
[0066] 1. Acquisition of time-stamped positioning data
[0067] 2. Selection of routes
[0068] 4. Filtering of aberrant paths
[0069] 5. Grouping of journeys
[0070] These steps will be detailed in the remainder of the description. These steps 2, 4 and 5 can be implemented by computer means, in particular a computer or a server, comprising at least one processor and a computer memory.
[0071] [Fig. 3] illustrates, schematically and in a non-limiting manner, the steps of the path grouping method according to the third embodiment of the invention (including the second embodiment of the invention of [Fig. 2]). The steps identical to [Fig. 2] are not detailed again. The method further comprises a step of selecting paths O / D before the pre-processing step PRE. In addition, the dotted arrow indicates the repetition of steps 2 to 5 for several path selections.
[0072] 1. Acquisition of time-stamped positioning data
[0073] In this step, time-stamped positioning data of a plurality of paths traveled are acquired. The time-stamped positioning data includes the location in space of points belonging to the trajectory / path as well as the time associated (timestamp) with this location.
[0074] According to one embodiment, the time-stamped positioning data may be acquired by measurements using a geolocation system, for example GPS for "Global Positioning System" which can be translated as global positioning system, or the Galileo system, or any similar system. The geolocation system may in particular be included in a smartphone or a connected object. Alternatively, the time-stamped positioning data may be acquired from connection data to a telephone network. This may be NSD or CDR data, corresponding in particular to the location of the antennas to which a mobile telephone was connected during the journey. For this type of data, the recorded geographical coordinates do not correspond to the actual positions of a user but rather to those of the antennas to which their mobile telephone was connected during a telephone event (call, SMS, internet connection).These paths can also contain artifacts such as echo phenomena between antennas when the phone connects to several antennas at the same time. Thus, . for these different achievements, this measurement step can be implemented using a mobile phone. 2. Route selection
[0075] During this optional step, journeys are selected from the plurality of journeys acquired in step 1. The selection of journeys concerns the collection of journeys having the same departure zone (also called geographical zone of origin) and the same arrival zone (also called geographical zone of destination). A zone or geographical zone is a given territory, for example a region, a department, a community of communes, a city, a district of a city, etc. A geographical zone therefore comprises a plurality of points (respectively a plurality of points of origin and a plurality of points of destination of journeys). By way of example only, it may involve determining journeys between a geographical residential zone and a geographical zone of professional activities.The geographical area of origin corresponds to the territory from which a trip begins, and the geographical area of destination corresponds to the territory from which a trip ends. For this step, we can select all the trips whose first position is in zone A (origin zone), and the last position is in zone B (destination zone).
[0076] Following this selection of journeys, the filtering and grouping steps can be implemented for the selection of journeys, which favors the number of journeys processed, and consequently reduces the calculation time and the requirements for computer resources (memory and processor). When this optional step is implemented, the different steps can be repeated for several selections of journeys, for example for several departure zones and / or for several arrival zones. Thus, groupings are constructed for each selection of journeys, therefore for each pair of origin zone - destination zone.
[0077] 3. Preprocessing of time-stamped positioning data
[0078] During this optional step, a preprocessing of the time-stamped positioning data acquired in step 1 is carried out, possibly only for the selection of paths carried out in step 2. For this preprocessing, the time-stamped positioning data of each path are converted into a vector having a predetermined number of path points. In other words, the time-stamped positioning data are converted into a vector with an identical number of path points for all the paths. This preprocessing promotes, thanks to this homogeneity, the implementation of the following steps, by limiting the calculation time as well as the computer resources (memory and processor) required. Indeed, each path, depending on its duration and the measurement mode, can have a different number of time-stamped positioning data.
[0079] According to one embodiment, the number of path points (after preprocessing) can be between 5 and 100, and preferably between 10 and 50. Thus, a good compromise is obtained between precision of the representation of the paths, calculation time and necessary computing resources.
[0080] For a journey, the time-stamped positioning data acquired in step 1 can be noted as follows: , ._ [, . x ( . i] where trace is the vector of a path, K the trace-- ..., x^, y^ J number of path positioning data, x; the longitude of point i, y; the latitude of point i, and f the time of point i.
[0081] According to an exemplary implementation, the preprocessing may consist of an interpolation of the time-stamped positioning data onto a preprocessing time vector of predefined length (number of path points), denoted with N the number of path points, where the instants are distributed linearly between 6 and 1%, such that:
[0082] = tl
[0083] tK = tN
[0084] 7
[0085] Then, we can implement the interpolation by forming the new traceinterpolated vector for each path in the following way:
[0086] traceitüerpoite := [ ( x (, J, ), ..., ( xM tN) ]
[0087] Where (¾ is an interpolation of the longitude and latitude of the trace vector at time 6. The interpolation for this preprocessing can be a linear interpolation between the two points of the acquired time-stamped positioning data which surround the time considered during the preprocessing. As an example, for an instant ta of the preprocessing time vector between the acquisition times f and ti+i, we can write respectively for the longitude and for the latitude:
[0088] ~
[0089] ~ _ 3^2^ a fi+ 4. Filtering outlier paths
[0090] During this step, aberrant paths are removed from among the paths acquired in step 1, possibly selected in step 2, and / or possibly preprocessed in step 3, by means of filtering the aberrant paths. An aberrant path is, for example, a path for which there have been errors in measuring the time-stamped positioning data, or which includes portions of aberrant trajectories: including large detours, loops, etc. This step facilitates spatiotemporal grouping during step 5.
[0091] According to one embodiment of the invention, aberrant paths can be identified by a data partitioning method based on a distance between the paths, in particular the DBSCAN method (from the English "density-based spatial clustering of applications with noise"), the aberrant paths being those which do not belong to a grouping formed by the data partitioning method. Other similar methods can be implemented for filtering aberrant paths.
[0092] The DBSCAN method allows to group sets of points into clusters in a hyperspace according to the following rules: • For a new point (in this case a path) Xi which is not assigned to any cluster, we look to see if there are points belonging to a cluster already identified within a distance e of Xj • If yes, X{ belongs to the nearest cluster as well as all points at a distance less than e from Xt • If not, we look at how many points there are located at a distance less than 6 from Xj • If there are more than a new cluster is created and all points at a distance less than é from Xj are assigned to the new cluster, • If there are fewer than s, Xj is assigned to the outlier group and can still be assigned to a cluster as long as points located at a distance less than s from Xj are neither assigned to a cluster nor to the outlier group.
[0093] The hyperparameters for tuning a DBSCAN data partitioning method are therefore e, which defines the minimum inter-cluster distance, nmin which defines the minimum number of samples present in a cluster and the distance measure used. In simplified terms, if we decrease the value of £, we increase the number of clusters and the number of outliers, if we decrease the value of nmin, we increase the number of clusters and we decrease the number of outliers. The DBSCAN method is interesting because it does not require an a priori number of clusters to be found, and it allows us to group everything that really resembles each other into clusters, as long as there is a certain density in the point cloud, and to separate what does not resemble anything else into outliers.
[0094] According to an embodiment option of the invention, in order to find the optimal values of the parameters of the DBSCAN method, one can opt for the silhouette score. The latter ensures cohesion of the traces within each cluster and separation from other clusters. The silhouette score values vary from -1 to 1. A value close to 1 indicates that the traces are closer to the traces of their cluster than to the traces forming the neighboring clusters. Conversely, a value close to -1 may reveal that some traces have been assigned to the wrong cluster. The silhouette score is calculated from the average intra-cluster distance (a) and the average distance between the closest clusters (b). It is expressed by:
[0095] Silhouette Score =
[0096] Such that, a is the average distance between each point within a cluster and b is the distance between a point in a cluster and the nearest cluster of which the point is not a part. In summary, the silhouette score provides a quantitative assessment of the quality of clustering and allows for clusters with good internal cohesion and separation from other clusters.
[0097] Then, the evaluation of the two parameters £ and nmin can be done in two stages. First, we can find the optimal value of e which corresponds to the best silhouette score by temporarily fixing a value of nmin of the same order of magnitude as what we expect to have. The value e found is used to calibrate in order to obtain the highest possible silhouette score.
[0098] According to an implementation of the invention, to efficiently group the traces according to their spatial similarity and consequently identify aberrant paths, one can use the Fréchet distance which gives a low distance to two paths which are very close to each other throughout the path, and a high distance if the traces move away from each other during the path, even if they are very close during almost the entire path. This distance is often illustrated as the minimum leash distance necessary for a master walking his dog where one trace is the path of the master and the other trace is the path of the dog.
[0099] For example, to operate the DBSCAN algorithm, the Fréchet distance between each pair of paths of the plurality of paths (possibly preprocessed and possibly selected) can be calculated. The Fréchet distance can be calculated using dynamic programming or can be approximated by a neural network (as described for example in the patent application whose filing number is: FR 2307300). 5. Grouping of journeys
[0100] In this step, the paths filtered in step 4 (without the aberrant paths) are grouped into at least two groups (clusters). In other words, at least two clusters are formed to classify the paths. For this, an agglomerative clustering method is implemented, which is a hierarchical clustering method which merges clusters in order of their degree of similarity until reaching a predetermined number of groupings, denoted here Nciustagg. This agglomerative clustering is illustrated in Figure 7, which is a dendrogram of a data set. A dendrogram is a diagram which on the abscissa has the indices N of the samples of the database (the numbers in parentheses indicate the number of data in a node, and the numbers without parentheses indicate the reference of the data), and on the ordinate the minimum distance D between the two clusters, with a tree which hierarchically connects the clusters. On the tree, the clusters closest to each other are merged at their level of minimum distance, starting from cluster with a single sample and going up to a single cluster grouping all the data.The dendrogram is a tool that provides a visual assessment of the number of relevant clusters to consider, and is the basis of hierarchical clustering methods. Agglomerative clustering consists of using the dendrogram to merge clusters one after the other until the required number of clusters is obtained ^clustMgg- .
[0101] According to the invention, the agglomerative clustering method is based on a spatiotemporal distance, which is a weighted sum of a spatial distance between two paths, and a so-called temporal distance which is a function of the difference in duration of the two paths. The spatial distance and the temporal distance are determined from the acquired time-stamped positioning data. Taking into account spatial and temporal data makes it possible to form, in a single step, clusters differentiating paths of different speeds and different trajectories.In other words, each grouping includes journeys taking similar trajectories at similar speeds, which makes it possible to distinguish a bicycle journey from a car journey in congested traffic which would take the same time, but which are not carried out at the same pace (the car will go quickly in some places and very slowly in others, whereas the bicycle will be much more regular).
[0102] According to one embodiment, the spatiotemporal distance can be determined using the following formula:
[0103] dspattemp — dspat + Àd(with dspatfemp the spatiotemporal distance, dt the temporal distance, 2 a weighting scalar, which emphasizes the importance of travel time.
[0104] The hyperparameters to be set for the agglomerative clustering method are the number of required clusters N^^gg and the weighting 2 which allows to give more or less importance to the travel time in the clustering. The choice of parameter 2 can depend on the type of the desired grouped information (clustering mainly temporal, spatial, or spatio-temporal).
[0105] As a non-limiting example, we can choose weighting 2 between 106 and 10 *. For each value of the weighting, we can then determine a silhouette score (as a reminder, calculated from the average intra-cluster distance and the average distance between the closest clusters) calculated for Nciust^gg = 1 to 40 groupings. The combination of and Ncilist4tgg giving the best silhouette score can then be chosen.
[0106] For the embodiment in which the time-stamped positioning data is preprocessed, the spatial distance can be determined using the following formula:
[0107] Xj) = (¾. ^), (^, ^.) )with x' le11 x'le path j, dspat the spatial distance, N the number of path points, dhaversine a geodesic distance between two points, y^) the coordinates of the k-th path point of path i, ^Xj^, ÿj^) 'cs coordinates of the k-th path point of path j.
[0108] For the embodiment in which the time-stamped positioning data is preprocessed, the time distance can be determined using the following formula:
[0109] with X; the path i, Xj the path j, dt the distance «r ( Xb Xj ) = — y — temporal, N the number of path points, the time of the last path point of path i, tu the time of the first path point of path j, tj^ the time of the last path point of path j, t^ the time of the first path point of path j.
[0110] For the embodiment for which the groupings are constructed for a selection of paths between an origin zone and a destination zone, steps 2 to 5 can be repeated for at least a second selection of paths, which has a different origin zone and / or destination zone.
[0111] Furthermore, the invention relates to a method for determining a mode of transport of a plurality of journeys. For this method, the following steps are implemented: - The grouping of each journey is determined by means of the grouping method according to any of the variants or combinations of variants described above; - The mode of transport of each group is determined based on the speed of the journeys of each group and / or by interpolation of a mode of transport known from at least one journey belonging to the group (i.e. if the mode of transport of at least one journey belonging to the group is known). grouping, then we apply this mode of transport to all the journeys of the grouping) and / or depending on the trajectory of the grouping; and - Each trip in the group is assigned the group's mode of transport.
[0112] This method has the advantage of determining the mode of transport of a large number of journeys in a simple manner, with reduced calculation time, and with limited computing resource requirements (memory and processor). In addition, due to the grouping of journeys by trajectory and speed, the determination of the mode of transport is more precise.
[0113] These steps can be implemented by computer means, in particular a computer or a server, comprising at least one processor and a computer memory.
[0114] The step of determining the mode of transport of each grouping may consist of: - Compare the speed of each group's journeys (which can be obtained by the distance of the journey trajectory and by the average duration of the group's journeys, or by measuring the speed simultaneously with the measurements of time-stamped positioning data) with a speed representative of the modes of transport: for example walking about 5 km / h, running about 10 km / h, cycling about 20 km / h, motorized vehicle in town between 30 and 50 km / h, etc., and / or - Identify at least one journey in the group for which the mode of transport is known, for example by means of a mode of transport acquisition step, and / or - Identify the typology of the paths taken by the group's trajectory, for example: if it is a highway, the mode of transport is a motorized vehicle, if it is a path, the mode of transport is a soft mobility mode of transport.
[0115] [Fig.4] illustrates, schematically and in a non-limiting manner, the steps of the method for determining a mode of transport according to an embodiment, including the steps of the embodiment of [Fig.3]. The steps described for [Fig.3] are not detailed again. The method further comprises a step of determining the mode of transport MDT.
[0116] This figure includes the optional steps of route selection and preprocessing, however, the method of determining the mode of transport can be implemented without these steps.
[0117] The invention also relates to a method for determining an average speed and / or a median speed and / or a flow rate of vehicles of a plurality of journeys. For this method, the following steps are implemented: - The grouping of each path is determined by means of the method according to any of the variants or combinations of variants described above; - An average speed and / or a median speed and / or a flow rate of vehicles in each group is determined using the time-stamped positioning data of the journeys belonging to the group.
[0118] This method has the advantage of determining an average and / or median speed and / or a flow rate of vehicles for a large number of journeys in a simple manner, with reduced calculation time, and with limited computing resource requirements (memory and processor). In addition, due to the grouping of journeys by trajectory and speed, the determination of a speed or flow rate is more precise.
[0119] These steps can be implemented by computer means, in particular a computer or a server, comprising at least one processor and a computer memory.
[0120] According to an embodiment option, the speed of the journeys of each group can be obtained by the distance of the trajectory of the journey, and by the average duration of the journeys of the group.
[0121] Alternatively, a median or average of the speeds of the group's journeys may be implemented.
[0122] Alternatively or additionally, the flow rate can be determined as a function of the number of journeys per grouping and as a function of the times of the time-stamped positioning data (for example, it is possible to identify the times of day favorable to congestion or fluid traffic).
[0123] [Fig.6] illustrates, schematically and in a non-limiting manner, the steps of the method for determining a speed or flow rate according to an embodiment, including the steps of the embodiment of [Fig.3]. The steps described for [Fig.3] are not detailed again. The method further comprises a step of determining a speed or flow rate V / D.
[0124] This figure includes the optional steps of path selection and preprocessing, however, the method of determining a speed or flow rate can be implemented without these steps.
[0125] Furthermore, the invention relates to a method for determining a quantity of pollutants emitted by a plurality of paths. For this method, the following steps are implemented: - The grouping of each path is determined by means of the method according to any of the variants or combinations of variants described above; - For each grouping, a pollutant emissions model is applied, which links the speed and trajectory of the grouping to a quantity of at least one pollutant emitted, thus obtaining a quantity of pollutant emissions per grouping; and - The quantity of at least one pollutant emitted by said plurality of paths is determined by means of a weighted sum of said quantities of pollutant emitted for each grouping.
[0126] This method has the advantage of determining a quantity of pollutants emitted from a large number of journeys in a simple manner, with reduced calculation time, and with limited computing resource requirements (memory and processor). In addition, due to the grouping of journeys by trajectory and speed, the determination of a quantity of pollutants emitted is more precise.
[0127] These steps can be implemented by computer means, in particular a computer or a server, comprising at least one processor and a computer memory.
[0128] The pollutant emissions model can be written in particular in the form:
[0129] Qpol^ro “ f jgro) with Qpdgro a 9uantlté of pollutant emissions from the group considered, vnro a speed of the vehicles for the group considered, ^rajgro 'a trajectory of the group considered, f a function corresponding to the model.
[0130] The speed may be determined as indicated above, or by any analogous determination. The function f may be obtained from a vehicle dynamic model, or by machine learning, or by any analogous means.
[0131] Then, we can determine a quantity of pollutants by the set of paths using a formula of the type:
[0132] Spoliât ^1=0 ^^polgro l
[0133] With the total quantity of pollutants emitted, Nciustjagg the number of groupings, (!,i the weighting of grouping 1, Qp^gro / the quantity of pollutants emitted for a journey of grouping 1.
[0134] The weighting of the grouping is advantageously proportional to the number of journeys within the grouping.
[0135] [Fig.5] illustrates, schematically and in a non-limiting manner, the steps of the method for determining a quantity of pollutants emitted according to a method of embodiment, including the steps of the embodiment of [Fig.3]. The steps described for [Fig.3] are not re-detailed. The method further comprises a step of determining a quantity of pollutants emitted EPO.
[0136] This figure includes the optional steps of route selection and preprocessing, however, the method of determining the mode of transport can be implemented without these steps.
[0137] In addition, the method according to the invention may comprise a step of displaying the mode of transport and / or the speed and / or the flow rate of vehicles and / or the quantity of pollutants emitted. During this step, the mode of transport and / or the speed and / or the flow rate of vehicles and / or the quantity of pollutants emitted determined on a road map (on a road graph) are displayed. This display may take the form of a note or a color code or a thickness of representation of the road. This display may be carried out on board the vehicle: on the dashboard, on a portable autonomous device, such as a geolocation device (GPS type), a mobile phone (smartphone type). It is also possible to display the mode of transport and / or the speed and / or the flow rate of vehicles and / or the quantity of pollutants emitted on a website.In addition, the mode of transport and / or the speed and / or the flow of vehicles and / or the quantity of pollutants emitted can be shared with public authorities (e.g. road manager) and public works companies. Thus, public authorities and public works companies can determine roads with a high flow of vehicles, a high quantity of pollutants emitted, a low speed, few trips in soft mobility, and adapt the roads to users (e.g. creation of new lanes, modification of signage, etc.). Examples
[0138] The characteristics and advantages of the method according to the invention will appear more clearly on reading the example below.
[0139] For this example, we consider 80 journeys between an origin zone noted ZO and a destination zone noted ZD. For these 80 journeys, we acquire time-stamped positioning data measured by connection to the communication network, called CDR data.
[0140] [Fig.8] illustrates, schematically and in a non-limiting manner, the journeys considered between the origin zone ZO and the destination zone ZD. It can be seen that certain journeys have portions of the journey far from the vast majority of journeys.
[0141] A filtering step is applied, using the DBSCAN method based on a Fréchet distance, to identify aberrant paths. The best silhouette score obtained is 0.53 for the parameters e = 1 -98 and nmin = 4. This operation made it possible to filter the 6 most aberrant traces.
[0142] [Fig.9] illustrates, schematically and in a non-limiting manner, according to the representation of [Fig.8], the set of the 6 most aberrant TA traces. The non-aberrant TNA traces are schematically represented with a single trajectory.
[0143] We note that the filtering makes it possible to identify the most aberrant traces having portions of the path far from the vast majority of the paths.
[0144] We then apply the step of grouping the 74 valid (non-aberrant) paths by an agglomerative grouping method based on a spatiotemporal distance (as exemplified in step 5). The best silhouette score obtained is 0.77 for 2 = KL and Nciust^gg = 7. We thus form 7 groups.
[0145] [Fig. 10] illustrates, schematically and in a non-limiting manner, according to the representation of [Fig.8], the trajectories of the 7 groupings. It can be observed that some overlap.
[0146] [Fig. 11] illustrates for these 7 groupings, the distance traveled D in km as a function of time T in s. We can observe that the 7 groupings have different curves (therefore different speeds).
[0147] Consequently, the method according to the invention makes it possible to clearly differentiate the journeys according to their trajectory but also their speed.
Claims
Claims
1. Method for grouping journeys taken by at least one user by any means of transport, characterized in that the following steps are implemented by computer means:
2.
3. a. Time-stamped positioning data of a plurality of journeys traveled are acquired (ACQ) from measurements using a geolocation device or measurements of connection to a telephone network; b. The aberrant paths are eliminated from said plurality of acquired paths, by a filtering method (FIL) applied to said time-stamped positioning data; and c. Grouping (CST) said plurality of acquired and filtered paths into at least two groupings using an agglomerative grouping method based on a spatiotemporal distance, said spatiotemporal distance being a weighted sum of a spatial distance between two paths and a temporal distance between two paths depending on the difference in duration between two paths, said spatial distance and said temporal distance being determined from said time-stamped positioning data, the groupings differentiating the speeds and the trajectories. The method of claim 1, wherein the method comprises a step of preprocessing (PRE) said time-stamped positioning data to determine a vector having a predetermined number of path points. Method according to claim 2, in which said spatial distance between two paths i and j is determined by means of the following formula: dspai ( Xp Xj ) — jV ^-i^-^haveraine ( ( ) s ( path i, path j, dspat the spatial distance, N the number of path points, dhm,ersitie a geodesic distance between two points, \Xjp~ yik^ the preprocessed coordinates of the k-th path point of path i, (Xjk- y^ the preprocessed coordinates of the k-th path point of path j.
4. Method according to one of claims 2 or 3, in which said time distance between two paths i and j is determined by means of the following formula: with Xile traJet i' Xjle traiet j' the XbXj)~ 2 time distance, N the number of path points, tiN the time of the last path point of path i, the time of the first path point of path j, tjjy the time of the last path point of path j, tj^ the time of the first path point of path j.
5. Method according to one of the preceding claims, in which said spatiotemporal distance is determined by means of the following formula: d spatietnp — ^spat~^~ dgpafjtemp the spatiotemporal distance, dt the temporal distance, 2 a weighting scalar.
6. Method according to one of the preceding claims, in which said method comprises a step prior to the filtering step (FIL) in which paths having the same geographical area of origin (ZO) and the same geographical area of destination (ZD) are selected, and the filtering (FIL) and grouping (CST) steps are applied to the selected paths.
7. Method according to claim 6, in which the filtering (FIL) and grouping (CST) steps are repeated for several selections of journeys with different geographical areas of origin (ZO) and / or destination (ZD).
8. Method according to one of the preceding claims, in which said aberrant paths are filtered (FIL) by means of a DBSCAN type data partitioning method, said aberrant paths being the paths not grouped by said DBSCAN type data partitioning method.
9. Method for determining a mode of transport (MDT) of a plurality of journeys, characterized in that the following steps are implemented: a. The grouping of each journey is determined by means of the method according to one of the preceding claims; and b. A mode of transport is determined for each grouping as a function of the speed of the journeys of each grouping and / or by interpolation of a mode of transport known from at least one path belonging to said grouping and / or depending on the trajectory of said grouping; and c. Each journey of said group is assigned (MDT) the determined mode of transport of said group.
10. Method for determining, for a plurality of journeys, an average speed and / or a median speed and / or a vehicle flow rate (V / D), characterized in that the following steps are implemented: a. The grouping of each path is determined by means of the method according to one of claims 1 to 8; and b. The said average and / or median speed and / or vehicle flow (V / D) of each group is determined by means of the said time-stamped positioning data of the journeys belonging to the said group.
11. Method for determining a quantity of at least one pollutant emitted by a plurality of paths (EPO), characterized in that the following steps are implemented: a. The grouping of each path is determined by means of the method according to one of claims 1 to 8; and b. For each grouping, a pollutant emissions model is applied, which links the speed and trajectory of said grouping to a quantity of at least one pollutant emitted; and c. The quantity of at least one pollutant (EPO) emitted by said plurality of paths is determined by means of a weighted sum of said quantities of pollutant emitted for each grouping.