Method for determining geometry of lane boundary and method for determining position of lane centerline
By using the spatial index system to compare initial clustering and distance thresholds, the location of lane boundaries and lane centerlines in the digital map is determined, which solves the problems of high data processing cost and high computational cost in the prior art, and achieves more efficient and accurate data processing.
Patent Information
- Application Number
- CN202411495982.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-20
- Filing Date
- 2024-10-25
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art faces the high cost of data processing and high computational cost when determining the geometry of lane boundaries and the location of lane centerlines in digital maps, especially when processing data containing exceptions and noise.
By obtaining data sets with multiple individual observations, initial clustering is performed using a spatial indexing system to reduce the processing volume; then, by calculating the distance between data points and comparing distance thresholds, a cluster of data sets related to the same lane boundary or centerline is determined, and the geometry of the lane boundary and the position of the lane centerline is determined.
This method effectively reduces the computational work required to determine lane boundaries and lane centerlines, improves the efficiency and accuracy of data processing, and reduces the impact of noise and anomalies on the results.
Smart Images

Figure CN120182931A_ABST
Abstract
Description
Technical Field
[0001] The technology described herein relates to techniques for digital mapping, and in particular, to methods for determining the geometry of lane boundaries within a road segment in a geographic area represented by a digital map. The technology described herein also relates to methods for determining the position of a lane centerline within a road segment in a geographic area represented by a digital map. Background Art
[0002] Lane boundary markings are along the edges of one or more lanes of a travel path (e.g., on a road or highway), and are typically physically indicated by (dashed or solid) lines on the surface of the travel path. The lane centerline indicates the midpoint between the two lane boundaries of the corresponding lane. In contrast to lane boundaries, the lane centerline is typically not physically indicated on the surface of the travel path.
[0003] It is useful to be able to accurately represent the geometry (e.g., geographic location, orientation, and / or extent) of lane boundaries and / or the position of the lane centerline in a digital map (e.g., and in particular, for so-called high-definition (HD) maps for navigation of advanced driver assistance systems and autonomous vehicles). For example, in order to functionally support an advanced driver assistance system (ADAS) or autonomous driving, an HD map can provide not only a more detailed three-dimensional view of the road geometry (as may be required for more traditional basic navigation support), but also a more detailed three-dimensional view of the lane geometry including an indication of the lane centerline.
[0004] HD maps are typically captured using sensor arrays (e.g., LIDAR, digital cameras, etc.) that can be set on vehicles (which can be dedicated mapping fleet vehicles, but at least some of this data can also be "crowdsourced" from other road users). Other methods for obtaining data for generating an HD map utilize GNSS data (e.g., GPS data) to track vehicle trajectories. Summary of the Invention
[0005] To determine the geometry of lane boundaries, data related to the lane boundaries is collected and processed. A large amount of data can be collected for processing, and the collected data can include anomalies and "noise" that can introduce inaccuracies into the determination of the actual geometry of the lane boundaries. Similarly, to determine the position of the lane centerline, (a large amount of) data related to the lane centerline is collected and processed. The processing of such a large amount of data (including identifying and mitigating aberrations within the data) can be computationally expensive and time-consuming.
[0006] The applicant has realized that there is room for improvement in methods for determining the geometry of lane boundaries and methods for determining the position of lane boundaries within a road segment in a geographic area represented by a digital map.
[0007] Accordingly, in a first aspect, there is provided a method of determining the geometry of a lane boundary within a road segment in a geographic area represented by a digital map, wherein the lane boundary divides the road segment into a set of one or more lanes, the method comprising:
[0008] Obtaining a plurality of data sets representing multiple separate observations of lane boundaries within the road segment, wherein different ones of the plurality of data sets may represent observations of the same or different lane boundaries within the road segment, and wherein each data set comprises a series of corresponding data points spaced along the road segment and representing the positions of the lane boundaries;
[0009] Identifying an initial candidate group of data sets from the plurality of data sets representing separate observations of lane boundaries within the road segment, for which it is to be further determined whether the data sets should be clustered together as being related to the same first lane boundary;
[0010] wherein the identification of the candidate group of data sets is performed using a spatial indexing system, wherein the geographic area containing the road segment is subdivided into a plurality of tiles, each tile representing a corresponding sub-region of the geographic area, and wherein the positions of the tiles are spatially indexed relative to each other such that it can be determined which tiles are adjacent to each other, and wherein data sets are identified as part of the initial candidate group of data sets based on correspondences of data points of data sets that fit within the same tile or n-level adjacent tiles of the spatial indexing system;
[0011] Determining a cluster of data sets related to the same first lane boundary from within the identified initial candidate group of data sets by calculating corresponding distances between corresponding data points of different data sets and comparing the calculated distances with a distance threshold; and
[0012] Using the cluster of data sets that have been determined to be related to the same first lane boundary to determine the geometry of the first lane boundary.
[0013] The present disclosure generally relates to processing data for generating and / or updating digital maps. Such digital maps typically consist of a set of road segments (or other navigable elements) and nodes that suitably graphically represent connecting the road segments together to form a basic road network. Various other features associated with the road network may also be included within the digital map. This may include (by way of example and does include in the described examples) lane models that describe lane geometry. In order to generate digital maps with higher accuracy, it may be necessary and is typically the case that large amounts of data are processed in combination with the features that are to be included in the map. The first aspect of the present disclosure particularly relates to determining lane geometry for inclusion in such digital maps.
[0014] Specifically, as will be further explained below, a spatial indexing system is used to perform an initial clustering of the acquired data. This initial clustering can then reduce the amount of processing required, specifically by allowing data related to observations of lane boundaries that are (too) far from the lane boundaries for which the geometry is to be determined to be removed from consideration, thus obviating the need to further process such data. Once the initial clusters have been determined based on the spatial indexing system, the geometry of the lane boundaries is determined based on the clusters.
[0015] Thus, the data to be processed includes data related to observations of physical lane boundaries. The data can be collected in any suitable and desired manner. It is possible and preferred to collect the data during a journey along a section of road containing the lane boundaries to be observed. It is possible and preferred to use sensors mounted on a vehicle traveling along the section of road to collect the data. For example, in a preferred embodiment, depending on requirements, the sensor data can include data captured by a digital camera, a LiDAR sensor, etc. However, generally speaking, the methods disclosed herein can be used to process any suitable and desired data related to observations of physical lane boundaries and can be used to determine lane geometry.
[0016] The data to be processed includes a plurality of data sets. Each data set is preferably associated with a unique identifier such that it can be distinguished from each of the other data sets in the plurality of data sets.
[0017] Each data set in the plurality of data sets contains a series of data points related to the observed lane boundary. The series of data points can indicate the geographical location of the observed lane boundary in any suitable and desired manner. In the simplest case, a data set contains two data points: one data point indicates the start of a line (segment) and the other data point indicates the end of the line (segment). The two data points can be in the form of, for example, a pair of coordinates indicating the geographical locations of the start and end points of the line (segment). A data set preferably does contain more than two data points related to the observed lane boundary. In an embodiment, the data set includes data points indicating the start of a line, the extent (length) of the line, and the direction in which the line extends. In another embodiment, the data points within the data set indicate the segments of a polyline (e.g., the data set includes one data point for each vertex of the polyline). The position (geographical location) of the lane boundary or a portion thereof can be determined from the data set related to the lane boundary.
[0018] The applicant has recognized that data related to the observation of physical lane boundaries may often be "noisy" (e.g., distorted or inaccurate) and / or incomplete, e.g., due to sensor noise and / or (intermittent) occlusion of the sensors during data collection. To mitigate noisy or incomplete data collection, it is possible and preferably (e.g.) to perform multiple observations of the same lane boundary during multiple trips along a road segment. As a result of the multiple observations, multiple data sets related to the same lane boundary are collected.
[0019] Thus, multiple data sets representing multiple observations of lane boundaries within a road segment are obtained, where different ones of the multiple data sets represent observations of the same or different lane boundaries within the road segment.
[0020] Obtaining the multiple data sets can be accomplished in any suitable and desired manner. For example, in some embodiments, obtaining the multiple data sets may include obtaining data that has been previously recorded, and in such cases, the data may have been pre-conditioned or processed into a desired format for processing. However, in other embodiments, obtaining the multiple data sets includes obtaining "raw" (unprocessed) sensor data. In such cases, after obtaining the multiple raw sensor data sets, the data can be processed into a desired format suitable for further processing. Processing the data from a first format into a desired format can be referred to as pre-processing of the data. Similarly, even when the data has been previously obtained and subjected to some pre-processing, further processing can be performed to condition the data into a desired format.
[0021] Thus, in a preferred embodiment, obtaining the multiple data sets includes obtaining multiple data sets in a first format and processing the obtained multiple data sets into a desired format. As mentioned above, the first format can be the format of raw sensor data. However, various other arrangements are possible. Processing the obtained multiple data sets may include resampling the data (e.g.) to obtain uniform sampling.
[0022] It may further be desirable to (pre-)process the data to obtain more or fewer data points within one (and each) data set. For example, in the case where a data set includes only two data points (i.e., the start and end points marking the observation of a lane boundary), it may be desirable to include additional data points along the line such that the distance between consecutive data points is reduced. By increasing the number of data points in a data set, the accuracy of the comparison of the data set with another data set can be improved.
[0023] Conversely, in cases where the data set includes a large number of data points (depending on the context), it may be desirable to resample the data such that a smaller number of data points are obtained for further processing. By reducing the number of data points in the data set, the amount of processing required for the data set can be reduced. One (and each) data set can be resampled at regular or irregular intervals. By providing regular intervals between data points, comparison between different data sets can be facilitated, as will be described in more detail below. When a data set related to observations of lane boundaries has been resampled at regular intervals, it is also easier to determine the length of the observations.
[0024] In a preferred embodiment, obtaining a plurality of data sets includes, for each of the plurality of data sets: determining that the series of data points within the data set are non-equidistantly spaced along a road segment; and resampling the data set to obtain a data set including a series of data points that are equidistantly spaced along the road segment.
[0025] The resampling interval can be selected in any suitable and desired manner. In a preferred embodiment, the resampling interval is selected based on the length of one or more of the observations of the lane boundaries and / or is commensurate with the expected average distance between adjacent lane boundaries. In a particularly preferred embodiment, the resampling interval is between 3 meters and 8 meters, such as 5 meters. However, any other suitable resampling interval can be used as needed.
[0026] It should be understood that the number of data points in the (and each) data set after resampling can be the same as, less than, or greater than the number of data points in the (and each) data set before resampling.
[0027] Alternatively or additionally, it may be desirable to divide a data set into two or more new data sets. Thus, in a preferred embodiment, obtaining a plurality of data sets includes, for each of the plurality of data sets: determining that the observation of the lane boundary represented by the data set extends beyond a geographical limit; and dividing the data set to form at least two data sets, where one of the new data sets is entirely located on one side of the geographical limit and the other of the new data sets is entirely located on the other side of the geographical limit.
[0028] The geographical limit can be a Universal Transverse Mercator (UTM) boundary. In an embodiment, when it is determined that a data set represents an observation of a lane boundary that extends across a UTM boundary, the data set is divided into a first new data set including those data points representing positions on one side of the UTM boundary and a second new data set including those data points representing positions on the other side of the UTM boundary. For example, this can be useful when providing different data sets to different processors of a distributed processing system, as described further below.
[0029] Data points within each data set may and preferably are identifiable. In other words, each data point can be associated with a unique identifier that enables it to be distinguished from other data points within the same data set and from other data points within other data sets. The ability to uniquely identify each data point within multiple data sets enables different ones of the multiple data sets to be compared with one another (e.g., to determine whether different ones of the multiple data sets should be grouped together).
[0030] When processing multiple data sets, it is useful to be able to determine whether corresponding data points from each data set correspond to one another. Corresponding data points may be co-located within a particular sub-region of a geographic area. Data points from different data sets may be considered to correspond to one another if they are the closest corresponding data points to one another from each data set.
[0031] Thus, in a preferred embodiment, based on the distance between a first data point of a first data set and a first data point of a second data set being shorter than the distance between the first data point of the first data set and any other data point of the second data set, the first data point of the first data set is determined to correspond to the first data point of the second data set.
[0032] The applicant has realized that by performing an initial "coarse" grouping of data sets that may be related to the same first lane boundary and then determining clusters of data sets that are indeed related to the same lane boundary, the computational effort required to determine the geometry of the lane boundary can be reduced. By first removing from consideration those data sets that represent observations of other (distant) lane boundaries other than the first lane boundary (i.e., on a row-by-row basis), the number of data sets for which point-by-point distance calculations are performed can be minimized.
[0033] Once multiple data sets have been obtained (whether pre-processed or not), initial candidate groups of data sets are identified. The initial candidate groups include data sets for which it is to be further determined whether the data sets should be clustered together as being related to the same first lane boundary. A spatial indexing system is used to identify the initial candidate groups, where a geographic area containing road segments is subdivided into a plurality of tiles, where each tile represents a corresponding sub-region of the geographic area.
[0034] The positions of the tiles within the spatial indexing system are spatially indexed relative to one another such that it can be determined which tiles are adjacent to one another and which tiles are far from one another. The tiles can be spatially indexed in any suitable and desired manner. For example, each tile may be assigned a specific identifier that enables its position within a wider network of tiles to be determined.
[0035] The tiles of the spatial index system can be configured in any suitable and desired manner. In a preferred embodiment, the tiles cover the entire geographical area in a mosaic manner. The tiles of the spatial index system can be regular or irregular (i.e., they may all have the same size and / or shape, or different ones among the tiles may have different sizes and / or shapes), provided that they appropriately cover the entire geographical area. It may be particularly useful for all tiles to have the same size and shape. In this way, the centers of each tile are equidistant from each other, which can facilitate distance calculation.
[0036] Thus, in a preferred embodiment, the spatial index system includes a set of tiles that are regular polygons. In a particularly preferred embodiment, the set of tiles are regular hexagonal tiles.
[0037] The spatial index system can be a hierarchical spatial index system. The hierarchical spatial index system includes several different resolution "levels", each level including a set (grid) of tiles, where the tiles associated with one (and each) level of the hierarchy have different sizes (resolutions) from the tiles associated with another (and each other) level of the hierarchy.
[0038] In a preferred embodiment, the hierarchical spatial index system includes "parent" and "child" units, which are configured such that one (and each) parent (i.e., lower resolution) unit has a given number of children (i.e., higher resolution) units that (substantially) fit within the area of the parent unit. The number of child units within a parent unit may and preferably does depend on the shape of the units within the hierarchical spatial index system.
[0039] The orientation of the units of one (and each) level within the hierarchy may be the same or different from the orientation of the units of another (and each other) level within the hierarchy.
[0040] In a particularly preferred embodiment, the spatial index system includes an H3 index system. The H3 index system is a hierarchical spatial index system that includes hexagonal tiles, where each parent grid of the tiles is oriented differently from its corresponding child grid of the tiles such that seven tiles of the child grid substantially fit within one tile of the parent grid. However, various other arrangements will be possible.
[0041] The resolution (e.g., tile size) of the spatial index system can be set (selected) in any suitable and desired manner.
[0042] In a preferred embodiment, the tile size of the spatial index system is selected based on the resampling interval (i.e., the distance between resampled data points in the dataset). In a preferred embodiment, the tile size of the spatial index system is selected such that consecutive data points in the dataset are located (or at least typically likely to be located) in the same tile or n-level adjacent tiles in the spatial index system. The value of n can be set to any suitable and desired integer. In a particularly preferred embodiment, n is 1, 2, 3, or 4. That is, the resampling interval and tile size are preferably such that it can be expected that most consecutive data points in the dataset will be located in adjacent tiles or in tiles that are 2, 3, or 4 tiles apart (i.e., there are 1, 2, or 3 tiles between the two tiles in which consecutive data points in the dataset are located).
[0043] Based on the counterparts among the data points of the dataset that fit the same tile or n-level adjacent tiles of the spatial index system, the dataset is identified as part of an initial candidate group of datasets. In other words, datasets that do not include corresponding data points that fit the same tile or n-level adjacent tiles as one or more other datasets are excluded from the initial candidate group.
[0044] In a particularly preferred embodiment, if at least two corresponding data points from two datasets fit the corresponding same tile or n-level adjacent tiles of the spatial index system, then the two datasets are determined to form part of the initial candidate group. That is, the first data point of the first dataset and the corresponding first data point of the second dataset fit the first same tile or n-level adjacent tiles, and the second data point of the first dataset and the second data point of the second dataset fit the second same tile (which may be the same as or different from the first same tile) or n-level adjacent tiles of the spatial index system.
[0045] By forming the initial candidate group based on the spatial index system, those datasets associated with observations of lane boundaries that are (too) far from the lane boundaries for which the geometry is to be determined can be removed from consideration. The tiling of the spatial index system makes this initial filtering step faster and easier because it is not necessary to compare each dataset point-by-point to determine the initial candidate group; instead, the datasets are considered row-by-row and those lines that are determined to be far from the line (lane boundary) to be determined can be removed from further consideration.
[0046] In a preferred embodiment, the initial candidate group includes at least a determined number of datasets. In other words, the initial candidate group is not identified unless a determined number of datasets are available for inclusion in the initial candidate group. The determined number of datasets can be any suitable and desired number of datasets. The determined number may (and preferably does) vary according to requirements. In a preferred embodiment, the initial candidate group includes at least 5 datasets, preferably at least 8 datasets, preferably at least 10 datasets, or preferably more datasets.
[0047] Once the initial candidate group has been identified, clusters of data sets associated with the same first lane boundary are determined by calculating the corresponding distances between corresponding data points of different data sets and comparing the calculated distances with a distance threshold. It should be understood that the determined clusters may include some or all of the data sets from the initial candidate group.
[0048] The distances between corresponding data points can be calculated in any suitable and desired manner. In one embodiment, the calculated distance is the Euclidean distance. In another embodiment, the calculated distance can be the haversine distance.
[0049] To reduce the amount of processing required to determine clusters of data sets associated with the same first lane boundary from the initial candidate group, the directed Hausdorff distance between two data sets can be determined. That is, instead of determining the distance from the first data set to the second data set and then subsequently determining the distance from the second data set to the first data set, the distance between the two data sets is determined only once.
[0050] To this end, the unique identifiers of the data sets can be used to filter the data to be processed. In one embodiment, a tuple is formed for each pair of data sets for which the distance is to be calculated. Each tuple includes a left value corresponding to the unique identifier of the first data set of the pair of data sets and a right value corresponding to the unique identifier of the second data set of the pair of data sets. The tuples are filtered such that any tuple in which the value of the identifier on the left side of the tuple is greater than the identifier on the right side of the tuple is removed from consideration. In this way, it can be ensured that any corresponding pair of data sets is only compared once. Of course, it should be understood that in alternative embodiments, the filtering can be performed based on the right value of the tuple rather than the left value. By filtering the data in this way before calculating the corresponding distances, the duplication of processing can be minimized.
[0051] To calculate the distance between a pair of data sets, the corresponding distances between the corresponding data points within the data sets are calculated and the distance between the pair of data sets is calculated based on the distances between the corresponding points.
[0052] In one embodiment, only the distance between the first data point of the first data set and the first data point of the second data set is calculated. Then, the calculated distance is used as the distance between the two data sets.
[0053] However, it may be desirable to consider multiple corresponding data points within each data set (e.g., to improve the accuracy of the calculated distance). Thus, in a preferred embodiment, determining a cluster of data sets related to the same first lane boundary from within the identified initial candidate group includes: calculating the Euclidean distance between a first data point of a first data set from the initial candidate group and a corresponding first data point of a second data set from the initial candidate group; calculating the Euclidean distance between a second data point of the first data set and a corresponding second data point of the second data set; and when the calculated distances are below a desired comparison distance threshold, determining that the first data set and the second data set should be clustered together as related to the same first lane boundary.
[0054] It may be desirable to define a minimum number of data points for which distances should be calculated. Thus, in a particularly preferred embodiment, determining a cluster of data sets related to the same first lane boundary from within the identified initial candidate group includes: determining a minimum number of corresponding data points for which Euclidean distances are to be calculated based on the number of data points within the first data set and / or the second data set; calculating the Euclidean distances between the determined number of corresponding pairs of data points from the first data set and the second data set; and when all the calculated distances are below a desired comparison distance threshold, determining that the first data set and the second data set should be grouped together as related to the same first lane boundary.
[0055] In one embodiment, the maximum of the calculated distances between corresponding points is determined as the distance between the data sets. In another embodiment, the minimum of the calculated distances between corresponding points is determined as the distance between the data sets.
[0056] Once the distance between a pair of data sets has been calculated, the calculated distance is compared to a distance threshold to determine whether the pair of data sets should be grouped to form a cluster. This comparison can be made in any suitable and desired manner.
[0057] In one embodiment, data pairs for which the calculated distance reaches or is below the threshold are grouped to form a cluster, while data pairs for which the calculated distance exceeds the threshold are not grouped to form a cluster.
[0058] In a preferred embodiment, a cluster includes a minimum number of data sets. In other words, a cluster is determined only if it is determined that at least the minimum number of data sets should be grouped to form a cluster. Conversely, if it is determined that fewer than the minimum number of data sets should be grouped to form a cluster, then no cluster is formed. The minimum number of data sets can be any suitable and desired number. The minimum number can (and preferably is) set according to requirements. In a particularly preferred embodiment, the minimum number of data sets is 5, 8, 10 or more. By requiring a minimum number of data sets to form a cluster, artifacts caused by noise and / or other data aberrations can be reduced or avoided.
[0059] Once clusters of data sets associated with the same first lane boundary have been identified, the identified clusters can be used to determine the geometry of the lane boundary.
[0060] The geometry of the lane boundary can be determined in any suitable and desired manner. In a preferred embodiment, determining the geometry of the lane boundary includes: generating a bounding box encompassing the data points of all the data sets within the identified clusters; and using the generated bounding box to determine the geometry of the lane boundary. In a particularly preferred embodiment, using the generated bounding box to determine the geometry of the lane boundary includes: determining the center line of the bounding box; and using the determined center line to determine the geometry of the lane boundary.
[0061] In an embodiment, the method as described herein further includes updating the digital map to include the determined geometry of the first lane boundary.
[0062] A second aspect of the present disclosure particularly relates to determining the position of a lane center line for inclusion in a digital map. The techniques described above in connection with determining the geometry of a lane boundary can be applied in a similar manner to determining the position of a lane center line. Specifically, the applicant has realized that the described clustering techniques can be used to reduce the amount of processing required to determine the position of a lane center line.
[0063] Accordingly, in a second aspect, there is provided a method of determining the position of a lane center line within a section of a geographical area represented by a digital map, wherein the section is divided into a group of one or more lanes each delimited by two lane boundaries, and wherein each lane center line indicates the midpoint between the two lane boundaries of the corresponding one of the one or more lanes, the method comprising:
[0064] obtaining a plurality of data sets representing multiple individual vehicle trajectories of vehicles traveling along respective portions of the section, wherein different ones of the plurality of data sets may represent different vehicle trajectories of the same or different lanes along respective portions of the section, and wherein each data set contains a series of respective data points spaced along a portion of the section and representing the trajectory of a vehicle traveling in a particular lane along the portion of the section;
[0065] identifying an initial candidate group of data sets from the obtained plurality of data sets, for which it is to be further determined whether the data sets should be clustered together as being associated with the same first lane center line;
[0066] A spatial indexing system is used to perform identification of an initial candidate group of data sets, wherein a geographical area containing road segments is subdivided into a plurality of tiles, each tile representing a corresponding sub-area of the geographical area, and wherein the positions of the tiles are spatially indexed relative to each other such that it is possible to determine which tiles are adjacent to each other, and wherein data sets are identified as part of the initial candidate group of data sets based on counterparts of data points of the data sets that fit within the same tile or n-level adjacent tiles of the spatial indexing system;
[0067] A first cluster of data sets related to the same first lane centerline is determined from the data sets within the identified initial candidate group of data sets by calculating the respective distances between corresponding data points of different data sets and comparing the calculated distances with a distance threshold; and
[0068] The position of the lane centerline is determined using the first cluster of data sets that have been determined to be related to the same first lane centerline.
[0069] Thus, the data to be processed includes data related to the vehicle trajectories of vehicles traveling along separate portions of a road segment. The data can be collected in any suitable and desired manner. It is possible and preferred to collect the data during a journey along the road segment containing the lane centerline whose position is to be determined. The data can be and preferably is position data, such as GNSS data, and particularly preferably GPS data. However, generally speaking, the methods disclosed herein can be used to process any suitable and desired data related to vehicle trajectories and can be used to determine the position of a lane centerline.
[0070] The data to be processed includes a plurality of data sets. Each data set can be and preferably is associated with a unique identifier such that it can be distinguished from each of the other data sets within the plurality of data sets.
[0071] Each data set within the plurality of data sets contains a series of data points related to the trajectory of a vehicle along a portion of a road segment. The series of data points can indicate the vehicle trajectory in any suitable and desired manner. In the simplest instance, a data set contains two data points: one data point indicating the start of a line (segment) representing the vehicle trajectory and one data point indicating the end of the line (segment) representing the vehicle trajectory. The two data points can be, for example, in the form of a pair of coordinates indicating the geographical locations of the start and end points of the line (segment). A data set can and preferably does contain more than two data points related to the vehicle trajectory.
[0072] In an embodiment, the data set includes data points indicating the start of a line, the extent (length) of the line, and the direction in which the line extends. In another embodiment, the data points within the data set indicate segments of a multi-segment line (e.g., the data set includes one data point for each vertex of the multi-segment line). The position (geographical location) of the lane centerline or a portion thereof can be determined from data sets related to the vehicle trajectory.
[0073] It is possible and preferably (e.g.) to determine a plurality of trajectories of vehicles traveling along the same part of a section during a plurality of trips along (parts of) the section. As a result of the multiple determinations of the trajectories, a plurality of data sets related to the same lane are collected.
[0074] Thus, a plurality of data sets representing a plurality of vehicle trajectories within (parts of) the section are obtained. Obtaining the plurality of data sets can be done in any suitable and desired manner. For example, in some embodiments, obtaining the plurality of data sets may include obtaining data that has been previously recorded, and in such cases, the data may have been pre-conditioned or processed into a desired format for processing. However, in other embodiments, obtaining the plurality of data sets includes obtaining "raw" (unprocessed) data. In such cases, after obtaining the plurality of raw sensor data sets, the data can be processed into a desired format suitable for further processing. Processing the data from a first format into a desired format may be referred to as pre-processing of the data. Similarly, even when the data has been previously obtained and subjected to some pre-processing, further processing can be performed to condition the data into a desired format.
[0075] Thus, in a preferred embodiment, obtaining the plurality of data sets includes obtaining a plurality of data sets in a first format and processing the obtained plurality of data sets into a desired format. As mentioned above, the first format can be the format of raw data. However, various other arrangements are possible. Processing the obtained plurality of data sets may include resampling the data (e.g.) to obtain a uniform sampling.
[0076] It may be further desirable to (pre-)process the data to obtain more or fewer data points within one (and each) data set. For example, in the case where a data set includes only two data points (i.e., marking the start and end points of a vehicle trajectory), it may be desirable to include additional data points along the line such that the distance between consecutive data points is reduced. By increasing the number of data points in the data set, the accuracy of the comparison of the data set with another data set can be improved.
[0077] Conversely, in the case where a data set includes a large number of data points (depending on the context), it may be desirable to resample the data such that fewer data points are obtained for further processing. By reducing the number of data points in the data set, the amount of processing required for the data set can be reduced. One (and each) data set can be resampled at regular or irregular intervals. By providing a regular interval between data points, the comparison between different data sets can be facilitated, as will be described in more detail below.
[0078] In a preferred embodiment, obtaining a plurality of data sets includes, for each data set of the plurality of data sets: determining that the series of data points within the data set are non-uniformly spaced along a road segment; and resampling the data set to obtain a data set including a series of data points that are uniformly spaced along the road segment. The resampling interval can be selected in any suitable and desired manner. The resampling interval can be selected to be commensurate with the expected average distance between the centerline of the lane or adjacent lanes. In a particularly preferred embodiment, the resampling interval is between 3 meters and 8 meters, such as 5 meters. However, any other suitable resampling interval can be used as needed.
[0079] It should be understood that the number of data points in the (and each) data set after resampling can be the same as, less than, or greater than the number of data points in the (and each) data set before resampling.
[0080] The data points within each data set can and preferably are identifiable. In other words, each data point can be associated with a unique identifier that enables it to be distinguished from other data points within the same data set and from other data points within other data sets. The ability to uniquely identify each data point within a plurality of data sets allows different ones of the plurality of data sets to be compared with each other (e.g., to determine whether different ones of the plurality of data sets should be grouped together).
[0081] When processing a plurality of data sets, it is useful to be able to determine whether the corresponding data points from each data set correspond to each other. The corresponding data points can be co-located within a particular sub-region of a geographic area. If the data points from different data sets are the closest corresponding data points to each other from each data set, then they can be considered to correspond to each other.
[0082] Thus, in a preferred embodiment, based on the distance between a first data point of a first data set and a first data point of a second data set being shorter than the distance between the first data point of the first data set and any other data point of the second data set, the first data point of the first data set is determined to correspond to the first data point of the second data set.
[0083] The applicant has realized that by processing in parallel data related to different sub-regions of a geographic area represented by a digital map, the determination of the position of one or more lane centerlines within the geographic area can be accelerated.
[0084] To this end, the data set is grouped based on a spatial indexing system, where the geographical area containing the road segments is subdivided into a plurality of tiles, each tile representing a corresponding sub-region of the geographical area including a given portion of the road segment. Thus, in an embodiment, all data sets located within a given tile (i.e., a given portion of the road segment) are processed together. In the case where the data set corresponding to the vehicle trajectory extends across multiple tiles of the spatial indexing system, the original data set is subdivided into corresponding multiple data sets such that each of the multiple data sets is contained within a single tile of the spatial indexing system.
[0085] The positions of the tiles within the spatial indexing system are spatially indexed relative to each other such that it is possible to determine which tiles are adjacent to each other and which tiles are far from each other. The tiles can be spatially indexed in any suitable and desired manner. For example, each tile can be assigned a specific identifier, which enables the determination of the position of the tile within the wider tile network.
[0086] The tiles of the spatial indexing system can be configured in any suitable and desired manner. In a preferred embodiment, the tiles, for example, cover the entire geographical area in a mosaic manner. The tiles of the spatial indexing system can be regular or irregular (i.e., they may all have the same size and / or shape, or different ones of the tiles may have different sizes and / or shapes), as long as they appropriately cover the entire geographical area. It may be particularly useful for all tiles to have the same size and shape. In this way, the centers of each tile are equidistant from each other, which can facilitate distance calculation.
[0087] Thus, in a preferred embodiment, the spatial indexing system includes a set of tiles that are regular polygons. In a particularly preferred embodiment, the set of tiles is regular hexagonal tiles.
[0088] The spatial indexing system can be a hierarchical spatial indexing system. The hierarchical spatial indexing system includes several different resolution "levels", each level including a set (grid) of tiles, where the tiles associated with one (and each) level of the hierarchical structure have different sizes (resolutions) from the tiles associated with another (and each other) level of the hierarchical structure.
[0089] In a preferred embodiment, the hierarchical spatial indexing system includes "parent" and "child" units, which are configured such that one (and each) parent (i.e., lower resolution) unit has a given number of children (i.e., higher resolution) units within the area that (substantially) fits the parent unit. The number of child units within a parent unit may and preferably does depend on the shape of the units within the hierarchical spatial indexing system.
[0090] The orientation of the units of one (and each) level within the hierarchical structure may be the same or different from the orientation of the units within another (and each other) level of the hierarchical structure.
[0091] In a particularly preferred embodiment, the spatial indexing system includes an H3 indexing system. The H3 indexing system is a hierarchical spatial indexing system that includes hexagonal tiles, where each parent grid of a tile is oriented differently from the corresponding child grids of the tile such that seven tiles of the child grids substantially fit one tile of the parent grid. However, various other arrangements will be possible.
[0092] The resolution (e.g., tile size) of the spatial indexing system can be set (selected) in any suitable and desired manner.
[0093] In a preferred embodiment, the spatial indexing system is a hierarchical spatial indexing system that includes a first level with a first tile size and a second level with a second, smaller tile size, where each portion of a road segment corresponds to an area covered by a corresponding tile of the first level of the spatial indexing system, where obtaining a plurality of data sets includes determining for each data set one or more groups of data points, where each group of data points includes data points that fit the corresponding tile of the first level, and where the tiles used to identify an initial candidate group of data sets are tiles of the second level.
[0094] Preferably, the tile size of the tiles of the first level of the spatial indexing system is selected based on the size of the road. For example, the tile size of the tiles of the first level of the spatial indexing system can be selected such that the total width of the road segment (i.e., from the leftmost lane boundary of the road segment to the rightmost lane boundary) fits within a single tile. That is, if the total width of the road segment from the leftmost lane boundary of the leftmost lane to the rightmost lane boundary of the rightmost lane is 40 meters, then the tile size of the tiles of the first level of the spatial indexing system is selected such that the side length of each tile is at least 40 meters.
[0095] The spatial indexing system can include a set of tiles of regular polygons, such as a set of tiles of regular hexagons. Preferably, the spatial indexing system includes an H3 indexing system.
[0096] In the case where the spatial indexing system includes an H3 indexing system, the first level can be the H11 level and the second level can be the H14 level. Of course, it should be understood that other levels of the H3 indexing system can be used depending on the requirements.
[0097] Once a plurality of data sets (whether preprocessed or not) have been obtained, an initial candidate group of the data sets is identified as described above in connection with the determination of the geometry of the lane boundaries.
[0098] In a preferred embodiment, the tile size of the second level of the spatial index system is selected based on the resampling interval (i.e., the distance between resampled data points in the dataset). In a preferred embodiment, the tile size of the second level of the spatial index system is selected such that consecutive data points in the dataset are located (or at least typically likely to be located) in the same tile or n-level adjacent tiles in the spatial index system. The value of n can be set to any suitable and desired integer. In a particularly preferred embodiment, n is 1, 2, 3, or 4. That is, the resampling interval and tile size are preferably such that it can be expected that most consecutive data points in the dataset will be located in adjacent tiles or in tiles that are 2, 3, or 4 tiles apart (i.e., there are 1, 2, or 3 tiles between the two tiles in which consecutive data points in the dataset are located).
[0099] Based on the counterparts among the data points of the dataset that fit within the same tile or n-level adjacent tiles of the spatial index system, the dataset is identified as part of an initial candidate group of datasets. In other words, datasets that do not include corresponding data points that fit within the same tile or n-level adjacent tiles as one or more other datasets are excluded from the initial candidate group.
[0100] In a particularly preferred embodiment, if at least two corresponding data points from two datasets fit within the corresponding same tile or n-level adjacent tiles of the spatial index system, then the two datasets are determined to form part of the initial candidate group. That is, the first data point of the first dataset and the corresponding first data point of the second dataset fit within the first same tile or n-level adjacent tiles, and the second data point of the first dataset and the second data point of the second dataset fit within the second same tile (which may be the same as or different from the first same tile) or n-level adjacent tiles of the spatial index system.
[0101] By forming the initial candidate group based on the spatial index system, those datasets related to vehicle trajectories in lanes that are not the lanes for which the centerlines are to be determined can be removed from consideration. The tiling of the spatial index system makes this preliminary filtering step faster and easier because it is not necessary to compare each dataset point-by-point to determine the initial candidate group; instead, the datasets are considered row-by-row and those lines that are determined to be far from the line (lane centerline) to be determined can be removed from further consideration.
[0102] In a preferred embodiment, the initial candidate group includes at least a determined number of datasets. In other words, the initial candidate group is not identified unless a determined number of datasets are available for inclusion in the initial candidate group. The determined number of datasets can be any suitable and desired number of datasets. The determined number may (and preferably does) vary according to requirements. In a preferred embodiment, the initial candidate group includes at least 5 datasets, preferably at least 8 datasets, preferably at least 10 datasets, or preferably more datasets.
[0103] Once the initial candidate group has been identified, clusters of data sets associated with the same first lane centerline are determined by calculating the corresponding distances between corresponding data points of different data sets and comparing the calculated distances with a distance threshold. It should be understood that the determined clusters may include some or all of the data sets from the initial candidate group.
[0104] The distance between corresponding data points can be calculated in any suitable and desired manner. In one embodiment, the calculated distance is the Euclidean distance. In another embodiment, the calculated distance can be the haversine distance.
[0105] To reduce the amount of processing required to determine clusters of data sets associated with the same first lane centerline from the initial candidate group, the directed Hausdorff distance between two data sets can be determined. That is, instead of determining the distance from the first data set to the second data set and then subsequently determining the distance from the second data set to the first data set, the distance between the two data sets is determined only once.
[0106] To this end, the unique identifiers of the data sets can be used to filter the data to be processed. In one embodiment, a tuple is formed for each pair of data sets for which the distance is to be calculated. Each tuple includes a left value corresponding to the unique identifier of the first data set of the pair of data sets and a right value corresponding to the unique identifier of the second data set of the pair of data sets. The tuples are filtered such that any tuple in which the value of the identifier on the left side of the tuple is greater than the identifier on the right side of the tuple is removed from consideration. In this way, it can be ensured that any corresponding pair of data sets is only compared once. Of course, it should be understood that in alternative embodiments, the filtering can be performed based on the right value of the tuple rather than the left value. By filtering the data in this way before calculating the corresponding distances, the duplication of processing can be minimized.
[0107] To calculate the distance between a pair of data sets, the corresponding distances between the corresponding data points within the data sets are calculated and the distance between the pair of data sets is calculated based on the distances between the corresponding points.
[0108] In one embodiment, only the distance between the first data point of the first data set and the first data point of the second data set is calculated. Then, the calculated distance is used as the distance between the two data sets.
[0109] However, it may be desirable to consider multiple corresponding data points within each data set (e.g., to improve the accuracy of the calculated distances). Thus, in a preferred embodiment, determining a cluster of data sets related to the same first lane centerline from within the identified initial candidate group includes: calculating the Euclidean distance between a first data point of a first data set from the initial candidate group and a corresponding first data point of a second data set from the initial candidate group; calculating the Euclidean distance between a second data point of the first data set and a corresponding second data point of the second data set; and when the calculated distances are below a desired comparison distance threshold, determining that the first data set and the second data set should be clustered together as related to the same first lane centerline.
[0110] It may be desirable to define a minimum number of data points for which distances should be calculated. Thus, in a particularly preferred embodiment, determining a cluster of data sets related to the same first lane centerline from within the identified initial candidate group includes: determining a minimum number of corresponding data points for which Euclidean distances are to be calculated based on the number of data points within the first data set and / or the second data set; calculating the Euclidean distances between the determined number of corresponding pairs of data points from the first data set and the second data set; and when all of the calculated distances are below a desired comparison distance threshold, determining that the first data set and the second data set should be grouped together as related to the same first lane centerline.
[0111] In one embodiment, the maximum of the calculated distances between corresponding points is determined as the distance between the data sets. In another embodiment, the minimum of the calculated distances between corresponding points is determined as the distance between the data sets.
[0112] Once the distance between a pair of data sets has been calculated, the calculated distance is compared to a distance threshold to determine whether the pair of data sets should be grouped to form a cluster. This comparison can be made in any suitable and desired manner.
[0113] In one embodiment, data pairs for which the calculated distance reaches or is below the threshold are grouped to form a cluster, while data pairs for which the calculated distance exceeds the threshold are not grouped to form a cluster.
[0114] In a preferred embodiment, a cluster includes a minimum number of data sets. In other words, a cluster is determined only if it is determined that at least the minimum number of data sets should be grouped to form a cluster. Conversely, if it is determined that fewer than the minimum number of data sets should be grouped to form a cluster, then no cluster is formed. The minimum number of data sets can be any suitable and desired number. The minimum number can (and preferably) be set according to requirements. In a particularly preferred embodiment, the minimum number of data sets is 5, 8, 10 or more. By requiring a minimum number of data sets to form a cluster, artifacts caused by noise and / or other data aberrations can be reduced or avoided.
[0115] Once clusters of data sets associated with the same first lane centerline have been identified, the identified clusters can be used to determine the position of the lane centerline.
[0116] The position of the lane centerline can be determined in any suitable and desired manner. In a preferred embodiment, determining the position of the lane centerline includes: generating a bounding box that encompasses the data points of all data sets within the identified clusters; and using the generated bounding box to determine the position of the lane centerline. In a particularly preferred embodiment, using the generated bounding box to determine the position of the lane centerline includes: determining the centerline of the bounding box; and using the determined centerline of the bounding box to determine the position of the lane centerline.
[0117] In an embodiment, the method further includes: identifying an intersection located in a sub-region of a geographical area represented by a digital map, the intersection allowing a vehicle passing through the intersection to take one of a plurality of different routes, wherein at least a part of each of two of the plurality of different routes overlaps; and performing a process to deduplicate the overlapping parts of the different routes.
[0118] In an embodiment, deduplicating the overlapping parts includes determining that a first part of a first route overlaps with a first part of a second route and combining the first part of the first route with the first part of the second route to form a combined part, the combined part forming a part of both the first route and the second route.
[0119] In an embodiment, the method as described herein further includes updating the digital map to include the determined position of the first lane centerline.
[0120] The present disclosure also extends to a device configured to perform one or more of the methods disclosed herein. Thus, a device can be provided that includes one or more processors configured to perform any one of the methods disclosed herein. In a particularly preferred embodiment, the methods disclosed herein are performed in a distributed processing system that includes a plurality of data processors configured to perform data processing operations. In a particularly preferred embodiment, the data processing operations are assigned to respective ones of the plurality of data processors such that processing related to an initial candidate group is performed by the same data processor.
[0121] The present disclosure also extends to software configured to cause a processor to perform one or more of the methods disclosed herein. Thus, a computer program product can be provided that includes a set of instructions that, when executed by one or more processors, will perform one of the methods disclosed herein.
[0122] That is, the methods according to the present disclosure can and preferably are implemented at least in part using software (e.g., a computer program).
[0123] Accordingly, it will be seen that, when viewed from another aspect, the present disclosure may comprise any combination of the following: computer software which is particularly adapted to carry out the methods described herein when installed on a data processing component; computer program elements which include computer software code portions for executing the methods described herein when the program element is run on a data processing component; and computer programs which include code components adapted to carry out all or some of the steps of one or more methods described herein when the program is run on a data processing system.
[0124] Accordingly, the methods described herein may suitably be embodied as a computer program product for use with a computer system. Such an embodiment may include a series of computer readable instructions fixed on a tangible, non-transitory medium (e.g., a computer readable medium such as a floppy disk, CD-ROM, ROM, RAM, flash memory or hard disk). It may also include a series of computer readable instructions which may be transmitted to the computer system via a modem or other interface device over a tangible medium (including but not limited to optical or analog communication lines) or wirelessly using wireless technologies (including but not limited to microwave, infrared or other transmission technologies). The series of computer readable instructions embody all or part of the functionality described hereinbefore.
[0125] Those skilled in the art will appreciate that such computer readable instructions may be written in several programming languages for use with many computer architectures or operating systems. Additionally, such instructions may be stored using any current or future memory technologies (including but not limited to semiconductor, magnetic or optical) or transmitted using any current or future communication technologies (including but not limited to optical, infrared or microwave). In view of this, the computer program product may be distributed as a removable medium with accompanying printed or electronic documentation (e.g., off-the-shelf software), preloaded on a computer system (e.g., on system ROM or a fixed hard disk), or distributed from a server or electronic bulletin board over a network (e.g., the Internet or the World Wide Web). BRIEF DESCRIPTION OF THE DRAWINGS
[0126] Several embodiments will now be described by way of example only and with reference to the drawings, in which:
[0127] Figure 1 A flowchart showing a method for determining the geometry of lane boundaries in an embodiment is shown;
[0128] Figures 2A to 2F Steps for determining the geometry of lane boundaries according to an embodiment are schematically shown;
[0129] Figure 3 Data processing steps according to an embodiment are schematically shown;
[0130] Figures 4A to 4CSchematically show the steps of the method for determining the lane centerline in the embodiment; and
[0131] Figure 5 Show a flowchart illustrating the method for determining the lane centerline in the embodiment. Detailed implementation manner
[0132] Figure 1 Show method 100 for determining the geometry of lane boundaries according to the present disclosure. Method 100 includes obtaining (step 102) a plurality of data sets representing multiple individual observations of lane boundaries within a road segment in a geographic area represented by a digital map. Different ones of the plurality of data sets may represent observations of the same or different lane boundaries within the road segment.
[0133] Figure 2A Schematically show a plurality of data sets 1, 2, 3, 4 representing multiple individual observations of lane boundaries within road segment 200 indicated by a dashed line. The plurality of data sets 1, 2, 3, 4 correspond to observations of lane boundaries spaced apart within the indicated road segment. Each data set among the plurality of data sets 1, 2, 3, 4 contains a series of corresponding data points spaced along the road segment and representing the positions of the lane boundaries. For illustrative purposes only, one data set 1 includes two data points 202, 204 indicating the corresponding starting and ending points of the observation associated with one data set 1. It should be understood that one (and each) data set may include more than two data points.
[0134] Each data set among the plurality of data sets 1, 2, 3, 4 is resampled at regular intervals to obtain data sets 1', 2', 3', 4', as Figure 2B illustrated. The individual data sets are hereinafter referred to as the first data set 1', the second data set 2', the third data set 3' and the fourth data set 4'. Each data point within the plurality of data sets 1', 2', 3', 4' is associated with a unique identifier that allows the corresponding data point to be distinguished from other data points within the same data set and from other data points within other data sets.
[0135] As Figure 2C shown, in this example, the data points of data set 1' are assigned unique identifiers 1.1, 1.2, 1.3 and 1.4. Similarly, the data points of data set 2' are assigned unique identifiers 2.1, 2.2, 2.3, 2.4 and 2.5; the data points of data set 3' are assigned unique identifiers 3.1, 3.2, 3.3, 3.4 and 3.5; and the data points of data set 4' are assigned unique identifiers 4.1, 4.2, 4.3 and 4.4.
[0136] Correspondents among the data points from each data set are determined based on the distances between the data points. Specifically, based on the distance between the first data point of the first data set being shorter than the distances between the first data point of the first data set and any other data points of the second data set, the first data point of the first data set is determined to correspond to the first data point of the second data set. In other words, the data point 1.1 of the first data set 1’ is determined to correspond to the data point 2.1 of the second data set 2’, the data point 3.1 of the third data set 3’, and the data point 4.1 of the fourth data set 4’.
[0137] Return reference Figure 1 In the flow chart shown in , an initial candidate group of data sets is identified (step 104) from multiple data sets 1’, 2’, 3’, 4’ for which it is to be further determined whether the data sets should be clustered together as related to the same first lane boundary.
[0138] Identifying the initial candidate group includes using a spatial indexing system in which the geographical area containing the road segments is subdivided into multiple tiles. This is illustrated in Figure 2D Although Figure 2D only 7 tiles of the spatial indexing system are shown for readability, it should be understood that the spatial indexing system includes more tiles and extends over the entire area of the road segment 200. Figure 2D The tiles 220, 222, 224, 226, 228, 230, 232 illustrated in may be tiles of an H3 indexing system.
[0139] The tiles of the spatial indexing system are spatially indexed relative to each other such that it can be determined which tiles are adjacent to each other. As illustrated in Figure 2D the central tile 220 is in the shape of a regular hexagon and is surrounded by 6 other tiles 222, 224, 226, 228, 230, and 232. Due to the spatial indexing of the tiles, it can be determined that the tile 220 is adjacent to each of the other 6 tiles 222, 224, 226, 228, 230, 232 (surrounded by the other 6 tiles 222, 224, 226, 228, 230, 232).
[0140] The resolution of the spatial indexing system (i.e., the tile size of the spatial indexing system) is selected based on the resampling interval such that consecutive data points within the data set are located (or at least likely to be located) in the same tile (e.g., the tile 228 containing the data points 4.1 and 4.2) or adjacent tiles (e.g., the tiles 220 and 226 containing the data points 1.1 and 1.2, 2.1 and 2.2, and 3.1 and 3.2, respectively).
[0141] From tile chunks 220 and 226, initial candidate groups of data sets 1', 2', and 3' can be identified, for which it will be further determined whether the data sets should be clustered together as related to the same first lane boundary. This is because each of the corresponding data points 1.1, 2.1, and 3.1 is contained within the same tile chunk 220, and because each of the corresponding data points 1.2, 2.2, and 3.2 is contained within the same tile chunk 226. The fourth data set 4' is not included in the initial candidate group of data sets because it does not meet the condition that the corresponding points are located in the same tile chunk as the corresponding points of the other data sets.
[0142] It should be understood that in another embodiment, the condition for determining the initial candidate group of data sets can be set such that the corresponding data points should be located in neighboring tile chunks of the spatial indexing system (or within 2, 3, or 4 tile chunks). In such a case, the fourth data set 4' as depicted in Figure 2D will also be included in the initial candidate group. In such a case, it may also be considered to select a higher resolution (i.e., a smaller tile size). Although this may improve accuracy, it may also increase the computational load and may pose an increased risk of false negatives.
[0143] In Figure 1 step 106 of the method illustrated in Figure 2E a cluster of data sets related to the same first lane boundary is determined from the initial candidate group by calculating the respective distances between the corresponding data points of different data sets and comparing the calculated distances with a distance threshold.
[0144] Of course, it should be understood that in other embodiments, not all of the data sets in the identified initial candidate group will form part of the determined cluster of data sets related to the same first lane boundary.
[0145] In Figure 1 step 108 of the method shown in Figure 2F the cluster is used to determine the geometry of the first lane boundary. As illustrated in
[0146] In Figure 1 step 110 of the method shown in
[0147] Figure 3Schematically shows the data processing steps according to an embodiment. To reduce the amount of data processing required to compute the distance between data sets, unique identifiers of the data sets are used to filter the data to be processed.
[0148] In one embodiment, a tuple in the format [A, B] is formed for each pair of data sets for which the distance is to be computed. Each tuple includes a first left value A corresponding to the unique identifier of the first data set of the pair of data sets and a second right value B corresponding to the unique identifier of the second data set of the pair of data sets.
[0149] As Figure 3 illustrated in Table 302 of
[0150] each combination of two unique identifiers is represented twice. In other words, the first data set 1 and the second data set 2 are combined in the first tuple [1, 2] and the second tuple [2, 1]. This is repeated for all combinations of the paired data sets. Figure 3 The shading in Table 304 of
[0151] is used to identify the tuples where the left value A is greater than the right value B. The identified tuples are filtered such that any tuple where the left value is greater than the right value is removed from consideration. The resulting Table 306 shows the simplified set of tuples for which the distance is to be computed.
[0152] Figures 4A to 4C Schematically shows the process steps for determining a lane centerline based on vehicle trajectories. Figure 4A A segment of a road network is shown, which includes three lanes at an intersection within the road network where vehicles can turn left or right depending on which lane they are traveling in. Also illustrated are multiple vehicle trajectories of vehicles traveling along the lanes. Two lanes allow only one possible path of travel: Vehicles in the left lane can only travel from point 402 to point 402a (turning left at the intersection), while vehicles in the rightmost lane can only travel from point 406 to point 406a (turning right at the intersection).
[0153] Vehicles traveling in the middle lane (i.e., starting from point 404) can turn left or right according to one of four possible paths. A vehicle starting in the middle lane and intending to turn left can travel from point 404 to point 404a and eventually enter the middle lane after the intersection, or can travel from point 404 to point 404b and eventually enter the rightmost lane after the intersection. Similarly, a vehicle starting in the middle lane and intending to turn right can travel from point 404 to point 404c and eventually enter the middle lane after the intersection, or can travel from point 404 to point 404d and eventually enter the leftmost lane after the intersection.
[0154] Figure 4B Shows a portion of a spatial indexing system superimposed on a road segment from Figure 4A . The spatial indexing system includes a number of hexagonal tiles (408 to 422) that fit together without overlap such that the road segment is divided into sub-regions corresponding to the tiles.
[0155] Group the data sets based on the spatial indexing system such that all data sets located within a given tile 408 to 422 are processed together. In the case where a data set corresponding to a vehicle trajectory extends across multiple tiles of the spatial indexing system, the original data set is subdivided into corresponding multiple data sets such that each of the multiple data sets is contained within a single tile of the spatial indexing system.
[0156] By processing each tile separately, the processing can be distributed to multiple processors to accelerate the determination of the lane centerline. That is, the more processors that can process parts of the data in parallel, the greater the throughput of the entire system (and thus, the faster the data can be processed).
[0157] Figure 4C Shows a number of clusters (424 to 462) that have been identified as being related to the lane centerline. The clustering method was described above in connection with Figures 2A to 2F and for the sake of brevity is not repeated here. Due to the spatial indexing nature of the tiles of the spatial indexing system, it is possible to determine which clusters are related to the same lane. For example, it can be determined that the identified clusters 436, 434, 430, and 424 are all related to the centerline of the left lane.
[0158] The central tile 416 located at the center of the intersection (see Figure 4B ) contains multiple overlapping routes. Once the clusters have been determined as described above, there will be two sections of the intermediate lane that have overlapping clusters. These are the clusters 450 and 460 as shown in Figure 4C .
[0159] To remove duplicates, the tile 416 is processed (deduplicated) after the clusters have been identified. Deduplication of the overlapping portions includes determining that a first portion of a first route overlaps with a first portion of a second route and combining the first portion of the first route with the first portion of the second route to form a combined portion that forms a portion of both the first route and the second route. In this way, the individual clusters 450, 460 are formed in place of the initially identified overlapping clusters.
[0160] Figure 5A flowchart schematically showing a method 500 according to aspects of the present disclosure. The method 500 includes: obtaining (step 502) a plurality of data sets representing a plurality of individual vehicle trajectories of a vehicle traveling along a corresponding portion of a road segment; identifying (step 504) an initial candidate group of the data sets for which it is to be determined whether the data sets are related to the same first lane centerline; determining (step 506) a cluster of data sets related to the same first lane centerline from the initial candidate group of the data sets by calculating corresponding distances between corresponding data points of different data sets and comparing the calculated distances with a distance threshold; using the determined cluster to determine (step 508) the position of the first lane centerline; and updating (step 510) a digital map to include the determined position of the first lane centerline.
[0161] It should be understood that although the data sets representing observations along lane boundaries that are substantially straight are illustrated in the figures, the techniques disclosed herein are also applicable to determining the geometry of non-straight (e.g., curved) lane boundaries. Thus, in an embodiment, at least one of the plurality of data sets represents a curved lane boundary (e.g., having a curvature typically found in roads and / or highways).
[0162] Although the embodiments described herein relate to determining the geometry of lane boundaries and / or the position of lane centerlines, it should be understood that the techniques disclosed herein can be more broadly applied to determining the geometry and / or position of any feature that can be included in a digital map and / or used to generate and / or update a digital map. For example, the same principles can be applied to determine whether a detection trajectory - i.e., a set of sequentially ordered position data reflecting the movement of a vehicle and / or a device - corresponds to the same path. It should be further understood that some of the techniques disclosed herein can be used for other forms of data comparison, e.g., for map-to-map comparison.
[0163] The foregoing detailed description has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the technology to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The described embodiments were chosen in order to best explain the principles of the technology and its practical applications, to thereby enable others skilled in the art to best utilize the technology in various embodiments and with various modifications as are suited to the particular use contemplated. The scope is intended to be defined by the appended claims.
Claims
1. A method for determining the geometry of lane boundaries within a road segment within a geographic area represented by a digital map, wherein the lane boundaries divide the road segment into a set of one or more lanes, the method comprising: obtaining a plurality of data sets representing multiple separate observations of lane boundaries within the road segment, wherein different ones of the plurality of data sets may represent observations of the same or different lane boundaries within the road segment, and wherein each data set comprises a series of respective data points spaced along the road segment and representing locations of the lane boundaries; identifying, from the plurality of data sets representing individual observations of lane boundaries within the road segment, an initial candidate group of data sets for which it is to be further determined whether the data sets should be clustered together as relating to the same first lane boundary, wherein the identifying of the candidate group of data sets is performed using a spatial indexing system, wherein the geographic area including the road segment is subdivided into a plurality of patches, each patch representing a respective sub-area of the geographic area, and wherein the positions of the patches are spatially indexed relative to each other so that it can be determined which patches are adjacent to each other, wherein a data set is identified as part of the initial candidate group of data sets based on a correspondence of the data points of the data sets of the same patch or n-level adjacent patches fitting the spatial indexing system; determining a cluster of data sets associated with the same first lane boundary from the data sets within the identified initial candidate group of data sets by calculating respective distances between corresponding data points of different data sets and comparing the calculated distances to a distance threshold; and The geometry of the first lane boundary is determined using the cluster of data sets that have been determined to be associated with the same first lane boundary.
2. The method according to claim 1, wherein obtaining a plurality of data sets comprises obtaining a plurality of data sets in a first format and processing the obtained plurality of data sets into a desired format, optionally wherein obtaining a plurality of data sets comprises, for each data set in the plurality of data sets: determining that the observation of the lane boundary represented by the data set extends beyond a geographic boundary; and dividing the data set to form at least two data sets, wherein one of the new data sets is located entirely on one side of the geographic boundary and the other of the new data sets is located entirely on the other side of the geographic boundary; and / or The obtaining of the plurality of data sets comprises, for each of the plurality of data sets: determining that the series of data points in the data set are not equidistantly spaced along the road segment; and resampling the data set to obtain a data set comprising a series of data points equally spaced along the road segment; and optionally, The tile size of the spatial index system is selected based on a resampling interval, and optionally the tile size of the spatial index system is selected so that consecutive data points in a data set are located in the same tile or n-level adjacent tiles in the spatial index system.
3. The method of any of the preceding claims, wherein determining a cluster of data sets associated with the same first lane boundary from the data sets within the identified initial candidate group comprises: calculating a Euclidean distance between a first data point of a first data set from the initial candidate group and a corresponding first data point of a second data set from the initial candidate group; calculating a Euclidean distance between a second data point from the first data set and a corresponding second data point from the second data set; When the calculated distance is below an expected comparison distance threshold, determining that the first data set and the second data set should be clustered together as being associated with the same first lane boundary, optionally wherein determining that the first data set and the second data set should be clustered together as being associated with the same first lane boundary comprises: determining a minimum number of corresponding data points for which a Euclidean distance is to be calculated based on the number of data points in the first data set and / or the second data set; calculating a Euclidean distance between the determined number of corresponding pairs of data points from the first data set and the second data set; and When all of the calculated distances are below a desired comparison distance threshold, it is determined that the first data set and the second data set should be grouped together as being associated with the same first lane boundary.
4. A method according to any of the preceding claims, wherein obtaining data comprises obtaining data from a plurality of separate trips along the road segment; and / or wherein the data in the plurality of data sets is sensor data from on-board sensors.
5. The method according to any one of the preceding claims, further comprising: generating a bounding box encompassing the data points of all data sets within the determined cluster; and Determining the geometry of the lane boundary using the generated bounding box, optionally wherein determining the geometry of the lane boundary using the generated bounding box comprises: determining a centerline of the bounding box; and The geometry of the lane boundary is determined using the determined centerline.
6. A method of determining the position of a lane centerline within a road segment within a geographic area represented by a digital map, wherein the road segment is divided into a set of one or more lanes each bounded by two lane boundaries, wherein each lane centerline indicates a midpoint between the two lane boundaries of a respective one of the one or more lanes, the method comprising: obtaining a plurality of data sets representing a plurality of individual vehicle trajectories of vehicles traveling along respective portions of the road segment, wherein different ones of the plurality of data sets may represent different vehicle trajectories along the same or different lanes of the respective portion of the road segment, wherein each data set comprises a series of respective data points spaced along the portion of the road segment and representing the trajectories of vehicles traveling along a particular lane of the portion of the road segment; identifying, from the obtained plurality of data sets, an initial candidate group of data sets for which it is to be further determined whether the data sets should be clustered together as being associated with the same first lane centerline; wherein the identifying of the initial candidate group of data sets is performed using a spatial indexing system, wherein the geographic area including the road segment is subdivided into a plurality of patches, each patch representing a respective sub-area of the geographic area, and wherein the positions of the patches are spatially indexed relative to each other so that it can be determined which patches are adjacent to each other, wherein a data set is identified as part of the initial candidate group of data sets based on a correspondence of the data points of the data sets of the same patch or n-level adjacent patches fitting the spatial indexing system; determining a first cluster of data sets associated with the same first lane centerline from the data sets within the identified initial candidate group of data sets by calculating respective distances between corresponding data points of different data sets and comparing the calculated distances to a distance threshold; and The position of the lane centerline is determined using the first cluster of data sets that have been determined to be associated with the same first lane centerline.
7. The method of claim 6, wherein obtaining a plurality of data sets comprises obtaining a plurality of data sets in a first format and processing the obtained plurality of data sets into a desired format, optionally wherein obtaining a plurality of data sets comprises, for each data set in the plurality of data sets: determining that the series of data points in the data set are not equidistantly spaced along the road segment; and The data set is resampled to obtain a data set comprising a series of data points equally spaced along the road segment.
8. The method of claim 6 or 7, wherein the spatial index system is a hierarchical spatial index system comprising a first level having a first tile size and a second level having a second smaller tile size, wherein each portion of the road segment corresponds to an area covered by a corresponding tile of the first level of the spatial index system, wherein obtaining a plurality of data sets comprises determining one or more sets of data points for each data set, wherein each set of data points comprises data points fitting a corresponding patch of the first level, and wherein the patches used to identify the initial candidate group of data sets are patches of the second level, optionally wherein a patch size of the patches of the first level of the spatial indexing system is selected based on a size of the road segments.
9. The method of any one of claims 6 to 8, wherein determining a cluster of data sets associated with the same first lane centerline from the data sets within the identified first initial candidate group comprises: calculating a Euclidean distance between a first data point of a first data set from the first initial candidate group and a corresponding first data point of a second data set from the first initial candidate group; calculating a Euclidean distance between a second data point from the first data set and a corresponding second data point from the second data set; When the calculated distance is below an expected comparison distance threshold, determining that the first data set and the second data set should be clustered together as being associated with the same first lane centerline, optionally wherein determining that the first data set and the second data set should be clustered together as being associated with the same first lane centerline comprises: determining a minimum number of corresponding data points for which a Euclidean distance is to be calculated based on the number of data points in the first data set and / or the second data set; calculating a Euclidean distance between the determined number of corresponding pairs of data points from the first data set and the second data set; and When all of the calculated distances are below a desired comparison distance threshold, it is determined that the first data set and the second data set should be grouped together as being associated with the same first lane centerline.
10. The method according to any one of claims 6 to 9, further comprising: identifying the road segment includes an intersection located in a sub-area of the geographic area represented by the digital map, the intersection allowing a vehicle passing through the intersection to take one of a plurality of different routes, wherein at least a portion of each of at least two of the plurality of different routes overlaps; and performing processing to deduplicate the overlapping portions of the different routes, and / or The data in the plurality of data sets are sensor data from vehicle-mounted sensors; and / or the data in the plurality of data sets are GNSS data, such as GPS data.
11. The method according to any one of claims 6 to 10, further comprising: generating a bounding box encompassing the data points of all data sets within the determined cluster; and using the generated bounding box to determine the position of the lane centerline; and optionally, wherein using the generated bounding box to determine the position of the lane centerline comprises: determining a centerline of the bounding box; and The lane centerline is determined using the determined centerline of the bounding box.
12. The method of any one of claims 1 to 5, further comprising updating the digital map to include the determined geometry of the first lane boundary; or The method of any one of claims 6 to 12, further comprising updating the digital map to include the determined lane centerline.
13. A method according to any of the preceding claims, wherein the first data point of the first data set is determined to correspond to the first data point of the second data set based on the distance between the first data point of the first data set and the first data point of the second data set being shorter than the distance between the first data point of the first data set and any other data point of the second data set.
14. A method according to any one of the preceding claims, wherein The spatial index system comprises a set of regular polygonal tiles, such as a set of regular hexagonal tiles, optionally wherein the spatial index system comprises an H3 index system, and / or The method is performed in a distributed processing system including a plurality of data processors configured to perform data processing operations, the method further comprising: Data processing operations are distributed to respective ones of the plurality of data processors such that processing associated with the initial candidate group is performed by the same data processor.
15. A computer program product comprising a set of instructions which, when executed by one or more processors, will perform the method as claimed in any of the preceding claims and / or an apparatus comprising one or more processors configured to perform the method as claimed in any of the preceding claims.