Method for storing and providing georeferenced vehicle data, computer-readable medium, and distributed system

EP4555424A1Pending Publication Date: 2025-05-21BAYERISCHE MOTOREN WERKE AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2023721611
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-15
Filing Date
2023-04-20
Publication Date
2025-05-21

AI Technical Summary

Technical Problem

Current methods for storing and querying georeferenced vehicle data in distributed systems, such as Apache Spark or Geospark, are inefficient due to data distribution and partitioning schemes that lead to suboptimal storage and querying, particularly for large datasets like those from vehicle fleets.

Method used

A method that uses a distributed system to receive and process georeferenced vehicle data by determining a subset with time and location information, training a machine learning model to predict zoom levels, and partitioning data using a scheme that includes time and location partitions, enabling efficient storage and querying through geohashes.

Benefits of technology

This approach allows for efficient and fault-tolerant storage and querying of georeferenced vehicle data, ensuring accurate data retrieval even if data is stored in non-optimal partitions, by using a trained machine learning method to predict zoom levels and determine geohashes for partitioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 1.1
    Figure 1.1
Patent Text Reader

Abstract

The invention relates to a method for storing and providing georeferenced vehicle data, the method comprising the following steps: A data memory of a distributed system receives the georeferenced vehicle data; a first computing unit of the distributed system determines a subset of the georeferenced vehicle data, wherein the subset comprises data with time information and location information of the georeferenced vehicle data; the first computing unit of the distributed system ascertains a partitioning schema comprising time partitions and location partitions depending on time information and location information of the subset of georeferenced vehicle data; the first computing unit of the distributed system trains a machine learning method to predict a zoom level of the location information of the georeferenced vehicle data with the ascertained partitioning schema and the subset of georeferenced vehicle data; a plurality of computing units of the distributed system partition data of the received georeferenced vehicle data using the trained machine learning method; the plurality of computing units of the distributed system store the partitioned data of the received georeferenced vehicle data in the data memory of the distributed system; a request message for georeferenced vehicle data is received, wherein the request message comprises time information and location information; and the georeferenced vehicle data of the request message is provided, wherein the georeferenced vehicle data is provided using the stored, partitioned data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Method for storing and providing georeferenced vehicle data, computer-readable medium, and distributed system

[0002] The invention relates to a method for storing and providing georeferenced vehicle data. The invention further relates to a computer-readable medium for storing and providing georeferenced vehicle data and a distributed system for storing and providing georeferenced vehicle data.

[0003] Modern vehicles can transmit a multitude of vehicle data to a backend server. The vehicle data is usually big data sets, which can contain many gigabytes of data. Partitioning methods are known from the state of the art to partition large data sets in order to enable more efficient data queries. Well-known software solutions include Apache Spark and Geospark. However, these known software solutions have the disadvantage that data must be distributed between computing units of a distributed system before the data can be processed on a computing unit of the distributed system. Furthermore, these known methods have the disadvantage that each computing unit creates its own partitioning scheme, which can lead to inefficient storage and inefficient querying of the stored data.

[0004] It is therefore an object of the invention to efficiently improve the storage and provision of georeferenced vehicle data. In particular, an object of the invention is to enable fault-tolerant storage of georeferenced vehicle data, which efficiently enables error-free retrieval of the georeferenced vehicle data.

[0005] This object is achieved by the features of the independent claims. Advantageous embodiments and further developments of the invention emerge from the dependent claims.

[0006] According to a first aspect, the invention is characterized by a method for storing and providing georeferenced vehicle data. The method can be a computer-implemented method. Preferably, the georeferenced vehicle data is stored and provided on a distributed system. The distributed system can be, for example, a cloud computing system, a grid computing system, or a cluster system.

[0007] The method comprises receiving the georeferenced vehicle data by a data storage device of a distributed system. For example, the georeferenced data can be received from a plurality of vehicles in a vehicle fleet. The method further comprises determining a subset of the georeferenced vehicle data by a first computing unit of the distributed system, wherein the subset comprises data with time information and location information of the georeferenced vehicle data, and determining a partitioning scheme comprising time partitions and location partitions as a function of time information and location information of the subset of the georeferenced vehicle data by the first computing unit of the distributed system.The method trains a machine learning method for predicting a zoom level of the location information of the georeferenced vehicle data using the determined partitioning scheme and the subset of the georeferenced vehicle data by the first computing unit of the distributed system, and partitions data of the received georeferenced vehicle data using the trained machine learning method by a plurality of computing units of the distributed system. The zoom level can comprise a subdivision of a world map into predefined areas of a predetermined size. The partitioned data of the received georeferenced vehicle data is stored in the data memory of the distributed system by the plurality of computing units of the distributed system.The method finally comprises receiving a request message for georeferenced vehicle data, wherein the request message comprises time information and location information, and providing the georeferenced vehicle data of the request message, wherein the georeferenced vehicle data is provided using the stored, partitioned data.

[0008] Advantageously, the method can efficiently store and query georeferenced vehicle data.

[0009] According to an advantageous embodiment of the invention, the size of the subset of georeferenced data can be selected such that the partitioning scheme for the subset of georeferenced data can be determined by the first computing unit of the distributed system and / or the machine learning method can be trained by the first computing unit of the distributed system. This allows the machine learning method to be trained with a partitioning scheme that is complete for the subset.

[0010] According to a further advantageous embodiment, the subset of georeferenced data can be determined depending on a temporal data pattern, wherein the temporal data pattern is preferably representative of a day, a week, and / or a month, and / or wherein the temporal data pattern is preferably representative of a difference between weekdays (Monday to Friday) and weekends (Saturday and Sunday). This allows a representative subset of the georeferenced data to be efficiently determined.

[0011] According to a further advantageous embodiment, the partitioning scheme can be determined using quadtrees. This allows the partitioning scheme to be efficiently calculated on a single processing unit of the distributed system.

[0012] According to a further advantageous embodiment, the partitioning scheme for time information and location information of a data set of the subset of georeferenced data can be determined as a time partition and a location partition, wherein the location partition is preferably uniquely determined by a geohash, and wherein the geohash preferably specifies a zoom level of the location partition. This allows a partitioning scheme to be specified with which the georeferenced vehicle data can be efficiently stored and retrieved.

[0013] According to a further advantageous embodiment, training the machine learning method can include determining hyperparameters of the machine learning method using cross-validation. This can efficiently improve the training of the machine learning method.

[0014] According to a further advantageous embodiment, the method may further comprise distributing the trained machine learning method from the first computing unit of the distributed system to the plurality of computing units of the distributed system. The trained machine learning method can thus be used on the plurality of computing units to divide and store the georeferenced data into partitions.According to a further advantageous embodiment, the partitioning of the data of the received georeferenced vehicle data using the trained machine learning method by the plurality of computing units of the distributed system can comprise determining a time partition of the partitioning scheme as a function of time information of the data set, predicting a zoom level of a location partition as a function of location information of a data set of the received, georeferenced vehicle data using the trained machine learning method by a computing unit of the plurality of computing units, determining a geohash using the location information and the predicted zoom level of the location partition by the computing unit of the plurality of computing units, and determining a location partition of the partitioning scheme as a function of the geohash.This allows georeferenced data sets to be efficiently divided into partitions with different zoom levels. Furthermore, the georeferenced vehicle data sets can be stored in partitions with different zoom levels in an error-tolerant manner.

[0015] According to a further advantageous embodiment, the storage of the partitioned data of the received georeferenced vehicle data in the data memory of the distributed system by the plurality of computing units of the distributed system can include storing the data set of the received georeferenced vehicle data in the specific time and location partition. This allows the georeferenced vehicle data to be stored efficiently.

[0016] According to a further advantageous embodiment, providing the georeferenced vehicle data of the request message can include determining one or more time partitions depending on the time information of the request message, determining geohashes of all zoom levels depending on the location information of the request message, determining location partitions of the determined geohashes, querying the georeferenced vehicle data of the time information and the location information of the request message in the determined one or more time partitions and the determined location partitions, and providing the queried, georeferenced vehicle data in response to the request message. This allows a request for vehicle data to be implemented efficiently.By determining geohashes of all relevant zoom levels of the location information, georeferenced vehicle data for which the trained machine learning algorithm predicted an incorrect zoom level can be correctly located and retrieved during the query. In other words, the query can provide the correct georeferenced vehicle data, even if it is stored in an incorrect partition.

[0017] According to a further aspect, the invention is characterized by a computer-readable medium for storing and providing georeferenced vehicle data, wherein the computer-readable medium comprises instructions which, when executed on one or more computing units of a distributed system, carry out the method described above.

[0018] According to a further aspect, the invention is characterized by a distributed system for storing and providing georeferenced vehicle data, wherein the distributed system is designed to carry out the method described above.

[0019] Further features of the invention emerge from the claims, the figure, and the description of the figures. All features and combinations of features mentioned above in the description, as well as the features and combinations of features mentioned below in the description of the figures and / or shown alone in the figure, can be used not only in the respective specified combination, but also in other combinations or even on their own.

[0020] In the following, a preferred embodiment of the invention is described with reference to the accompanying drawing. Further details, preferred embodiments and developments of the invention will emerge from this. In detail, schematically

[0021] Fig. 1 shows an exemplary method for storing and retrieving georeferenced vehicle data in a distributed system.

[0022] In detail, Fig. 1 shows an exemplary method 100 for storing and retrieving georeferenced vehicle data in a distributed system. Vehicle data often includes time data and georeferenced data. For example, georeferenced data can include latitudinal and longitudinal coordinates of a vehicle's position. In addition, the vehicle data can include further use case-specific data. The vehicle data, in particular the georeferenced vehicle data, can be used for location-based services, map-based functions, and / or driver assistance systems of a vehicle. The georeferenced vehicle data can be big data datasets with several hundred gigabytes to several terabytes of data. The georeferenced vehicle data can be stored by partitioning according to time and location. Partitioning according to location can be done, for example, using geohashes.To enable partitioning with geohashes, a world map can be divided into different geometric areas, such as rectangles or hexagons, as well as into different zoom levels. A zoom level can specify the size of a geometric area. Each geohash of a zoom level n is completely filled by geohashes of a zoom level n+1 in the same geometric area. To encode a geohash, combinations of one or more letter digits and / or one or more numbers can be defined, for example, to uniquely identify the area of ​​the geohash. The zoom level can specify the size of a partition based on the size of the geometric area of ​​the zoom level. By specifying the size of the partition, the amount of data that is sorted into this partition from the total amount of georeferenced vehicle data can be specified.

[0023] In order to be able to specify a number of zoom levels, preferably all georeferenced vehicle data must be taken into account. Due to the size of the data set of georeferenced vehicle data, only a subset of the georeferenced vehicle data can be used to define the zoom levels. Furthermore, the subset of georeferenced vehicle data can include a subset of attributes of all or some subsets of the georeferenced vehicle data. To enable the most efficient distribution of the georeferenced vehicle data into partitions, method 100 proposes using a machine learning method to partition the georeferenced vehicle data.

[0024] In detail, the method 100 can receive 102 the georeferenced vehicle data through a data storage of a distributed system. The georeferenced vehicle data can be received from one vehicle or a plurality of vehicles, for example, a vehicle fleet or a subset of a vehicle fleet. The method 100 can determine 104 a subset of the georeferenced vehicle data through a first computing unit of the distributed system. The subset can include data with time information, for example, a date, and location information, for example, a position of a vehicle, of the georeferenced vehicle data. The subset preferably does not contain any use case-specific data. This allows the size of the georeferenced vehicle data to be efficiently reduced.The size of the subset of georeferenced vehicle data is preferably selected such that the subset can be used on a single computer of the distributed system to train the machine learning method. For example, the size of the subset of georeferenced vehicle data can be selected such that specifications regarding memory consumption and computing time are met when training the machine learning method. The size of the subset of georeferenced vehicle data preferably comprises a few gigabytes of data. The subset of georeferenced vehicle data is further selected such that the distribution of the data in the subset corresponds to the distribution of all georeferenced vehicle data with regard to time information and / or location information. This can avoid errors when training the machine learning method.For example, the subset of georeferenced vehicle data can only include data on time information and location information. Use-case-specific data from the georeferenced vehicle data that is not relevant for training the machine learning process cannot be further considered for determining the subset of georeferenced vehicle data. In addition, for example, only a date of time information can be used to determine the subset. A timestamp of the time information cannot be further considered for determining the subset.

[0025] The method 100 can determine 106 a partitioning scheme comprising time partitions and location partitions as a function of time information and location information of the subset of the georeferenced vehicle data by the first computing unit of the distributed system. For example, the method 100 can determine a partitioning scheme, in particular an optimal partitioning scheme for the subset of the georeferenced vehicle data, using one or more known methods. A known method for calculating a partitioning scheme can, for example, determine an optimal partitioning scheme for the subset of the georeferenced vehicle data using quadtrees. The determined partitions of the partitioning scheme can be used as labels for a multi-class classification problem to be solved using the machine learning method.

[0026] The method 100 can train 108 a machine learning method for predicting a zoom level of the location information of the georeferenced vehicle data using the determined partitioning scheme and the subset of the georeferenced vehicle data by the first computing unit of the distributed system. The machine learning method can be a known machine learning method. Various machine learning methods can be evaluated using expert knowledge or methods such as cross-validation to determine a suitable machine learning method. The goal of training the machine learning method is to predict a zoom level of the location information for each piece of time information, in particular each new date, and each piece of location information, for example latitudinal longitudinal coordinates. A geohash can then be determined for the predicted zoom level.For this purpose, the labels of the determined partitions can be used as the output of the machine learning process. A label can correspond to a zoom level of the location information. The subset of georeferenced vehicle data can serve as input for training the machine learning process.

[0027] The method 100 can partition 110 the received georeferenced vehicle data using the trained machine learning method by a plurality of computing units of the distributed system. For each data set of the received georeferenced vehicle data, the time information and the location information are used as input to the trained machine learning method to predict a zoom level. This is done analogously to the training of the machine learning method as described above. Using the zoom level and the location information, an associated geohash can be calculated. The geohash can be used to uniquely determine the location partition. Furthermore, the time partition can be determined using the time information, for example, using a timestamp of each data set of the georeferenced vehicle data.

[0028] The trained machine learning method can perform a probabilistic prediction of a zoom level for each datum in a data set of georeferenced vehicle data. This can result in the machine learning method not always predicting the optimal zoom level for each datum, and the data set of georeferenced vehicle data not being stored in the optimal partition. By storing the georeferenced vehicle data in a fault-tolerant storage scheme, subsequent queries for georeferenced vehicle data can be enabled without loss. A fault-tolerant storage scheme can be implemented by partitioning the georeferenced vehicle data using geohashes. For geohashes, a geohash of zoom level n maps to a set of geohashes of zoom level n+1.Furthermore, for geohashes, all geohashes of zoom level n+1 are completely contained within the geohash of zoom level n and fill it. When querying for georeferenced vehicle data partitioned with geohashes, all relevant geohashes of all zoom levels can be queried, thus also retrieving records of georeferenced vehicle data stored at a different zoom level. A lossless query of the georeferenced vehicle data is thus possible.

[0029] The method 100 may store 112 the partitioned data of the received georeferenced vehicle data in the data storage of the distributed system by the plurality of computing units of the distributed system. Using the time partition and the location partition, each data set of the georeferenced vehicle data may be processed and stored independently by a computing unit of the plurality of computing units of the distributed system. An example of a storage scheme for storing the data sets of the georeferenced vehicle data may be structured as follows:

[0030] Root folder /

[0031] - date=2022-03-07

[0032] - geohash=g12

[0033] All data records from 07.03.2022 and Geohash=g12

[0034] - geohash=g1234

[0035] All data records from 07.03.2022 and Geohash=g1234

[0036] - geohash=f567

[0037] All data records from 07.03.2022 and Geohash=f567

[0038] - date=2022-03-08

[0039] - geohash=g12

[0040] All data records from 08.03.2022 and Geohash=g12

[0041] - geohash=g1234

[0042] All data records from 08.03.2022 and Geohash=g1234

[0043] - geohash=f567

[0044] All data records from 08.03.2022 and Geohash=f567 io

[0045] The method 100 may receive 114 a request message for georeferenced vehicle data, wherein the request message comprises time information and location information. Furthermore, the method 100 may provide 116 the georeferenced vehicle data to the request message, wherein the georeferenced vehicle data is provided using the stored, partitioned data. An exemplary request message comprises a query for all data of the georeferenced vehicle data from a specific location as location information and from a specific time interval as time information. The specific time interval of the exemplary request message may include the days 07 / 03 / 2022 and 08 / 03 / 2022. In the exemplary system schema, the days of the time interval may be mapped directly to the time information of the time partitions. The specific location of the exemplary request message may be Munich.For the specific location Munich, the method can determine the relevant geohashes for all zoom levels. For example, the specific location Munich can include the geohash g12 for a first zoom level of the specific location and the geohash g1234 for a second zoom level of the specific location. The geohashes can be mapped directly to the partitions of the storage schema. The result of the query includes all data of the georeferenced vehicle data stored in the determined partitions of the storage schema.

[0046] The machine learning algorithm can be retrained based on a change in the distribution of the georeferenced vehicle data. To do this, new training data can be obtained, as described above, and the machine learning algorithm can be trained using the new training data.

[0047] Advantageously, the method can efficiently determine a partitioning scheme for georeferenced vehicle data, which enables fault-tolerant storage of the georeferenced vehicle data.

[0048] List of reference symbols

[0049] 100 procedures

[0050] 102 Receiving georeferenced vehicle data

[0051] 104 Determining a subset of the georeferenced vehicle data

[0052] 106 Determining a partitioning scheme

[0053] 108 Training a machine learning algorithm

[0054] 110 Partitioning of georeferenced vehicle data

[0055] 112 Saving the partitioned data

[0056] 114 Receiving a request message

[0057] 116 Providing georeferenced vehicle data

Claims

Patent claims 1. A method for storing and providing georeferenced vehicle data, the method comprising: Receiving the georeferenced vehicle data by a data storage of a distributed system; Determining a subset of the georeferenced vehicle data by a first computing unit of the distributed system, wherein the subset comprises data with time information and location information of the georeferenced vehicle data; Determining a partitioning scheme comprising time partitions and location partitions depending on time information and location information of the subset of the georeferenced vehicle data by the first computing unit of the distributed system; Training a machine learning method for predicting a zoom level of the location information of the georeferenced vehicle data with the determined partitioning scheme and the subset of the georeferenced vehicle data by the first computing unit of the distributed system; Partitioning data of the received georeferenced vehicle data using the trained machine learning method by a plurality of computing units of the distributed system; Storing the partitioned data of the received georeferenced vehicle data in the data storage of the distributed system by the plurality of computing units of the distributed system; Receiving a request message for georeferenced vehicle data, wherein the request message includes time information and location information; and Providing the georeferenced vehicle data of the request message, wherein the georeferenced vehicle data is provided using the stored partitioned data.

2. The method according to claim 1, wherein a size of the subset of the georeferenced data is selected such that the partitioning scheme for the subset of the georeferenced data can be determined by the first computing unit of the distributed system and / or the machine learning method can be trained by the first computing unit of the distributed system.

3. Method according to one of the preceding claims, wherein the subset of georeferenced data is determined as a function of a temporal data progression; and wherein the temporal data progression is preferably representative of a day, a week, and / or a month; and / or wherein the temporal data progression is preferably representative of a difference between weekdays, days Monday to Friday, and weekends, days Saturday and Sunday.

4. Method according to one of the preceding claims, wherein the partitioning scheme is determined using quadtrees.

5. The method according to any one of the preceding claims, wherein the partitioning scheme for a time information and a location information of a data set of the subset of the georeferenced data determines a time partition and a location partition, wherein the location partition is preferably uniquely determined by a geohash, and wherein the geohash preferably specifies a zoom level of the location partition.

6. The method according to any one of the preceding claims, wherein training the machine learning method comprises determining hyperparameters of the machine learning method by means of cross-validation.

7. A method according to any one of the preceding claims, the method further comprising: Distributing the trained machine learning method from the first computing unit of the distributed system to the plurality of computing units of the distributed system.

8. The method according to any one of the preceding claims, wherein partitioning the data of the received georeferenced vehicle data using the trained machine learning method by the plurality of computing units of the distributed system comprises: Determining a time partition of the partitioning scheme depending on time information of the data set; Predicting a zoom level of a location partition as a function of location information of a data set of the received, georeferenced vehicle data with the trained machine learning method by a computing unit of the plurality of computing units; Determining a geohash using the location information and the predicted zoom level of the location partition by the computing unit of the plurality of computing units; and Determine a location partition of the partitioning scheme depending on the geohash.

9. The method according to any one of the preceding claims, wherein storing the partitioned data of the received georeferenced vehicle data in the data memory of the distributed system by the plurality of computing units of the distributed system comprises: Storing the dataset of received georeferenced vehicle data in the specified time and location partition.

10. The method according to any one of the preceding claims, wherein providing the georeferenced vehicle data of the request message comprises: Determining one or more time partitions depending on the time information of the request message; Determining geohashes of all zoom levels depending on the location information of the request message; Determining location partitions of the determined geohashes; Querying the georeferenced vehicle data, the time information, and the location information of the request message in the specific one or more time partitions and the specific location partitions; and Providing the requested, georeferenced vehicle data in response to the request message.

11. A computer-readable medium for storing and providing georeferenced vehicle data, the computer-readable medium comprising instructions which, when executed on one or more computing units of a distributed system, carry out the method according to any one of claims 1 to 10.

12. A distributed system for storing and providing georeferenced vehicle data, wherein the distributed system is designed to carry out the method according to one of claims 1 to 10.