Method for anonymizing vehicle data

By classifying vehicles into similar classes using static and dynamic vectors and transmitting anonymized data, the method effectively protects vehicle data while enabling use of external services, reducing data transfer times and enhancing security.

JP2025515771AActive Publication Date: 2025-05-20MERCEDES BENZ GROUP AG
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024566452
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-05-17
Filing Date
2023-04-24
Publication Date
2025-05-20
Estimated Expiration
2043-04-24

AI Technical Summary

Technical Problem

Existing methods for anonymizing vehicle data for external service functions do not adequately protect personal data while allowing continued use of these services.

Method used

A method involving determining two characteristic quantities for each vehicle, a static and a dynamic vector, using machine learning clustering to classify vehicles into similar classes, and transmitting anonymized data from similar vehicles to external service functions, combined with spatial dimension reduction and encryption to protect individual vehicle data.

Benefits of technology

This approach effectively anonymizes vehicle data, ensuring data protection by replacing individual vehicle information with similar fleet data, reducing data transfer times, and enhancing security against central failures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025515771000001_ABST
    Figure 2025515771000001_ABST
Patent Text Reader

Abstract

The invention relates to a method for anonymizing vehicle data for utilizing a service function external to the vehicle (FES), the method according to the invention comprises the steps of: in a fleet (10) of homogenous vehicles (11, 12, 13, 14) comprising a heterogeneous group of vehicle users, determining for each of the vehicles (11, 12, 13, 14) a first characteristic quantity which is at least indirectly dependent on an integer number n of recorded vehicle sensor values ​​and a second characteristic quantity which is at least indirectly dependent on an integer number m of current vehicle sensor values, determining a similarity for at least one characteristic quantity of the plurality of characteristic quantities for a number or all of the vehicles (11, 12, 13, 14) of the fleet (10) and accordingly determining an anonymization of the vehicle data for the service function external to the vehicle (FES), the method according to the invention comprises the steps of: determining a first characteristic quantity which is at least indirectly dependent on an integer number n of recorded vehicle sensor values ​​and a second characteristic quantity which is at least indirectly dependent on an integer number m of current vehicle sensor values, determining a similarity for at least one characteristic quantity of the plurality of characteristic quantities of a number or all of the vehicles (11, 12, 13, 14) of the fleet (10) and accordingly determining an anonymization of the vehicle data for the service function external to the vehicle (FES), the method according to the invention comprises the steps of: determining a first characteristic quantity which is at least indirectly dependent on an integer number m of current vehicle sensor values ​​and a second ... The vehicles (11, 12, 13, 14) are classified into similar or dissimilar classes based on the similarity of at least one characteristic value, and when an external service function (FES) is requested, instead of a characteristic value of the requesting vehicle (11) containing information relevant to the provision of the service, a corresponding characteristic value of a vehicle (12) of a class of similar vehicles (12), a characteristic value mathematically determined from the corresponding characteristic values ​​of a plurality of vehicles (12) of a class of similar vehicles (12), or an artificially generated similar characteristic value is transmitted.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a method for anonymizing vehicle data for utilizing service functions outside the vehicle, and further to a vehicle designed to utilize said method. [Background technology]

[0002] In today's vehicles, a large amount of data from vehicles of a vehicle fleet is periodically exchanged, for example with a vehicle manufacturer, a server outside the vehicle of the management of the vehicle fleet, or with a server of a service function provider outside the vehicle. The technical progress and standard installation of vehicle and interior sensor systems and the exchange of information detected by these sensor systems allow for an excellent monitoring of people, their data and their status at points outside the vehicle. The data may include personal data of the vehicle user, which allows an estimation of the person who uses the vehicle, or the corresponding vehicle data is considered to be the case. In particular, such data can be used to create a user profile, to create a kind of "fingerprint" of the respective vehicle, etc. The data must therefore be protected more or less strongly from the point of view of data protection. In this case, the people who use the vehicle are concerned about the use of their data, but at the same time they also want to continue to use existing services.

[0003] Patent Document 1 describes a method for protecting personal data of a vehicle occupant, in which the emotional state of the vehicle occupant is detected in the vehicle as personal data, and the masked emotional state is evaluated in a device outside the vehicle. Subsequently, the emotional state is associated with an emotional group including a plurality of emotional patterns to protect the vehicle occupant, thereby depersonalizing the emotional state of the vehicle occupant.

[0004] The patent application EP 1 099 536 B1, which has not yet been published at the time of filing, describes a method for anonymizing movement data of road users equipped with location detection devices. The main objective here is to monitor the traffic flows, in which however the transmission of data that must be protected must be prevented, while the presence of data of sufficient accuracy must ensure a reliable assessment of the traffic flows.

[0005] Patent Document 3 describes a method for anonymizing vehicle data for using a service function outside the vehicle. In this method, in addition to original vehicle data representing a route section traveled by the vehicle, another artificial vehicle data representing an artificial travel section is generated. The original vehicle data and the artificial vehicle data are stored together or transmitted together. Patent document 4 discloses a method for anonymizing and transmitting a first value of a driving parameter of a vehicle to an external data receiving unit. A vehicle receives another value for the driving parameter transmitted to the vehicle from another vehicle. From the first value and the another value a second value for the driving parameter is calculated such that the first value cannot be reconstructed by the external data receiving unit. The second value is transmitted to the external data receiving unit. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] DE102020003188A1 [Patent Document 2] DE10202100378B3 [Patent Document 3] DE102015213393A1 [Patent Document 4] DE102015226650A1 Summary of the Invention [Problem to be solved by the invention]

[0007] It is an object of the present invention to provide an improved method for anonymizing vehicle data, in particular when using service functions outside the vehicle. [Means for solving the problem]

[0008] According to the invention, this problem is solved by a method with the features of claim 1, and here in particular with the features recited in the characterizing part of claim 1. Advantageous configurations and developments of the method according to the invention emerge from the claims dependent on claim 1. Claim 13 also describes a vehicle suitable for carrying out the method.

[0009] The method according to the present invention is used to anonymize vehicle data for using external service functions. It is assumed that in a large vehicle fleet of identical vehicles and the associated large heterogeneous group of vehicle users, there is a high probability of similar vehicle usage profiles. Furthermore, external service functions are used, and when using the service functions, a number of pieces of information are transmitted from the vehicles to the provider of the external service functions.

[0010] Sensors installed inside and outside the vehicle continuously determine vehicle sensor values ​​which are representative of the vehicle's state (e.g. position, speed, traveled section, route information, aging, etc.) as well as the state of the driver and passengers (e.g. attention level, driving style, etc.). According to the invention, for each vehicle in a fleet of homogenous vehicles comprising a heterogeneous group of vehicle users, a first characteristic quantity is determined which is at least indirectly dependent on an integer number n of recorded vehicle sensor values ​​and a second characteristic quantity is determined which is at least indirectly dependent on an integer number m of current vehicle sensor values.

[0011] That is, two characteristic quantities are determined for each vehicle of the fleet. The first characteristic quantity is a characteristic set of up to n vehicle sensor values ​​and is determined from recorded vehicle sensor values. This first characteristic quantity is approximately constant over time (load set data) and can therefore be referred to here as a static characteristic quantity. In limiting cases, the characteristic quantity may also correspond directly to a single sensor value. Alternatively, the second characteristic quantity is a characteristic set of up to m vehicle sensor values. The second characteristic quantity is determined from current vehicle sensor values ​​and can therefore be referred to here as a dynamic characteristic quantity, since it is determined from current vehicle sensor values ​​and changes dynamically over time (e.g. position, speed, acceleration, etc.). This characteristic quantity may also correspond directly to a single sensor value in limiting cases.

[0012] The first characteristic quantities, especially if they are based on a large number of vehicle sensor values, allow the identification of each vehicle, so to speak like a fingerprint, since they are based on individual vehicle and utilization characteristics and do not change much over time. These first characteristic quantities are therefore particularly worthy of protection from the point of view of data protection. In contrast, the second characteristic quantities reflect the current, up-to-date data of the vehicle and its utilization. Identification of the vehicle is only possible indirectly (e.g. by a combination of geolocation and local image information).

[0013] Here, a similarity is determined for a first characteristic value of a number of vehicles or all vehicles of the fleet, and the vehicles are accordingly classified into similar or dissimilar classes on the basis of the similarity of at least one characteristic value. In particular, vehicles having similar vehicle sensor values ​​in the first characteristic value and / or the second characteristic value can be identified. In a preferred configuration of the invention, for this purpose, a machine learning clustering mechanism (e.g. k-means, Mean-Shift clustering or EM (expectation maximisation) clustering) is used to generate a similarity for the first characteristic value and / or the second characteristic value. Here, the similarity in at least the second characteristic value is determined continuously, since the second characteristic value changes over time in the fleet.

[0014] Here, when a service function outside the vehicle is requested, instead of the characteristic quantities of the requesting vehicle containing information relevant to the provision of the service, corresponding characteristic quantities of a vehicle of a similar vehicle class, characteristic quantities arithmetically determined from corresponding characteristic quantities of a number of vehicles of a similar vehicle class, or similar characteristic quantities artificially generated are transmitted.

[0015] This creates a "privacy-layer" which in the simplest case uses sufficiently similar information of other vehicles in the fleet instead of the vehicle's actual information for the corresponding external vehicle service function. In one alternative configuration, the relevant information can be determined via a group of relevant information of similar vehicles, e.g. averaged or generated artificially.

[0016] That is, when a service external to the vehicle is called up by a first vehicle, information from characteristic quantities of similar vehicles that is relevant to the service function is transmitted, data that has been determined or artificially generated from those characteristic quantities (e.g. by averaging).

[0017] Besides, typically, the information transmitted to the provider of the service function outside the vehicle is not generally all necessary to make the service function available, but a large amount of information unrelated to the service function is transmitted to the provider of the service function outside the vehicle, and the additional information can be monetized by the provider of the service function outside the vehicle as added value. Typically, the information is based on other characteristic values. Therefore, according to a very advantageous development of the method according to the invention, instead of the other characteristic values ​​of the requesting vehicle, corresponding characteristic values ​​of a vehicle from a class of dissimilar vehicles, characteristic values ​​that are mathematically determined from corresponding characteristic values ​​of a plurality of vehicles from a class of dissimilar vehicles, or similar characteristic values ​​that are artificially generated are transmitted.

[0018] This creates a "privacy layer" that replaces the vehicle's information with information from similar vehicles related to service functions outside the vehicle, and irrelevant information from dissimilar vehicles, thus hiding the vehicle's original information worthy of protection.

[0019] Artificially generated characteristic quantities are possible in either case, but they play a crucial role especially in the case of less relevant information and, according to an advantageous configuration of the invention, can be generated via generative methods of machine learning.

[0020] For this purpose, models can be trained that are optimized to generate the most "realistic" possible characteristic values ​​and these models can be used for the random characteristic values, in particular the second characteristic value. This "realism" is achieved by using appropriate functions that are either generally defined or defined depending on the service functions that the vehicle wants to use. This means that a content-specific focus on the n or especially m vehicle sensor values ​​can also be placed.

[0021] According to a very advantageous embodiment of the method according to the invention, the characteristic quantities which are at least indirectly dependent on an integer number of vehicle sensor values, i.e. the first and second characteristic quantities, can be formed as a set of n or m vehicle sensor values, i.e. the characteristic quantities consist of corresponding sets or, insofar as a specific arrangement of the sensor values ​​in a predefined order is desired, of n-tuples or m-tuples.

[0022] In an alternative configuration of the respective characteristic quantity, the respective characteristic quantity can be formed as a vector based on the respective vehicle sensor values ​​in an n-dimensional or m-dimensional space.

[0023] That is, two vectors are determined for each vehicle of the fleet: the first characteristic is a characteristic n-dimensional vector determined from the recorded vehicle sensor values, and as the second vector corresponding to the second characteristic an m-dimensional vector determined from the current vehicle sensor values ​​is considered.

[0024] In particular the first vector allows for a very simple identification of each vehicle, like a "fingerprint", and is therefore particularly worthy of protection, whereas the second vector essentially reflects the current value of the vehicle in use and is here of relatively lower protection value from a data protection point of view.

[0025] According to another very preferred configuration of the method according to the invention, as a development of the configuration as vectors, the two characteristic quantities can also be realized in such a way that the respective characteristic quantities are formed as a mapping of the respective vectors into dimensions of a lower order or into dimensions of at most the same order. This mapping may in particular represent a spatial dimensionality reduction. For example, by means of a suitable mapping, such as a principal component analysis, also called PCA (English: Principal Component Analysis), or via a mapping by a deep neural network, a corresponding value with n' or m' dimensions can be generated from a vector with n or m dimensions. In this case, the number of dimensions is in particular smaller or at most the same as the previous number of dimensions, i.e. the number of dimensions does not increase. By mapping a vector with n or m dimensions into a space with n' or m' dimensions, the vector is compressed and encrypted in an appropriate manner, so that it is no longer possible to directly infer the previous vector or the vehicle sensor values ​​on which it is based from the new values. That is to say, if a provider of a service function outside the vehicle can process such vectors or in particular the values ​​generated and mapped from them for the provision of the service function, a certain degree of anonymization can already be achieved by itself. Additionally, when this is combined with the above-mentioned "privacy layer" functionality, very good data protection is achieved by the transmission of related similar information and, if necessary, by the transmission of less relevant and dissimilar information.

[0026] According to a very preferred embodiment of the method according to the invention, furthermore, a machine learning clustering mechanism is used to determine the similarity in order to determine classes of similar and dissimilar characteristic quantities, the characteristic quantities being classified on the basis of a predefined value and on the basis of a comparison of the determined similarity with the predefined value. That is to say, this advantageous embodiment of the method according to the invention uses a machine learning clustering mechanism. These clustering mechanisms may include, for example, k-means, Mean-Shift clustering or EM (Expectation Maximisation) clustering. Via such mechanisms, the similarity for the first characteristic quantity and / or the second characteristic quantity can be generated, for example, as a vector or a mapping of the respective vectors. On the basis of the predefined value, the determined similarity can be classified via the predefined value. If the similarity is, for example, between 0% and 100%, a distinction can be made via a predefined value of, for example, 80%, between dissimilar information between 0% and 80% and sufficiently similar information between 80% and 100%.

[0027] Here, according to a very advantageous embodiment of the method according to the invention, the predetermined value can be varied depending on the external service function of the vehicle. Another possibility for varying or parameterizing the value is, for example, providing the data of the relevant vehicle sensor or its absolute value, because, for example, a different similarity can be used at low speeds than at correspondingly high speeds. Here, as an example, a weather application is considered, in which a relatively low similarity in terms of the exact position is sufficient, so that a position determination of a few kilometers or even a few hundred meters is sufficient for accuracy. However, if the external service function of the vehicle is a navigation system, such information on the position is of course not sufficient, so a much higher similarity of the value is required for the navigation system.

[0028] In one further highly preferred variant of the method according to the invention, the characteristic values ​​are at least partially exchanged between the vehicles and a central data center, with the aggregation and evaluation of the information taking place in the central data center. That is, information from and to the vehicles of the fleet can be exchanged via a data center, in particular a cloud, and the information can be aggregated and evaluated in the data center / cloud. Alternatively or additionally, according to the invention, the characteristic values ​​are at least partially exchanged, aggregated and evaluated between the vehicles of the fleet. In this second approach, the information can be aggregated and evaluated in a distributed manner in the vehicle fleet by information exchange between the vehicles of the fleet. Here, advantageously, the vehicles can communicate directly with other vehicles in the vicinity in a distributed manner, in particular with vehicles with similar dynamic values ​​(i.e. vehicles that are traveling in the same place, for example in the same direction, and therefore have similar current values ​​(e.g. position, speed, acceleration, etc.).

[0029] The first solution reduces data transfer times and latency, while the distributed information processing allows for more secure data (since not all fleet information is stored in a central place) and is robust against failures of individual nodes, especially the central one. Now, the two solutions can be combined, so that one part acts as a central solution and one part as a distributed solution.

[0030] According to a very advantageous configuration of this idea, in order to determine the difference between relevant and less relevant information for the external-vehicle service function, information with similar first characteristic quantities and information with respective dissimilar second characteristic quantities can be transmitted from at least some of the vehicles of the fleet to the external-vehicle service function, and the service function is subsequently evaluated. Here, the vehicle fleet is used to test the external-vehicle service function as to which information is relevant for the result of the service function (profiling). For this purpose, a number of combined queries are transmitted to the provider of the external-vehicle service function with the first and second characteristic quantities covering the possible value range of the query. The responses to the queries transmitted from the provider of the external-vehicle service function are analyzed for similarity in the data center / cloud and / or in the fleet.

[0031] For this purpose, a vehicle of the fleet first sends a query to the provider of the external vehicle service function, the query having similar data in one of two characteristic quantities, but different data in the other characteristic quantity, so that it can be determined whether only one of the two characteristic quantities is sufficient for utilizing the external vehicle service function.

[0032] Accordingly, in a highly advantageous configuration of the method, individual vehicle sensor values ​​are retained and further vehicle sensor values ​​are randomized in order to determine the relevant and less relevant vehicle sensor values ​​for the use of the respective vehicle exterior service function, i.e., individual data or data groups are retained from the characteristic quantities based on the above-mentioned query and the remaining data in the query are randomized in order to arrive at a final group of information relevant for the vehicle exterior service function, i.e., the "privacy layer" further realizes the use of a value range as far away as possible from the respective characteristic quantities.

[0033] That is to say, the method according to the invention is suitable for use in vehicles of a vehicle fleet consisting of substantially homogenous vehicles, i.e. for example in fleets of passenger cars for private use, passenger cars used as public vehicles, vehicles of a brand or brand group, commercial vehicles, etc., where the vehicle according to the invention is equipped with a number of sensors and at least one communication interface designed to implement the method together with other vehicles and / or with an external data center.

[0034] Further advantageous configurations of the method according to the invention as well as of a vehicle designed for carrying out the method according to the invention will become apparent from the exemplary embodiments which are explained in more detail below with reference to the respective figures. [Brief description of the drawings]

[0035] [Figure 1] 1 illustrates a basic scenario for applying the method according to the invention; [Diagram 2] 1 illustrates an exemplary two-dimensional representation of vehicle information. [Diagram 3] 3 illustrates an exemplary one-dimensional compression of the representation of FIG. 2. [Figure 4] 1 shows a first possible embodiment for implementing the method according to the invention; [Diagram 5] 2 shows a second possible embodiment for implementing the method according to the invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0036] 1 shows a first vehicle 11, which is assumed to have a number of sensors, not shown, as well as a communication interface 1, via which a communication connection can be established with a cloud 2, referenced 2, via which a provider of a service function FES external to the vehicle offers its service functions. Between the vehicle 11 and the cloud 2 of the provider of the FES there is a privacy-layer, referenced 3, in which vehicle data for the use of the FES by the vehicle 11 is anonymized. For this purpose, vehicle data from other vehicles, referenced 12, 13, 14, in a fleet of homogenous vehicles, referenced 11, 12, 13, 14, comprising a heterogeneous group of vehicle users, is used.

[0037] In such a fleet 10 of homogenous vehicles 11, 12, 13, 14, and the associated large heterogeneous group of vehicle users, there is a very high probability of similar vehicle data in the form of usage profiles, which may be utilized to realize the privacy layer 3 suggested herein.

[0038] Vehicle sensors installed inside and outside each vehicle 11, 12, 13, 14 of the vehicle fleet 10 continuously determine vehicle sensor values, which contain information about the state of the vehicles 11, 12, 13, 14, such as the vehicle's position, speed, traveled sections, route information, but also about ageing, wear and tear, service waiting periods, etc. Furthermore, via the sensors, the state of the driver of the vehicle 11, 12, 13, 14 or of the passengers can also be determined, for example by detecting the driving style, but also by detecting the attention level of the driver of the vehicle via an attention monitoring module. In this regard, for example, the blinking frequency and / or other body data are evaluated.

[0039] Here, two characteristic vectors are determined for each vehicle 11, 12, 13, 14. The vectors represent the inventive characteristic quantities that depend on the sensor values, which are described here by way of example, and which in principle can also be replaced by a set of sensor values. However, in the following example, the vectors substantially correspond to the characteristic quantities.

[0040] That is, two vectors are determined for each vehicle of the fleet: one characteristic static n-dimensional value WK is determined from recorded vehicle sensor values ​​and is approximately constant over time, which may be, for example, load set data; a second vector is determined from current vehicle sensor values, e.g. position, speed, acceleration, etc., as a dynamic m-dimensional value WZ; this second vector changes correspondingly continuously over time.

[0041] In particular, the first vector WK allows the identification of each vehicle 11, 12, 13, 14 of the vehicle fleet 10, since it is based on the individual vehicle data and the usage characteristics of the vehicle users. The first vector WK is particularly worthy of protection from the point of view of data protection, since it is less time-varying. In contrast, the second vector WZ reflects the current values ​​of the vehicles 11, 12, 13, 14 and the vehicle usage. Identification of the vehicles 11, 12, 13, 14 via these values ​​is only possible indirectly, for example by combining position data with local image information. The second vector WZ is therefore less worthy of protection from the point of view of data protection than the first vector.

[0042] In particular, these two vectors WK, WZ are processed with a spatial dimension reduction by suitable mapping, for example PCA, deep neural network, etc. That is, the n-dimensional WK_n becomes an n'-dimensional space WK', where n'<=n. Thus, the n-dimensional space WK is mapped to the n'-dimensional space WK', which is compressed and encrypted in a predetermined manner. A similar process is performed for WZ.

[0043] In particular, vehicles 11, 12, 13, 14 with similar sensor values ​​in WK' and / or WZ' space can be identified. In an exemplary specific configuration, a machine learning clustering mechanism (e.g., k-means, Mean-Shift clustering, or expectation maximization (EM) clustering) is used to generate similarities in WK and / or WZ space, where the WZ similarity values ​​and the WZ' similarity values ​​for the fleet 10 change over time, so that the WZ similarity or the WZ' similarity is continually identified.

[0044] Correspondingly, in Fig. 2 an exemplary two-dimensional WK representation of similar and dissimilar vehicle information is shown. Similar vehicle information is symbolically represented by dashed areas marked with crosses. These similarities are determined using the above-mentioned clustering mechanism. Here, the similarities are based on a two-dimensional first vector WK merely for illustrative purposes, but of course the corresponding second vector WZ can be represented in a similar manner. The number of dimensions can be much higher than the two dimensions shown here, but two dimensions are particularly well suited for graphic display.

[0045] Correspondingly, in Fig. 3, a one-dimensional compression of what is shown in Fig. 2 can be seen, i.e., a spatial dimension reduction or mapping WK' is shown accordingly. Here again, similar vehicles 11, 12, 13, 14 or vehicle information are again shown between the dashed lines and marked with a cross. Meanwhile, the vehicle data outside these dashed lines, i.e. the vehicle data to the right and left of the dashed lines, without additional marking, are dissimilar vehicle data.

[0046] In a first variant as illustrated in Fig. 4, information from and to the vehicles 11, 12, 13, 14 of the fleet 10 can be exchanged via a data centre / cloud 4, in which it can be aggregated and evaluated. This central data centre 4 represents a trusted instance, which here can be, for example, a back-end server of the vehicle manufacturer.

[0047] When a vehicle-external service function FES is called by a first vehicle 11, instead of the vehicle information 5 of that vehicle 11, vehicle information 6 derived from the WK or WZ of similar vehicles 12, which is relevant for the service, or information 6 determined from those WK or WZ (for example by averaging) is transmitted. If the vehicle information 5 of a vehicle 11 is not relevant for the service, vehicle information 7 derived from the WK and WZ of non-similar vehicles 13, or a random value, is transmitted instead of the vehicle information 5 of that vehicle 11. Hereby, a "privacy layer" is realized, which replaces the information 5 of the vehicle 11 with the information 6 of similar vehicles 12, which is relevant for the FES, and with non-similar irrelevant information 7, thus hiding the original information 5 of the vehicle.

[0048] To improve the artificial generation of random WZ, generative methods of machine learning (e.g., generative adversarial networks, Encoder-Decoder networks) can be used. For this, models optimized to generate target vectors as "realistic" as possible can be trained and used for random WZ values. This "realism" is achieved by using appropriate features that are defined generally or according to the services that the vehicle driver wants to use. That is, the content focus of the m-dimension can be placed.

[0049] In another embodiment, it is conceivable that the FES is only allowed access to the space WK' or WZ' and the mapping rules, so that the information in WK' is present in a compressed state (see FIG. 3) and a non-reversible mapping is used, so that the user-specific values ​​of WK cannot be reproduced without further knowledge. In this case, it is important for the dimensionality reduction of the space WK or WZ to WK' or WZ' to combine the information of a large number of vehicles 11, 12, 13, 14 to achieve a sufficiently good compression to identify similar vehicles 11, 12 in this space and thus guarantee anonymity.

[0050] Alternatively or additionally, the information 5 can be aggregated and evaluated in a distributed manner in a vehicle fleet 10 by information exchange between the vehicles 11, 12, 13 of the fleet. This is shown in figure 5, which is similar to figure 4. While the first solution reduces data transfer times and latency, distributed information processing allows for greater data security (since not all information of the fleet 10 is stored in a central place) and is robust against failures of individual nodes, especially the central node.

[0051] In the following, the vehicle fleet 10 is used to test the FES as to which information is relevant for the result of the service function. For this purpose, a number of queries combined are sent to the provider or cloud 2 with values ​​WK and WZ covering the possible value range of the query. The responses to the query transmitted from the FES are analyzed for similarity in the data center / cloud 4 and / or in the fleet 10. For this purpose, queries are sent to the FES from a group of vehicles 11, 12, 13, 14 of the fleet 10, which have similar values ​​in one of the two value ranges (WK or WZ), but different values ​​in the respective other value range. This makes it possible to determine whether values ​​of only one category are sufficient for the use of the FES. On this basis, individual values ​​or value groups of WK and WZ are kept and the remaining values ​​in the query are randomized in order to arrive at a final group of vehicle information relevant for the use of the FES. That is to say, it is also possible to use a value range as far away as possible from WK.

[0052] In the use case of a weather app provided by an FES, the information relevant to the application of the FES is determined by a combined fleet query. Here, the location data and the destination address stored in the navigation system as well as the arrival time predicted by the navigation system are identified as relevant information. Now, if a first vehicle 11 uses the weather app, instead of the location data of the first vehicle 11, the location data of a second vehicle 12 of a class with similar location data, or the location data generated from the averaging of the vehicles 12 of the class with similar location data, is transmitted. Similarly, instead of the destination address of the first vehicle 11, the destination address of a third vehicle 12 of a class with similar destination addresses, or the destination address averaged with the number of similar third vehicles 13, is transmitted to the weather app. Similarly, the arrival time to the destination of a fourth vehicle 12 with a similar arrival time, or the averaged arrival time of the class of similar vehicles 12, is used. For data not relevant to the use of the FES, data from (or averaged from) other vehicles 13 in the class that do not have similarity to the first value of the first vehicle 11 can be used.

Claims

1. 1. A method for anonymizing vehicle data for use in a vehicle external service function (FES), comprising: In a fleet (10) of homogenous vehicles (11, 12, 13, 14) comprising a heterogeneous group of vehicle users, a first characteristic quantity is determined for each of said vehicles (11, 12, 13, 14), the first characteristic quantity being at least indirectly dependent on an integer number n of recorded vehicle sensor values ​​and a second characteristic quantity being at least indirectly dependent on an integer number m of current vehicle sensor values, a similarity is determined for at least one characteristic value of a plurality of characteristic values ​​of a plurality of or all of the vehicles (11, 12, 13, 14) of the fleet (10), and the vehicles (11, 12, 13, 14) are classified accordingly into similar or dissimilar classes on the basis of the similarity of the at least one characteristic value; The method according to claim 1, characterized in that when a service function (FES) external to the vehicle is requested, instead of a characteristic value of the requesting vehicle (11) containing information relevant to the provision of the service, a corresponding characteristic value of a vehicle (12) of a class of similar vehicles (12), a characteristic value mathematically determined from the corresponding characteristic values ​​of a plurality of vehicles (12) of the class of similar vehicles (12), or an artificially generated similar characteristic value is transmitted.

2. 2. The method according to claim 1, characterized in that if the service function (FES) external to the vehicle requires less relevant information for the execution of the service function, which is based on the respective further characteristic values, then instead of the characteristic values ​​of the requesting vehicle (11), corresponding characteristic values ​​of a vehicle (13) from a class of dissimilar vehicles (13), characteristic values ​​mathematically determined from corresponding characteristic values ​​of a plurality of vehicles (13) from the class of dissimilar vehicles (13), or similar characteristic values ​​that are artificially generated are transmitted.

3. The method according to claim 1 or 2, characterized in that the random characteristic quantities are generated via generative methods of machine learning.

4. 4. The method according to claim 1, 2 or 3, characterized in that each said characteristic quantity is formed as a set of n or m vehicle sensor values.

5. 4. The method according to claim 1, 2 or 3, characterized in that each of the characteristic quantities is formed as a vector (WK, WZ) based on the respective vehicle sensor values ​​in an n-dimensional or m-dimensional space.

6. 6. The method according to claim 5, characterized in that each of said characteristic quantities is formed as a mapping of each of said vectors (WK, WZ) to values ​​(WK', WZ') having a smaller number of dimensions or at most the same number of dimensions.

7. 7. The method according to claim 1, wherein a machine learning clustering mechanism is used to determine similarities in order to determine classes of similar and dissimilar features, the features being classified based on a predefined value and based on a comparison of the determined similarities with the predefined value.

8. 8. The method according to claim 7, characterized in that the predetermined value is parameterized or defined depending on an associated vehicle sensor value, an absolute value of the vehicle sensor value, and / or the vehicle external service function (FES).

9. 9. The method according to claim 1, characterized in that the characteristic quantities are at least partially exchanged between the vehicles (11, 12, 13, 14) and a central data center (4), and the aggregation and evaluation of the information takes place in the central data center (4).

10. 10. The method according to claim 1, wherein the characteristic quantities are at least partially exchanged, aggregated and evaluated between the vehicles (11, 12, 13, 14) of the fleet (10).

11. 11. The method according to claim 1, further comprising transmitting information having similar first characteristic quantities and information having respectively dissimilar second characteristic quantities from at least some of the vehicles (11, 12, 13, 14) of the fleet (10) to the vehicle external service function (FES) in order to determine a difference between relevant and less relevant information for the vehicle external service function (FES), after which the service function is evaluated.

12. 12. The method of claim 11, further comprising: retaining individual vehicle sensor values; and randomizing the individual vehicle sensor values ​​to determine relevant and less relevant vehicle sensor values ​​for each of the vehicle exterior service functions (FES) usages.

13. A vehicle (11, 12, 13, 14) equipped with a number of sensors and at least one communication interface (1), said vehicle (11, 12, 13, 14) being designed to carry out the method together with other vehicles (11, 12, 13, 14) and / or an external data center (4).

Citation Information

Patent Citations

  • Procedures for protecting the personal data of a vehicle occupant

    DE102020003188A1

  • Method and system for storing and transmitting measurement data from measuring vehicles

    EP3492872A1

  • Communication system, terminal device, privacy protection device, privacy protection method, and program

    JP2017151942A

  • Autonomous Vehicle Systems

    JP2022524920A

  • Method and device for anonymizing position and movement data of a motor vehicle

    DE102015213393A1