Vehicle Data Anonymization Through Similarity-Based Fleet Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for anonymizing vehicle data for vehicle-external services do not adequately protect personal data while enabling continued use of such services, particularly in fleets with diverse user groups where similar vehicle use profiles exist.
Innovation Solution
A method utilizing machine learning clustering mechanisms to categorize vehicles into similar and dissimilar classes based on static and dynamic sensor data variables, replacing sensitive information with aggregated or artificially generated data from similar vehicles, and employing spatial dimension reduction and generative machine learning to create a privacy layer.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If vehicle data is transmitted to external services, then service functionality is enabled, but personal data protection is compromised
Solution Approach 1:
The patent introduces a data intermediary layer that processes vehicle data before transmission. This intermediary anonymizes the data by removing or generalizing personally identifiable information while preserving the operational characteristics needed for service functionality. The intermediary acts as a buffer between the vehicle data and external services, enabling service usage without direct exposure of personal data.
Solution Approach 2:
The patent creates anonymized copies of vehicle data that retain the essential patterns and characteristics needed for service analysis. Instead of transmitting original data containing personal information, the system generates synthetic or aggregated copies that preserve statistical properties and operational patterns while eliminating identifiable personal data. These copies can be used by external services without compromising individual privacy.
2Measurement precision
If detailed vehicle sensor data is collected, then data accuracy for analysis is improved, but data protection requirements increase
Solution Approach 1:
The patent segments vehicle data into different categories based on their sensitivity and utility. Critical operational data that requires high precision for service functionality is separated from personally identifiable information. This segmentation allows the system to apply different processing levels: high-fidelity processing for operational parameters and anonymization or aggregation for personal data, thereby maintaining necessary accuracy while reducing protection overhead.
Solution Approach 2:
The patent transforms data parameters by applying mathematical transformations, aggregation functions, or generalization operations. Continuous sensor data may be aggregated into statistical summaries (mean, variance, ranges) that preserve analytical value while reducing granularity of personal information. Time-series data may be transformed into frequency-domain representations or pattern descriptors that maintain diagnostic capability while obscuring individual usage patterns.
Data Source
AI summary
Vehicle data for the use of vehicle-external services is anonymized. Within a fleet of vehicles of the same type with a diverse group of vehicle users, a first variable, which is at least indirectly dependent on an integer number n of recorded vehicle sensor values, and a second variable, which is at least indirectly dependent on an integer number m of current vehicle sensor values, is recorded for each of the vehicles. Similarities are determined for at least one of the variables of several or all vehicles in the fleet, after which the vehicles are then categorized into a class that is similar in terms of the similarity of the at least one variable or into a dissimilar class. If a vehicle-external service is requested, a computed variable from the corresponding variables of several vehicles from the class of the similar vehicles or an artificially generated similar variable is transmitted.

