Method and system for real-time vehicle health monitoring and predictive maintenance

The system addresses the limitations of existing driver behavior monitoring by using machine learning and real-time analytics to provide accurate, adaptive vehicle health monitoring and predictive maintenance, ensuring safety and efficiency across diverse conditions.

WO2026154516A1PCT designated stage Publication Date: 2026-07-23SRIVASTAVA NAVROOP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SRIVASTAVA NAVROOP
Filing Date
2026-01-16
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Existing driver behavior monitoring systems fail to capture complex dynamics of vehicle-driver interaction, lack robustness in detecting subtle shifts, and are not adaptable to different drivers or environments, leading to inadequate safety and maintenance predictions.

Method used

A vehicle health monitoring and predictive maintenance system using machine learning algorithms and real-time data analytics, incorporating a hybrid anomaly-detection engine with autoencoders, graph-neural networks, and digital-twin reconstruction to provide accurate maintenance insights and alerts, adaptable to varying sensor configurations and driving conditions.

Benefits of technology

Enables precise prediction of vehicle failures, personalized monitoring, and timely alerts, improving safety and reducing operational costs by continuously adapting to different drivers and environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IN2026050072_23072026_PF_FP_ABST
    Figure IN2026050072_23072026_PF_FP_ABST
Patent Text Reader

Abstract

This invention provides a vehicle health monitoring and predictive maintenance system that integrates real-time data analytics, anomaly detection, and hybrid machine learning models. The system utilizes data from vehicle sensors (OBD-II, GPS, and accelerometers) to monitor health, detect irregularities, and predict maintenance needs. A unique combination of ARIMA, SVR, and XGBoost models enables precise forecasting of failures and maintenance requirements, while Isolation Forest ensures efficient anomaly detection. The invention adapts to individual driving behaviours and environmental conditions, generating user- specific maintenance insights and cost-saving recommendations. The generated insights are delivered through user-friendly reports, promoting proactive maintenance strategies that enhance vehicle safety, reduce downtime, and optimize performance. This scalable solution is applicable across industries, including fleet management, autonomous vehicles, and personal vehicle maintenance. By combining IoT, advanced predictive modelling, and real- time insights, this invention addresses limitations in traditional vehicle diagnostics and sets a new standard for predictive maintenance systems.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] FIELD OF INVENTION

[0002] The present invention pertains to the field of automotive diagnostics and predictive maintenance systems.

[0003] BACKGROUND OF THE INVENTION

[0004] Prior to the present invention, systems for monitoring driver behavior were typically based on limited sensor data such as steering angle, vehicle speed, or camera-based systems. These systems often failed to account for the complex dynamics of driver-vehicle interaction over time and under various driving conditions. Traditional methods relied on simple rule-based algorithms or models that did not adequately subtle changes in driver behavior, such as drowsiness, distraction, or stress.

[0005] US9663047B2 involves a method for inferring the behavior or state of the driver through the use of inertial sensors, including accelerometers and gyroscopes, integrated into an autonomous computation device within the vehicle. This method evaluates driver behavior by comparing the data from these sensors to reference information associated with normal driving behaviors for a given road and vehicle condition. Such systems predict abnormal driver states, such as drowsiness or distraction, by comparing descriptive functions derived from sensor data to established reference values. These functions include standard deviations in lateral position, frequency of zigzags, and variations in speed or acceleration. In some implementations, additional sensors like cameras, temperature sensors, and GPS may be used to enrich the data collected from the inertial sensors.

[0006] Despite their advancements, the prior art systems have several limitations. First, they often rely on relatively basic sensors and models that do not fully capture the complex dynamics of the vehicle-driver interaction. The reference information used in these systems is often limited to data collected from a specific driver or vehicle under certain conditions, which may not generalize well to different drivers, vehicles, or environments. Furthermore, there is a lack of robust systems capable of detecting subtle and complex shifts in driver behavior that can precede dangerous driving conditions.

[0007] The limitations of existing driver behavior monitoring systems raise the need for the present invention.OBJECT OF THE INVENTION

[0008] An object of the present invention is to provide a vehicle health monitoring and predictive maintenance system that leverages machine learning techniques and real-time data analytics to improve vehicle reliability, safety, and efficiency.

[0009] SUMMARY OF THE INVENTION

[0010] The present invention discloses a vehicle health monitoring and predictive maintenance system that leverages software and machine learning algorithms to transform vehicle sensor data into actionable maintenance insights. The system is designed to operate with minimal hardware requirements and relies on data collected from onboard diagnostics (OBD-II) systems, utilizing a multi-step machine learning pipeline to deliver accurate predictions for potential failures and maintenance needs.

[0011] A core aspect of the system is a hybrid anomaly-detection engine that combines forecast residuals, autoencoder-based reconstruction residuals, digital-twin reconstruction errors, and graph-neural network sensor-risk scores into a mask-gated feature vector. A learned fusion function maps this vector to a calibrated anomaly probability, and a decision threshold is chosen on validation data so that the operational false-positive rate respects a predefined bound. This provides the technical effect of predictable alerting behavior under noisy and partially missing telemetry.

[0012] BRIEF DESCRIPTION OF FIGURES

[0013] Figure 1 illustrates the overview architecture of the system.

[0014] Figure 2 represents the flow of data from sensors to predictive models.

[0015] Figure 3 discloses the anomaly detection module.

[0016] Figure 4 represents the workflow of hybrid predictive model.

[0017] Figure 5 illustrates the architecture of the system comprising various interconnected components.

[0018] Figure 6 presents the process through which the system identifies irregularities.

[0019] Figure 7 outlines the design of the API-driven backend.

[0020] Figure 8 outlines the flow of machine learning model development, deployment, monitoring, and retraining within the platform.

[0021] Figure 9 presents a detailed overview of the data flow, processing, analysis, and presentation layers.Figure 10 showcases the components of the system.

[0022] Figure 11 illustrates an exemplary circuit-level block diagram of a vehicle-mounted OBD-II telemetry device showing major electrical subsystems and their interconnections.

[0023] Figure 12 illustrates an exemplary functional architecture of the vehicle-mounted telemetry device, grouped into functional subsystems.

[0024] DETAILED DESCRIPTION OF THE INVENTION

[0025] The present invention provides a comprehensive system for vehicle health monitoring and predictive maintenance, leveraging advanced machine learning algorithms and Internet of Things (loT) technologies to analyze and process vehicle sensor data. This system enables the prediction of vehicle failures, identifies maintenance needs in advance, and ensures optimal vehicle performance, reducing the occurrence of unexpected breakdowns and minimizing operational costs. It is a next-generation vehicle health intelligence system that leverages advanced Al to offer predictive diagnostics, prescriptive maintenance forecasting, and driver behavior analysis. It is designed to support heterogeneous vehicle data from a variety of telematics systems and sensor setups, ensuring robust performance across commercial and consumer fleets.

[0026] The objective of the present invention is to provide a vehicle health monitoring and predictive maintenance system that ingests multi-source automotive telemetry and computes model-based residuals that are mask-gated to handle missing sensors, fuses these residuals to produce calibrated anomaly probabilities with decision thresholds selected to satisfy a target falsepositive rate, generates component-level health scores, remaining-useful-life estimates, and prescriptive maintenance recommendations and supports a consumer cloud-only implementation in which the vehicle device performs no machine-learning inference, with an optional edge-enhanced implementation for commercial deployments that provides low-latency checks while the cloud remains authoritative.

[0027] The present invention relates to a telematics-based vehicle health monitoring system that performs mask-gated residual fusion of multi-model anomaly features under varying sensor availability and heterogeneous telemetry schemas, while enforcing a deployment-level falsepositive rate constraint.

[0028] In an embodiment of the present invention discloses a system (200) for real-time vehicle health monitoring and predictive maintenance, comprising:i. a vehicle-mounted telemetry device (202) comprising:

[0029] a. a microcontroller or microprocessor (302);

[0030] b. an OBD-II interface (304) coupled to a vehicle communication bus (306) and configured to acquire engine (308) and diagnostic parameters;

[0031] c. a positioning receiver (310) configured to determine geographic location;

[0032] d. an inertial sensor (312) configured to measure linear acceleration and angular velocity of the vehicle; and

[0033] e. a wireless communication module (314) configured to connect to a mobile data network; and

[0034] ii. a processing unit (204) located external to the vehicle;

[0035] wherein the vehicle-mounted telemetry device is a physical electronic apparatus configured to be removably coupled to an on-board diagnostics (OBD-II) port of a vehicle and to directly receive, during operation of the vehicle, electrical signals representing engine operating parameters and diagnostic status from a vehicle communication bus,

[0036] wherein the microcontroller or microprocessor is configured to synchronize, at predefined sampling intervals, the engine and diagnostic parameters received through the OBD-II interface with geographic position signals from the positioning receiver and motion signals from the inertial sensor to generate time-aligned telemetry data corresponding to real-world vehicle operation,

[0037] wherein the wireless communication module is configured to transmit the time-aligned telemetry data over a mobile data network to the processing unit external to the vehicle,

[0038] wherein the processing unit comprises physical processors and a non-transitory memory storing executable instructions which, when executed by the processors, cause the processing unit to perform predictive analysis and anomaly detection on the telemetry data so as to derive component-level health indicators and maintenance requirement signals, and

[0039] wherein the derived component-level health indicators are used to electronically generate and transmit machine-readable maintenance alerts and reports to a user device,

[0040] thereby achieving a technical effect of continuous electronic monitoring of physical vehiclesubsystems, early detection of deviations in vehicle operating behavior, and reduction of unexpected mechanical failures through processor-controlled analysis of sensor-derived vehicle telemetry.

[0041] The telemetry devices for different vehicle models provide different subsets of sensor channels, and the processing unit is configured to accommodate such variation by mapping available sensor channels from each vehicle into a common feature representation that indicates presence or absence of predefined feature slots using validity indicators and applying learned fusion model to the common feature representation without retraining separate models for each distinct vehicle sensor schema.

[0042] The microprocessor further comprises a memory storing firmware instructions that, when executed by the microcontroller or microprocessor, cause the apparatus to periodically sample signals from the OBD-II interface, the positioning receiver, and the inertial sensor to generate telemetry data; perform local pre-processing operations including formatting and basic validation of the telemetry data; and transmit the telemetry data to a remote processing unit for hybrid predictive analytics, anomaly detection, and report generation, without performing the hybrid predictive analytics and anomaly detection models on the apparatus.

[0043] The wireless communication module comprises a cellular modem with integrated Global Navigation Satellite System (GNSS) capability and is configured to transmit telemetry data to the remote processing unit using a publish-subscribe messaging protocol.

[0044] The processing unit comprises a processor and a memory storing instructions that, when executed by the processor, cause the processing unit to receive the telemetry data from the telemetry device; pre-process the telemetry data by performing noise filtering, outlier removal, scaling, time alignment, and handling of missing values to produce cleaned time series data; apply a hybrid predictive analytics model comprising an Autoregressive Integrated Moving Average (ARIMA) model, a Support Vector Regression (SVR) model, and an Extreme Gradient Boosting (XGBoost) model to the cleaned time-series data to generate predictions of vehicle operating variables; compute residual signals based on differences between the predictions and corresponding observed telemetry values; provide features derived from the cleaned time-series data and the residual signals to an anomaly detection module comprising an Isolation Forest estimator configured to assign anomaly scores to the telemetry data;determine, based on the anomaly scores and the predictions, a health state of a vehicle component and a predicted maintenance requirements; and generate and transmit to a user device a report including the health state and the predicted maintenance requirements.

[0045] The processing unit is configured to:

[0046] i. receive, for each of time windows, multivariate telemetry comprising sensor measurements from a vehicle, the sensor measurements including a subset of OBD-II parameters and inertial measurements;

[0047] ii. for each time window, apply at least two predictive models of different model families selected from temporal forecasting models and reconstruction-based models to the telemetry to obtain model outputs, the model families including a temporal forecaster configured to predict future or current sensor values from historical telemetry and a reconstruction-based model or digital-twin model configured to reconstruct sensor values from contemporaneous telemetry;

[0048] iii. compute, for the time window, residual signals for the temporal forecaster and reconstruction-based model by comparing the model outputs to corresponding observed values;

[0049] iv. compute, for each sensor channel, a validity indicator that encodes whether the corresponding measurement in the time window is present and within predefined plausibility bounds;

[0050] v. compute residual-based features that are normalised by the number of valid channels, including feature obtained by aggregating residual magnitudes over only those sensor channels whose validity indicators denote valid data;

[0051] vi. construct, for the time window, a feature vector comprising summary statistics of the residual signals over the window, summary statistics of reconstruction-error features over the window and statistic indicative of a proportion of valid measurements in the window;

[0052] vii. provide the feature vector to the learned fusion model that outputs a scalar anomaly score or calibrated anomaly probability for the time window; and

[0053] viii. select anomaly decision threshold for the learned fusion model using a validation dataset that contains predominantly non-fault telemetry such that a false-positive rate measured on the validation dataset does not exceed a predefined budget and apply the anomaly decision threshold during deployment to generate anomaly alerts.The processing unit is further configured to detect temporal patterns in the residual signals and the anomaly scores that persist over multiple time windows; and escalate a maintenance recommendation when such persistent patterns exceed thresholds associated with a likelihood of future failure.

[0054] The processing unit is deployed in a cloud computing environment and is configured to manage telemetry and analytics for a plurality of vehicles, including fleet vehicles, while maintaining per-vehicle histories of telemetry, predictions, anomalies, and maintenance recommendations.

[0055] The telemetry data further comprises driver behavior features, including braking intensity, acceleration variability, and cornering characteristics, and the processing unit is configured to incorporate the driver behavior features as inputs to the hybrid predictive analytics model and / or the anomaly detection module.

[0056] The telemetry device and the processing unit communicate using a publish-subscribe messaging protocol over a mobile or wireless network, and data in transit is protected using authenticated, encrypted communication.

[0057] The system further comprising a user-facing application executing on a user device, the userfacing application being configured to receive the report generated by the processing unit; display summary health status, predicted maintenance events, and driver behavior insights; and generate real-time alerts when anomaly scores or predicted failure probabilities exceed thresholds.

[0058] The system further comprising an edge processing node located within or near at least one of the vehicles, the edge processing node being configured to locally perform a subset of the preprocessing and anomaly detection steps to support lower-latency alerts in environments with intermittent or limited network connectivity, and to forward summarized results and raw data to the cloud-based processing unit for long-term analysis and model retraining.

[0059] The system is housed in a compact enclosure configured to plug into an OBD-II port of the vehicle and is powered from the vehicle electrical system while being configured to operate within an automotive temperature and vibration range.The reconstruction-based model comprises a surrogate model of a vehicle subsystem, trained to approximate physical relationships between a plurality of sensor channels forthat subsystem, and the reconstruction-error features include multi-sensor deviations from the surrogate model outputs.

[0060] The learned fusion model comprises a linear model parameterized by a weight vector and a bias term applied to the feature vector, followed by a non-linear activation function that maps a result to an anomaly likelihood value, and wherein parameters of the learned fusion model are obtained by training on labelled historical data comprising examples of nominal vehicle operation and examples of known faults.

[0061] The processing unit is configured to calibrate raw outputs of the learned fusion model using a calibration function trained on held-out validation data to produce calibrated anomaly probabilities, and to choose the anomaly decision threshold as a quantile of the calibrated anomaly probabilities on negative validation samples corresponding to a target false-positive rate.

[0062] The processing unit is further configured to model correlations between sensor channels as a graph with nodes representing sensors and edges representing learned or engineered dependencies, to compute node-level or graph-level anomaly scores using a graph-based model, and to include summary statistics of the anomaly scores as additional components of the feature vector. In the present invention the hardware components selected for this loT device are for vehicle diagnostics.

[0063] In the present invention the components are interconnected through various protocols: the microcontroller connects via UART to both the module (eg: BG96 module) and the IC (eg: MCP2562FD IC), while the sensor (eg: BMI088 sensors) is connected via I2C. The power supply distributes stable 2-4V and 4-6V to all components. Flexible PCB antennas are embedded in the housing for GPS and cellular communication, reducing external bulk. There can be multiple processors in the system of the present invention.

[0064] In the present invention the loT device is designed to integrate real-time data collection,processing, and communication functionalities while maintaining a compact form factor. It supports multiple protocols including CAN-FD, LTE, and MQTT / HTTP, making it suitable for modem vehicles. The device combines cellular and GNSS functionalities into a single module, streamlining the design. It also captures motion data using the 6-axis accelerometer / gyroscope for anomaly detection.

[0065] The design of the device is optimized for space efficiency. All components are placed on a 4-layer PCB to minimize the footprint and reduce noise. Additionally, flexible PCB antennas are embedded within the housing, contributing to a sleeker overall design. Power management is a key consideration, with efficient DC-DC conversion ensuring low power consumption, a critical feature for loT applications that require long battery life and minimal energy demands.

[0066] The present invention introduces the integration of inertial sensors with advanced algorithms that evaluate multiple descriptive functions of driver behavior, including but not limited to lateral position, acceleration, and angular velocity. These descriptive functions are compared with reference values that have been specifically tailored to the vehicle's operational environment and the driver's historical behavior, providing a much more accurate representation of what constitutes "normal" driving for a given driver and context.

[0067] Additionally, the present invention improves on the prior art by allowing for continuous calibration and adaptation of the reference values used for comparison. These values are dynamically generated from an initial characterization process and are updated based on ongoing data collection, making the system adaptable to different drivers and driving conditions. This allows the system to provide more personalized monitoring and early warning systems, rather than relying on fixed, one-size-fits-all reference data.

[0068] Overall, the present invention offers a more comprehensive, adaptive, and accurate solution for inferring driver behavior, combining inertial sensors and environmental context to ensure greater safety and more timely alerts regarding potential driver impairment.

[0069] The present invention has been incorporated with a comprehensive driver behavior analysis module that continuously monitors and profiles driver interactions with the vehicle. This subsystem analyzes longitudinal control patterns such as acceleration, braking, cornering, idling, and throttle modulation to assess driving style characteristics.Machine learning-based behavioral classifiers are employed to detect aggressive, inefficient, or unsafe driving behaviors. Metrics such as harsh braking frequency, rapid acceleration episodes, prolonged idling, and sharp cornering are quantified and benchmarked against optimized safe-driving standards. Behavior-based risk scores are generated for each driver profile, and these scores are integrated into the predictive maintenance and health forecasting models. For instance, drivers exhibiting consistently aggressive behaviors are associated with accelerated brake wear, increased fuel consumption, and higher tire degradation rates. This enables not only vehicle health prediction but also behavior-driven prescriptive recommendations for improving driving habits, optimizing maintenance intervals, and extending overall vehicle lifespan. In fleet applications, aggregated behavior insights facilitate driver training programs, insurance risk profiling, and operational cost reduction strategies.

[0070] In another embodiment of the present invention discloses a method for real-time vehicle health monitoring and predictive maintenance, comprising the following steps:

[0071] i. receiving telemetry data transmitted over a wireless network from the telemetry device mounted in a vehicle, the telemetry data comprising OBD-II parameters, location information, and inertial measurements;

[0072] ii. pre-processing the telemetry data by performing noise filtering, outlier removal, scaling, time alignment, and handling of missing values to produce cleaned time series data;

[0073] iii. applying a hybrid predictive analytics model comprising an ARIMA model, an SVR model, and an XGBoost model to the cleaned time-series data to generate predictions of vehicle operating variables;

[0074] iv. computing residual signals between the predictions and corresponding observed telemetry values;

[0075] v. providing features derived from the cleaned time-series data and the residual signals to the anomaly detection module comprising an Isolation Forest estimator configured to assign anomaly scores;

[0076] vi. determining, based on the anomaly scores and the predictions, a health state of the vehicle component and predicted maintenance requirements; and

[0077] vii. generating a report including the health state and the predicted maintenance requirements and transmitting the report to a user device.The pre-processing in step ii comprises imputing missing values in the telemetry data using historical statistics, interpolation, or model -based estimates, and discarding samples that violate predefined physical plausibility constraints.

[0078] The method further comprising deriving from the telemetry data, driver behavior features including braking intensity, acceleration variability, or cornering events and using the driver behavior features as additional inputs for the hybrid predictive analytics models and the anomaly detection module when performing steps (iii) through (vi).

[0079] The anomaly detection module is configured to receive as input a feature vector comprising the residual signals, current telemetry values, and summary statistics over a time window; and output anomaly scores that indicate the degree of deviation of current operating behavior from learned normal behavior of the vehicle.

[0080] Determining the health state of the vehicle component in step vi comprises computing a health score as a function of the anomaly scores and residual signals and comparing the health score to thresholds mapped to discrete health categories.

[0081] The report includes component-level health indicators, predicted time-to-service or mileage-to-service for the vehicle component and driver behavior scores and recommendations to reduce wear and improve safety.

[0082] The method further comprising:

[0083] i. maintaining historical records of telemetry, predictions, and anomaly scores for each vehicle;

[0084] ii. periodically retraining any of the ARIMA, SVR, and XGBoost models using the historical records to improve prediction accuracy for that vehicle or for a group of similar vehicles; and

[0085] iii. updating parameters of the Isolation Forest estimator based on newly observed normal operating data to maintain alignment with changing operating conditions.

[0086] The method further comprising, prior to training the hybrid predictive analytics model and the anomaly detection module:

[0087] (a) identifying historical telemetry segments associated with rare failure events;

[0088] (b) generating synthetic variants of such telemetry segments by perturbing or resamplingthe identified segments while preserving their characteristic failure patterns; and (c) augmenting a training dataset with the synthetic variants to increase an effective number of failure examples used for training.

[0089] In the present invention as part of the system enhancements, the XGBoost component can be replaced with Long Short-Term Memory (LSTM) networks to better capture temporal dependencies and long-range correlations in vehicle sensor data.

[0090] In the enhanced hybrid framework, ARIMA models are responsible for capturing linear temporal patterns in individual sensor time series, while Support Vector Regression (SVR) handles non-linear patterns with high precision for small and medium datasets. The LSTM models introduce deep learning-based sequence modeling, capable of learning complex temporal behaviors, sensor drift patterns, and latent degradation trends that traditional models may fail to capture.

[0091] This three-pronged hybrid approach, combining statistical, machine learning, and deep learning methodologies, improves prediction accuracy for vehicle operational parameters, component wear levels, and performance deterioration rates. It also enhances resilience against data noise and missing values, ensuring more reliable forecasting outcomes for maintenance planning and anomaly detection.

[0092] The system employs an ensemble learning mechanism that intelligently combines the outputs from multiple independent predictive models to derive a unified health risk score and anomaly prediction. The models contributing to the ensemble include transformer-based multi -sequence regression models, variational autoencoder (VAE) based anomaly reconstruction scoring, digital twin models predicting multi-sensor reconstruction residuals, traditional time-series forecasting models (ARIMA, SVR, LSTM) and graph neural network (GNN)-based node-level risk predictions.

[0093] Each model’s outputs are normalized and aggregated using dynamically weighted ensemble techniques based on recent validation performance. The ensemble is designed to maximize robustness, ensuring that transient failures or temporary degradations in individual models do not cascade into system-wide inaccuracies. If a particular model fails to produce an output (e.g., missing data, corrupted inputs), the ensemble fallback mechanism re-weights predictionsaccordingly, ensuring continuous operation. The resulting aggregated health score and sensor failure risk scores are more accurate, resilient to sensor anomalies, and better suited for deployment in both consumer vehicles and commercial fleet environments. Additionally, the ensemble enables cross-verification of sensor degradation signals, thereby reducing false positives and enhancing early fault detection capabilities.

[0094] The system in the present invention integrates graph neural networks (GNNs), specifically utilizing Graph Attention Network (GAT) architectures, for sensor failure risk prediction. In this subsystem, each vehicle sensor feature is represented as a node in a fully connected or partially connected graph structure. The edges between nodes encode relationships and interactions between sensor readings, capturing complex interdependencies that traditional feature modeling cannot easily represent.

[0095] The GNN learns to propagate information across sensor nodes, updating feature embeddings based on both individual sensor status and the contextual status of neighboring sensors. Attention mechanisms are employed to assign higher importance to more influential sensor relationships during prediction. The trained GNN model outputs node-level anomaly risk scores, quantifying the likelihood of failure or abnormal behavior for each sensor independently. This allows the system to identify localized degradation patterns even when global vehicle-level indicators remain normal.

[0096] Furthermore, the GNN model supports dynamic feature graphs, enabling adaptation to varying sensor configurations across different vehicle types. It significantly enhances failure prediction sensitivity, especially for subtle, localized anomalies that would otherwise evade detection in traditional sequential modeling approaches.

[0097] The system is also designed with built-in fallback mechanisms to ensure continuous functionality even when advanced deep learning models such as Variational AutoEncoders (VAEs) or Graph Neural Networks (GNNs) are unavailable or have not been trained for a particular vehicle profile. In such cases, the system automatically transitions to traditional machine learning techniques such as Principal Component Analysis (PCA) and unsupervised clustering algorithms (e.g., KMeans, DBSCAN) to approximate anomaly detection and latent feature extraction. PCA-based dimensionality reduction identifies the principal axes of variance in the sensor dataset, enabling the system to capture major trends and deviations in alow-dimensional latent space. Anomalies are flagged based on reconstruction errors or deviations from principal subspaces. Clustering methods provide an unsupervised segmentation of operational patterns, where abnormal driving sessions or rare operational modes can be detected as outliers relative to dominant clusters.

[0098] By incorporating these fallback techniques, the system of the present invention ensures minimal performance degradation and operational continuity even when resource-intensive or highly specialized deep learning models are not accessible, thus enhancing system resilience and deployment flexibility.

[0099] To enhance usability, the system includes user-friendly reporting via a mobile application. This app visualizes predictions and maintenance recommendations in an intuitive format, making it easier for users to manage vehicle maintenance. Finally, the design prioritizes energy efficiency. The advanced PMIC and low-power modules ensure that the device is loT- ready, while maintaining minimal energy demands, which is essential for long-term, sustainable operation.

[0100] A key innovation of the system lies in its ability to dynamically adapt to vehicles with differing sensor configurations, data schemas, and operational characteristics, without requiring manual reconfiguration or full model retraining. Upon data ingestion, the system identifies the available set of sensor features for the specific vehicle profile. It dynamically adjusts internal processing pipelines — including input dimensions for neural networks, feature selection layers, and normalization statistics — based on the detected schema. Models are trained using either a feature-masking strategy (zeroing out missing features during inference) or using fallback imputations (using mean / median values or synthetic extrapolations). Additionally, profile-based adapters are utilized to map features between the training schema and the inference-time schema. This approach ensures that the predictive and anomaly detection models remain fully operational even when new vehicles introduce sensor additions, removals, or replacements. It future-proofs the system against evolving automotive technologies without compromising performance or requiring disruptive retraining cycles.

[0101] The system architecture of the present invention is inherently modular and scalable, allowing for flexible deployment in both consumer and commercial settings. For consumers, the solution is implemented via an loT-based OBD-II hardware device that collects sensor data from thevehicle and streams it to the backend platform. For commercial users, such as fleet operators and original equipment manufacturers (OEMs), the same backend functionalities can be delivered entirely via APIs without the need for hardware integration.

[0102] The architecture enables selective activation of services such as predictive maintenance analytics, real-time anomaly detection, sensor failure prediction, and driver behavior analysis through secured API endpoints. This dual-mode offering — device-based for consumers and API-based for commercial clients — ensures wide applicability of the invention across different market segments without dependency on hardware presence. This hybrid approach is designed to maximize adoption flexibility while minimizing barriers to integration.

[0103] The system integrates an advanced Digital Twin architecture wherein a deep learning model (typically a multi-target regression neural network) is trained to reconstruct multiple sensor outputs simultaneously from a compressed latent representation. Unlike conventional singleoutput predictive models, the Digital Twin reconstructs the full set of critical vehicle parameters — such as engine temperature, fuel consumption, battery health, RPM, and torque — from past operational data. The reconstruction error (difference between predicted and actual sensor values) for each feature serves as a strong indicator of system health. Consistent or increasing reconstruction errors in specific features signify underlying performance degradation or emerging faults. Multi-target reconstruction also enables cross-feature consistency checks — for example, anomalous behavior in engine torque predictions can be cross-referenced against RPM and throttle position anomalies for triangulated diagnosis. This self-supervised anomaly detection mechanism enhances early warning capabilities, improves fault localization, and provides a richer understanding of vehicle degradation patterns without requiring manual rule-based diagnosis.

[0104] To enable seamless scalability across diverse vehicle types, the system incorporates dynamic input dimension adaptation capabilities. During model initialization and inference, the input layers and feature transformation blocks dynamically adapt based on the detected feature set size of the incoming vehicle data. This is achieved via several techniques like dynamic resizing of model input layers to match feature dimensions, zero-padding or feature masking when the live feature set is smaller than training set, selective feature imputation based on historical averages if essential features are missing and feature adaptation layers trained to map between feature sets across profiles. In effect, this ensures that whether the vehicle has 50, 100, or 500+sensor inputs, the system can flexibly accommodate the varying schema without requiring architecture changes or retraining from scratch. This dynamic adaptation ensures robustness across evolving sensor technologies, retrofitted vehicles, or datasets collected under different telematics standards, and eliminates hardcoded feature dependencies that would otherwise cripple cross-vehicle generalization.

[0105] A Variational AutoEncoder (VAE) model is implemented within the system to learn latent representations of normal vehicle behavior from multi-sensor data. The encoder projects input sensor data into a lower-dimensional probabilistic latent space, and the decoder reconstructs the original inputs. Reconstruction errors are computed and monitored continuously. Significant deviations in reconstruction loss beyond adaptive thresholds are used to flag anomalous behavior or early signs of potential failures. This VAE-based module supplements both anomaly detection and predictive analytics capabilities of the system, providing an unsupervised learning layer that enhances sensitivity to unknown or emerging failure patterns. The system employs a hybrid ensemble model that aggregates predictive outputs from multiple sources, including Transformer networks, Variational AutoEncoders (VAE), Graph Neural Networks (GNN), Digital Twin models, and traditional time-series models such as ARIMA, SVR, and LSTM. To ensure continuous and reliable operation even under partial failures or unavailability of individual models, a dynamic fallback mechanism is integrated. If any model experiences degraded performance, missing output, or becomes temporarily unavailable, the system automatically reroutes the ensemble weighting to compensate using alternative model predictions. This fallback design ensures uninterrupted real-time prediction, anomaly detection, and risk scoring.

[0106] During the anomaly risk estimation process using Graph Neural Networks (GNN) and ensemble scoring, a valid masking mechanism is employed to enhance reliability. Predictions with insufficient confidence or missing outputs (e.g., NaN values) are automatically filtered out based on a validity mask, ensuring that only robust and validated prediction samples are used for final health scoring, risk trend forecasting, and anomaly flagging. The valid mask approach improves overall system trustworthiness, minimizes false positives and false negatives, and ensures that incomplete or uncertain model outputs do not skew operational decision-making.

[0107] The system supports dynamic generation and maintenance of separate model profiles forconsumer vehicles and commercial fleets. This modular profiling allows the system to automatically adjust predictive thresholds, anomaly detection sensitivity, behavioral scoring parameters, and maintenance recommendation strategies based on the operational environment of the vehicle. Consumer profiles are optimized for individual driving patterns and maintenance needs, while fleet profiles are tuned for high-utilization, operational efficiency, and predictive maintenance at fleet scale. This flexibility enables the platform to serve diverse user segments without requiring fundamental architectural changes.

[0108] The system incorporates a hybrid, multi-model anomaly detection architecture designed to achieve superior accuracy, robustness, and early fault detection compared to standalone methods. Specifically, the anomaly detection pipeline combines outputs from multiple specialized models, including isolation forest (unsupervised anomaly detection based on data point isolation), variational autoencoder (VAE) latent reconstruction error scoring, transformer-based prediction deviation scoring, digital twin reconstruction error analysis and graph neural network (GNN) node-level risk scores. Each model contributes an independent anomaly likelihood or risk score. These scores are normalized and aggregated through an ensemble strategy — typically using a weighted or rule-based aggregation mechanism that emphasizes models with higher local confidence.

[0109] The hybrid ensemble approach provides several critical advantages like reduces reliance on any single model, minimizing false positives / negatives, captures complementary anomaly patterns (temporal, relational, statistical) and improves resilience against data noise or incomplete feature sets. This hybridized anomaly detection engine enables proactive identification of faults, driver behavior anomalies, sensor malfunctions, and unusual vehicle operating conditions far earlier than conventional threshold-based methods.

[0110] In addition to mechanical health monitoring, the system provides advanced driving behavior analysis to offer deeper insights into vehicle operation and driver habits. Key aspects analyzed include acceleration patterns (aggressive / abrupt acceleration indicating aggressive driving), braking behavior (frequency and intensity of harsh braking), steering stability (abrupt steering inputs, directional drift detection), idling time and engine load patterns and fuel efficiency trends linked to driving habits. These behavior profiles are extracted by continuously analyzing real-time telemetry data and applying statistical modeling, rolling window analysis, and ML-driven classification. Drivers or fleet managers are provided with categorized behavior insights(e.g., "aggressive", "efficient", "unstable") along with quantified impact on vehicle health, wear-and-tear, and maintenance needs. Behavior-based risk scoring is integrated with predictive maintenance models, enabling the system to provide more accurate forecasts based on both mechanical degradation and driving patterns, which is a key differentiator over traditional diagnostics solutions.

[0111] Beyond predictive analytics (forecasting future failures), the system includes a prescriptive analytics layer, enabling actionable recommendations based on risk assessment outputs. Once anomalies, degradations, or behavior issues are detected root cause estimation is performed based on deviation patterns across multiple sensors, risk severity levels are computed using ensemble anomaly scores and health score trends and recommended actions are dynamically generated, such as "Inspect braking system — consistent braking anomalies detected.", "Schedule battery check — voltage irregularities increasing.", "Change driving style — aggressive acceleration detected, leading to increased fuel usage.". Prescriptive actions are tailored for consumers [Simple actionable advice (e.g., "schedule a checkup", "improve driving style")] and fleet managers [Maintenance scheduling integration, operational risk prioritization]. This prescriptive layer adds decision intelligence to the system, moving beyond simply alerting to offering what should be done to improve or mitigate risks, thereby enhancing ROI for users.

[0112] The present invention leverages ensemble model outputs over a moving temporal window to forecast vehicle health risk trends and estimate Remaining Useful Life (RUL) in a days-based metric. The operational flow is as follows: Time-series predictions from the ensemble model are tracked continuously, a risk curve is generated by applying a sliding window average (e.g., 100-sample rolling mean), trend detection algorithms monitor whether risk is increasing, decreasing, or stable and based on risk thresholds crossing predefined danger levels, the system estimates the time-to-failure or remaining safe operational days. This enables extremely powerful early warnings such as "Risk increasing — possible major failure in ~15 days.", "Risk stable — no immediate action required.". Such time-to-failure forecasts are critical for scheduled predictive maintenance (fleet efficiency) and early vehicle checkup recommendations (consumer safety). It also supports proactive spare parts inventory management and service scheduling for commercial operators, delivering direct cost savings.In present invention discloses components of Vehicle Health Monitoring and Predictive Maintenance System, consisting of a data collection module (100) interfacing with OBD-II, GPS, and accelerometers to gather real-time vehicle data, a data transmission module (200) using lightweight MQTT protocols for secure and low-latency communication, a data processing module (300) for noise filtering, normalization, and feature engineering of collected data, a predictive analytics module (400) to forecast maintenance requirements and failure probabilities, an anomaly detection module (500) to identify irregularities in sensor data, a report generation module (600) providing user-friendly maintenance recommendations and insights.

[0113] Figure 1 illustrates the system architecture of a predictive vehicle maintenance system. It depicts the flow of data from various vehicle sensors (OBD-II, GPS, accelerometers) through a series of interconnected modules. Raw sensor data is initially collected by the Data Collection Module (100). This data is then securely transmitted to the Data Processing Module (300) via the Data Transmission Module (200), which employs the MQTT protocol for efficient communication. The Data Processing Module preprocesses the raw data to ensure its quality and extracts meaningful features. Subsequently, the Predictive Analytics Module (400) utilizes a combination of machine learning models (ARIMA, SVR, XGBoost) and an anomaly detection technique (Isolation Forest) to analyze the processed data and forecast future maintenance requirements. Finally, the Reporting Module generates comprehensive reports that provide actionable insights to vehicle owners, including maintenance predictions, driving behavior analysis, and recommendations for optimizing vehicle performance and minimizing costs. The invention consists of several interconnected modules that work together to collect, process, and analyze real-time data from various vehicle sensors, including OBD-II interfaces, GPS devices, and accelerometers. The primary goal is to provide actionable insights that predict potential vehicle failures and suggest proactive maintenance steps, thereby increasing the reliability of vehicles in various environments, from individual ownership to large-scale fleet management.

[0114] Figure 1 can also be described as the figure to illustrate the comprehensive architecture of the system, designed for real-time vehicle health monitoring, predictive maintenance, anomaly detection, and driver behavior analysis. The architecture is modular, scalable, and supports both consumer and fleet applications through a hybrid edge-cloud computing model.Sensor data is captured from multiple sources within the vehicle, including the OBD-II interface (engine metrics), GPS (location data), accelerometer and gyroscope (motion tracking), and optional environmental or contextual sensors. Data flows securely through authenticated channels using MQTT or HTTP protocols, ensuring minimal latency and high reliability. The incoming data is initially processed by an edge device that performs filtering, de-noising, and temporary buffering. Lightweight inference models run locally to support critical real-time tasks even when connectivity is limited. This layer ensures continuous monitoring and resilience in low-connectivity environments. The system supports integration with camera and microphone streams for advanced safety applications. Dashcam feeds are processed for scene recognition and event detection, while in-cabin microphones support audio anomaly classification. These modalities enhance behavioral monitoring and incident detection.

[0115] Once filtered, the raw data is passed to a transformation pipeline that performs data type harmonization, sensor embedding, and feature extraction. This stage generates structured feature matrices suitable for downstream machine learning pipelines and anomaly scoring mechanisms. The system employs a hybrid anomaly detection pipeline that includes Isolation Forests, Variational AutoEncoders (VAEs), and Graph Neural Network (GNN)-based risk scoring. Concurrently, it builds dynamic driver behavior profiles based on braking, acceleration, cornering, and idling patterns, using classification models to determine behavioral risk scores. These profiles directly influence predictive maintenance decisions. The prediction core integrates multiple models to maximize forecasting accuracy and fault sensitivity. These include ARIMA for modeling linear temporal patterns in sensor data, SVR for capturing complex nonlinear relationships, LSTM and Transformer models for sequential pattern learning and temporal dependencies, Digital Twin models for multi-sensor reconstruction and early fault triangulation and Graph Neural Networks (GNNs) especially Graph Attention Networks (GATs), to model interdependencies between sensor nodes and deliver node-level anomaly risk scores. All model outputs are unified through a weighted ensemble mechanism that adapts to individual model confidence, ensuring continuity even during partial model failures.

[0116] The architecture supports seamless transitions between edge-based and cloud-based processing. While low-latency tasks (e.g., emergency alerts) are handled on the edge, cloud resources process trend analysis, fleet-wide insights, and long-term model training. This designenables robust scalability for both single-vehicle and fleet-level deployments. Predicted risks and detected anomalies trigger an intelligent decision engine that generates real-time alerts and maintenance recommendations. This module calculates risk severity, identifies likely root causes, and prescribes actions tailored to the user context — simple suggestions for consumers and actionable diagnostics for fleet managers. Insights are presented through an intuitive mobile app and web dashboard. Users can view vehicle health scores, predicted failures, driver behavior summaries, and recommended actions. The interface supports real-time alerts and periodic reporting, with dashboards designed for both individual and fleet-level users. The system incorporates a closed-loop feedback process where maintenance outcomes and real-world sensor feedback are used to retrain and fine-tune models. Combined with dynamic input dimension adaptation, this allows the system to support vehicles with diverse sensor configurations without requiring rearchitecture or retraining.

[0117] Figure 2 provides a detailed data flow diagram illustrating the process of collecting, processing, and analyzing vehicle data for predictive maintenance. The diagram showcases how raw data is ingested from various sources, including the vehicle's On-Board Diagnostics (OBD-II) interface, environmental sensors, and GPS modules. This raw data is then stored and undergoes a series of preprocessing steps, including data cleaning and feature engineering. The processed data is subsequently used for anomaly detection, predictive modeling, and health analysis. The results of these analyses are then utilized to generate reports and insights for vehicle owners and fleet managers.

[0118] The data collection module (100) collects real-time data from vehicle sensors like the OBD- II interface for engine and performance metrics, GPS for location data, and accelerometers for driving behavior metrics such as braking and acceleration. The OBD-II system gathers critical vehicle data, such as engine performance, fuel efficiency, temperature, RPM (revolutions per minute), and more. GPS sensors capture location data, while accelerometers track the vehicle's movement, acceleration, and braking behavior. These data sources are collected in real-time to provide a comprehensive view of the vehicle's current operating state.

[0119] The system ensures seamless integration with these sensors to provide continuous monitoring of vehicle health. The data collected is processed to remove noise, standardize formats, and handle any missing or inconsistent values before further analysis.To facilitate real-time monitoring, the invention employs a data transmission module (200) transmits the collected data securely using MQTT protocols, with support for Bluetooth, GPRS, or Wi-Fi, ensuring low-latency and secure data flow. The MQTT protocol enables efficient, low-latency communication between the vehicle's sensors and the central server or cloud platform. This module ensures that sensor data is transmitted securely and in real-time, providing the necessary inputs for predictive analysis and anomaly detection without significant delay.

[0120] For increased flexibility, the system allows the use of different transmission mediums, including Bluetooth, GPRS, or Wi-Fi, depending on the specific requirements of the vehicle and the surrounding environment. Security measures, such as AES-256 encryption, are employed to protect the data during transmission, ensuring that sensitive information remains secure and private.

[0121] The data processing module (300) pre-processes the raw sensor data by removing noise, scaling values, and handling missing data to prepare it for analysis. It is responsible for preprocessing the raw data collected from the vehicle’s sensors. This module ensures that the data is cleaned, normalized, and transformed into a usable format for further analysis. The data processing module ensures secure data handling through AES 256 encryption and OAuth 2.0 authentication mechanisms for user access control. Key processing tasks include noise removal which filters out irrelevant data or sensor errors that may affect the accuracy of predictions, data scaling which standardizes values across different sensors (e.g., temperature, pressure, speed) to make them comparable and handling missing values which uses data imputation techniques to fill in missing or incomplete data points. Once processed, the cleaned data is passed on to the predictive analytics module (400) for in-depth analysis.

[0122] The predictive analytics module (400) is the heart of the system, responsible for making accurate predictions about the vehicle's health and potential maintenance needs. The predictive analytics module integrates ARIMA for capturing temporal trends in time-series data, SVR for refining residuals to handle non-linear relationships in the data and XGBoost for enhancing prediction accuracy through error correction. ARIMA helps to detect patterns that indicate normal wear or potential failure. SVR enables the system to detect more complex patterns and interactions between different vehicle components. XGBoost provides a more robust overall system.Together, these models improve the system’s ability to accurately predict potential maintenance issues, from minor component wear to more severe failures, by addressing both linear and non-linear trends in the data.

[0123] The present invention supports real-time anomaly detection and immediate risk alerting mechanisms, designed for both consumer and fleet applications. Sensor data streams are continuously monitored through lightweight inference modules. Each incoming data batch is processed through a series of predictive models (Transformer, VAE, GNN, Digital Twin, Traditional) and ensemble risk aggregation mechanisms. Anomalies are detected based on deviation thresholds (Z-score deviations, reconstruction errors, ensemble probability scores), validated against moving averages and statistical baselines to reduce false positives. When an anomaly exceeding configurable severity thresholds is detected, real-time alerts are generated. Alerts can trigger multiple actions like immediate in-app notifications to vehicle owners (consumer mode), cloud-based log events for fleet managers (commercial mode), triggered health report generation and prescriptive maintenance recommendations (if anomaly trend persists).

[0124] The real-time pipeline ensures that operational risks, potential failures, and health degradations are detected proactively, enabling users to take preventive action before critical failures occur. The anomaly detection module (500) employs isolation forest to detect anomalies. Anomalies, such as sudden spikes in RPM or coolant temperature, can indicate impending failures or irregular vehicle behavior that requires immediate attention. Traditional anomaly detection methods often struggle with high-dimensional, sparse sensor data, leading to delayed detection of faults. Isolation Forest efficiently isolates anomalies by identifying outliers in high-dimensional datasets. This technique excels in real-time applications, enabling the system to flag critical issues quickly. For example, if a vehicle experiences an unexpected rise in engine temperature, the system will immediately detect this anomaly and send an alert to the user or fleet manager. This allows for early intervention, minimizing the risk of severe breakdowns.

[0125] The predictive analytics module (400) incorporates feature engineering to calculate derived metrics, including braking intensity, acceleration variability, and comer handling, enhancing prediction accuracy. These metrics can include braking intensity that is how frequently and forcefully the driver applies the brakes, which could indicate potential issues with the brakingsystem, acceleration patterns that is how aggressively the vehicle accelerates, which could point to inefficient engine performance and location-specific data what is environmental factors such as terrain type or weather conditions, which can affect vehicle performance and maintenance needs.

[0126] Figure 2 can also be described as the figure that illustrates the comprehensive pipeline through which vehicle telemetry data is collected, processed, analyzed, and converted into actionable insights. It captures the interaction between various components — including external data sources, processing units, model inference engines, and visualization tools — working in unison to support real-time vehicle health monitoring and predictive maintenance.

[0127] The data pipeline begins with multiple external entities feeding raw telemetry into the system. These include the GPS module, which provides geolocation data; environmental sensors, which record ambient conditions; the vehicle's OBD-II interface, which supplies detailed engine and performance metrics; and the edge device, which functions as a localized loT unit collecting and buffering all incoming data. Additionally, the vehicle owner interacts with the system via mobile or dashboard interfaces, providing input or feedback for closed-loop learning. Once data is ingested, it is routed to various storage components, each dedicated to a specific data state. Raw data storage houses the unprocessed input; edge storage contains locally processed results and predictions; model storage holds deployed and updated machine learning models; processed data storage retains cleaned and feature-engineered datasets; and reports storage archives health diagnostics and usage summaries. These data stores allow seamless access and efficient retrieval across modules during inference, visualization, and reporting.

[0128] The system then moves into the data processing phase, which comprises several tightly coupled modules. The data collection and ingestion unit standardizes incoming formats, after which the data preprocessing module filters noise, scales values, and handles missing information. Feature engineering enriches the data by deriving secondary metrics such as acceleration variability or brake pressure trends. Simultaneously, a model training and update mechanism prepares learning datasets from this preprocessed information to refine existing models. An edge model training module generates lightweight versions of these models for constrained deployment environments, ensuring low-latency predictions at the vehicle level. Next, thecleaned and feature-rich data flows into the predictive modeling hub, which is composed of multiple specialized models. Graph Neural Networks (GNNs) capture relational dependencies between sensors, flagging node-level degradation. The Support Vector Regression (SVR) model detects nonlinear correlations in driving behavior and system health. LSTM networks identify long-range temporal patterns linked to performance decline, while ARIMA models forecast mechanical deterioration based on historical trends. These models operate in ensemble, producing robust and context-aware predictions, which are further optimized through a feedback-driven tuning process.

[0129] In parallel, the anomaly detection subsystem identifies unexpected patterns that may indicate failure or erratic sensor behavior. This subsystem includes an Isolation Forest model for statistical anomaly detection in high-dimensional datasets and a Transformer-based anomaly detector, which excels at recognizing irregularities in time-series sequences. Together, they provide highly sensitive and early alerts for mechanical or behavioral abnormalities. The outputs of these predictive and anomaly detection models are integrated into downstream modules for user delivery. The report generation engine compiles maintenance forecasts, health scores, and driver behavior insights into readable documents. Advanced dashboard visualization tools present live data and historical trends through intuitive charts and widgets. The real-time insight generator ensures that incoming data triggers updates to the dashboards without delay, offering near-instantaneous visibility into vehicle health. The health analysis module consolidates prediction and anomaly signals, computing composite risk scores and determining recommended interventions. A robust feedback loop closes the system by integrating user input and field observations. Feedback on prediction accuracy, anomaly alerts, and health reports is fed back into the data processing and model training layers. This loop enables continuous learning, ensuring the system remains adaptable to changes in driving behavior, vehicle conditions, and evolving data patterns.

[0130] Figure 3 details the Anomaly Detection Pipeline, outlining the sequence of steps involved in identifying and handling abnormal data within the predictive maintenance system. The process begins with the collection of raw data from various vehicle sensors. Subsequently, this data undergoes a preprocessing stage where noise is removed, missing values are handled, and data scaling is performed. The preprocessed data is then fed into an Isolation Forest model, a machine learning algorithm specifically designed to detect anomalies by isolating them in the data space. Detected anomalies are flagged, alerting the system to potential issues such as sensormalfunctions or unusual driving behavior. These flagged anomalies are then further analyzed and used to refine the predictive models, improving the overall accuracy of the system.

[0131] By extracting these tailored features, the system provides personalized insights into each vehicle's condition based on its unique usage patterns. This allows for more accurate predictions and better-informed maintenance recommendations, tailored to individual vehicles or specific driving conditions.

[0132] The reporting module generates comprehensive, user-friendly reports that summarize vehicle health scores, predicted maintenance needs, cost-saving recommendations, and emission reduction strategies. These reports are designed to be easily understood by both technical and non-technical users, ensuring that fleet managers, insurance companies, rental agencies, and individual vehicle owners can make informed decisions regarding vehicle maintenance. Reports are generated in a PDF format and include key information such as: current vehicle health status, predicted maintenance needs (e.g., engine repairs, brake replacements), estimated timelines for upcoming maintenance and recommendations based on driving patterns or environmental conditions

[0133] Figure 4 illustrates the hybrid predictive modeling system employed in the vehicle health monitoring system. The system operates by combining the strengths of multiple machine learning models to achieve accurate and robust predictions of vehicle health. Raw sensor data is initially collected and preprocessed, which includes steps such as data cleaning, feature engineering, and standardization. This processed data is then fed into the model pipeline, where it undergoes a series of analyses. Time series analysis is performed to identify trends and patterns in the data, and initial predictions are generated. The residuals, or errors, from these initial predictions are then used to refine the model through a process of error correction and feature ranking. This iterative process continues until the final predictions are generated. The system also incorporates mechanisms for residual update and pattern learning to further enhance the accuracy and robustness of the predictions.

[0134] Figure 4 can also be described as the figure that represents the internal architecture of the system’s machine learning pipeline, showcasing how diverse models and data processing techniques are orchestrated to generate robust predictions for vehicle health monitoring and predictive maintenance. The diagram flows from input to preprocessing, through multipleparallel model pathways, to an integrated hybrid prediction engine and output layer, enabling high-accuracy insights.

[0135] The process begins with the input layer, which aggregates a wide variety of data sources. These include feature data derived from preprocessed streams, processed data from storage, quality metrics, raw sensor data, edge-processed data, and historical data collected over time. These input streams form the foundation for modeling efforts, contributing temporal, contextual, and quality-based perspectives on vehicle performance. The data is then funneled through the data preprocessing block, which performs several essential preparation tasks. First, data cleaning removes corrupt, missing, or irrelevant data points. This is followed by time-series processing, which organizes the data into chronologically structured sequences. Feature standardization ensures that variables are normalized and scaled appropriately for model training, and data splitting divides the dataset into training, validation, and test subsets to ensure generalization and performance benchmarking. Post-preprocessing, the data is routed to multiple parallel modeling pipelines, each tailored to address different aspects of vehicle health prediction. The VAE model pipeline starts with latent feature learning, allowing the system to compress complex input data into a lower-dimensional representation. This supports VAE anomaly detection and predictive reconstruction, enabling unsupervised learning of typical operational patterns and early identification of deviations.

[0136] The Isolation Forest module specializes in outlier detection and generates anomaly scores, effectively identifying statistical anomalies in the dataset. Meanwhile, the Graph Neural Networks (GNNs) pipeline focuses on relationship modeling across sensor data. This allows the system to capture interdependencies and generate failure pattern detection and advanced predictions at the node level, which is particularly useful for localized sensor degradation. The LSTM (Long Short-Term Memory) pipeline handles sequential pattern learning by analyzing temporal patterns in the input data. This model outputs temporal forecasts and final predictions, making it effective in capturing long-term trends. Similarly, the ARIMA model executes time analysis to generate predictions and residuals based on linear dependencies and historical trends. The Transformer model is included for its capability in temporal dependency learning across longer sequences. It also contributes to anomaly detection and predictive enhancements, adding precision to the ensemble, especially under complex operating scenarios.

[0137] Each of these model pathways feeds into the Hybrid Ensemble Engine, which performs modeloutput combination using confidence-based or rule-based strategies. The result is a final hybrid prediction that combines the strengths of all contributing models to offer a more accurate and resilient forecast of vehicle health status. In parallel, a Digital Twin Interface complements the hybrid engine by simulating real-time sensor behavior. This module handles real-time vehicle health modeling, behavior prediction, and anomaly scoring integration. It also supports pattern learning, residual update, and prediction refinement, closing the loop with a self-supervised reconstruction-based assessment of health degradation. The output layer of the system consolidates all predictions into combined results, along with calculated performance metrics, and generates recommendations and real-time insights. These outputs are delivered to userfacing components, such as dashboards or mobile applications. Additionally, this layer supports health update feedback, behavior update predictions, and health model accuracy tuning, ensuring the system continues to evolve based on operational data and user feedback.

[0138] Figure 5 depicts the architecture of the system comprising various interconnected components. Raw data is collected from multiple sources, such as devices and user interactions, and stored in multiple databases, including user, device, and engine health databases. A central pipeline processes this raw data, performing tasks like data cleaning, transformation, and enrichment. The processed data is then utilized for various analytical purposes, including predictive analysis, anomaly detection, and location-based services.

[0139] It illustrates the integrated data and service infrastructure that supports core functionalities such as engine health monitoring, predictive analytics, user management, and real-time streaming for connected automotive systems. This architecture leverages distributed components, secure communication layers, and scalable pipelines to enable low-latency, resilient, and extensible operations. At the heart of the system is the Engine Health module, which interfaces with a local database and responds to GET requests routed through an authentication system and web application firewall (WAF). The user accesses this system via a secure Auth service, which verifies credentials and enforces access policies before directing requests through a load balancer. The engine health module retrieves or stores vehicle performance data in the local and replicated engine health database clusters, comprising master-slave replication for fault tolerance and load distribution. Adjacent to this, a centralized User Database links to both a Hardware Database and a Subscription Database, tracking physical device associations and subscription details respectively. The user database serves multiple purposes, including authorization, user profile tracking, and association with real-time analytics.Engine health data and user inputs are streamed to a predictive analytics pipeline running in the cloud, where long-term performance trends are modeled and used for proactive diagnostics. This data pipeline processes vehicle signals in batch or stream form and feeds them back to cloud services or to real-time dashboards for user and administrator access. The system interacts with a distributed device cluster, which includes multiple edge-connected devices deployed in vehicles or field locations. These devices serve as data originators for engine telemetry and are actively monitored for events like accident detection, which automatically triggers alerts and downstream analytics. Data from devices is also forwarded to a Kafka-based message queue, which acts as the backbone for data replication, retention policy enforcement, streaming analytics, and high-throughput fraud detection or map generation tasks. Kafka streams are fed into processing pipelines that distribute data across nodes and services, enhancing the system’s ability to scale horizontally and process large volumes in near realtime.

[0140] A cluster of nodes (Node 1 through Node 4) forms the core of the system's distributed computing environment. These nodes exchange information and host various analytics, storage, or decision-making services in a replicated, resilient architecture. Data collected and aggregated here is further analyzed or stored based on retention and usage policies. The system also integrates with location services, allowing it to track devices geographically and trigger contextual responses. This is particularly useful for emergency contact coordination, where real-time location data supports automated dispatch, alert generation, or user notification based on detected incidents. In summary, this architecture illustrates a modular, distributed, and fault-tolerant system design for real-time engine health analysis and predictive maintenance. It combines secure authentication, scalable message queues, multi-node processing, and cloudbased analytics pipelines into a cohesive platform for intelligent automotive operations.

[0141] Figure 6 presents the comprehensive end-to-end process through which the system identifies irregularities in vehicle behavior, driver patterns, and sensor data. It depicts a modular and feedback-integrated architecture, starting from data acquisition to anomaly scoring and decision delivery to the user, with multiple inference and refinement loops enhancing accuracy.

[0142] The process begins with the input sources feeding real-time telemetry into the system. These sources include the GPS module, accelerometer, environmental sensors, and edge devices. Thedata collector module aggregates and synchronizes these incoming data streams, which are then passed to the edge device for preliminary processing and local anomaly checks. This layer ensures that low-latency decisions can be made even before full cloud-side inference. Once collected, the raw data flows into the preprocessing and prediction modeling engine, where core time-series forecasting techniques are applied. The system runs parallel applications of three statistical and machine learning models: ARIMA, LSTM, and SVR. Each of these models contributes to predicting expected sensor behavior, identifying potential deviations, and generating forecast-based residuals. These residuals are then compared against real sensor data to isolate anomalies.

[0143] To further enrich prediction accuracy, the system performs contextual signal enrichment. This subsystem executes several derived-feature extraction tasks, including trip marker extraction, location identification, vehicle speed analysis, and dynamic adjustment modeling. The results feed into both real-time edge inference and later-stage anomaly detection layers. Additionally, fallback imputation handles missing or inconsistent values, ensuring continuity in downstream analysis. The core anomaly detection engine receives predicted values, residuals, and enriched features. This engine consists of several submodules working in tandem. The pipeline begins with edge-based diagnostics for lightweight local checks. This is followed by traditional machine learning models, such as the Isolation Forest, which flags statistical outliers, and VAE-based anomaly detection, which identifies deviations from learned latent patterns. Further processing is done using threshold scoring based on reconstruction errors, and finally, Graph Neural Networks (GNNs) evaluate inter-sensor dependencies to detect complex structural anomalies across the data graph.

[0144] Following anomaly scoring, the pipeline performs risk analysis to assess the severity and category of detected anomalies. This includes separate pathways for analyzing vehicle component anomalies and driver behavior anomalies, ensuring precise categorization and appropriate follow-up actions. The results are compiled into a final anomaly report that includes health impact assessments and possible behavioral correlations. These insights flow into the user delivery interface, where the system provides vehicle health scores, live anomaly alerts, and recommendation messages. These include actionable suggestions for maintenance or driving behavior modification. The system also supports real-time updates and predictive feedback, which are factored into the maintenance recommendation engine, allowing prescriptive outputs to be continuously refined. Finally, a feedback loop captures userresponses and actual vehicle outcomes to improve model reliability. This feedback is routed back into both the prediction models and anomaly detection modules, facilitating adaptive threshold updates and continuous learning. Additionally, updates are issued to improve the health model accuracy and adjust the update prediction mechanisms, enabling personalized, context-aware monitoring.

[0145] Figure 7 outlines the design of the API-driven backend that powers the platform. It illustrates how multiple APIs, services, and databases interact to deliver secure, scalable, and modular support for mobile and web applications, internal microservices, and external fleet or OEM platforms. This architecture ensures seamless integration of health monitoring, anomaly detection, billing, reporting, and user management functionalities.

[0146] External interfaces such as the Mobile Application, Web Application, Fleet Management Systems, OEM Partner Platforms, and internal components like the Monitoring Service and Internal Microservices initiate various API requests, including prediction queries, anomaly alerts, health data lookups, report generation, and authentication. These requests are funneled through a centralized API Gateway, which acts as the primary traffic controller and security checkpoint. Within the API Gateway, the architecture is divided into a collection of modular services, each accessible via a specific RESTful API. The Report Generation API interfaces with the Report Service, which processes requests and stores output in the Report Storage system for future retrieval. The Admin & Configuration API allows access to administrative features, which are executed through the Admin Configuration Service.

[0147] The Anomaly Detection API links directly with the Anomaly Analysis Service, which evaluates incoming data for abnormal behavior and stores the results in the Anomaly Database. For vehicle-specific diagnostics, the Vehicle Health API interacts with the Vehicle Data Service, which in turn communicates with the Vehicle Database to read or write diagnostic metrics. All data traffic is protected via End-to-End Encryption, ensuring secure communication between clients and backend systems. Monetization and plan management are handled through the Subscription & Billing API, which interfaces with the Billing & Subscription Service and its corresponding Billing Database. This subsystem manages billing history, user subscriptions, and service tiers.

[0148] For predictive capabilities, the Predictive Analytics API sends data to the Predictive ModelingService, which generates outcomes and updates the Prediction Database. These insights are consumed by other services such as dashboards or notification engines. The Notification API manages real-time user communication by routing instructions to the Notification Management Service, which delivers alerts through the integrated Notification System. To maintain secure, role-based access, the Authentication API facilitates login, session validation, and access permissions via the Auth Service and Role-Based Access Control (RBAC) system. User credentials are stored securely in the User Database. Fleet and user management are also embedded into the architecture. The Fleet Management API provides access to the Fleet Analytics Service, which generates insights and stores metrics in the Fleet Database. Simultaneously, the User Profile API integrates with the User Management Service to allow updates to user data, syncing changes with the User Database. In summary, the API architecture ensures a secure, efficient, and modular approach to vehicle health monitoring. It supports extensive scalability across consumer and enterprise applications through clearly defined APIs and services, all orchestrated under a unified gateway.

[0149] Figure 8 outlines the end-to-end flow of machine learning model development, deployment, monitoring, and retraining within the platform. This lifecycle ensures continuous improvement of predictive accuracy, resilience to data drift, and adaptability to evolving vehicle profiles and user environments. The process begins with data sources, including sensor data streams, historical vehicle data, and edge device data. These are processed through the data ingestion and cleaning pipeline, where raw telemetry is standardized, erroneous records are filtered, and missing values are flagged. Once cleaned, the data passes through a series of transformation layers such as feature engineering and adaptation, dynamic schema handling, and fallback imputation, ensuring compatibility across diverse vehicle configurations and sensor sets.

[0150] Following preprocessing, the pipeline initiates model training. The Model Training Engine comprises multiple modeling techniques tailored to capture various patterns and relationships within vehicle data. It includes training of Transformer models for complex sequential dependencies, VAE models for latent anomaly detection, GNNs for inter-sensor relationship modeling, LSTM networks for time-series forecasting, Isolation Forests for statistical outlier detection, ARIMA models for linear time-series trend analysis, and SVR models for nonlinear regressions. Each model is trained using the most up-to-date data streams and user feedback. Once models are trained, they undergo validation in the Model Validation Engine. This layer performs cross-validation, model metric evaluation, and critical checks like bias and varianceanalysis and drift detection. These steps ensure that only high-quality models are promoted to deployment.

[0151] The deployment process is managed by the Model Deployment Manager, which facilitates seamless distribution of inference models to two primary environments: Cloud Prediction Service for centralized computing and Edge Node Prediction Service for low-latency, local predictions. These services handle real-time requests and continuously monitor performance metrics. Simultaneously, the Model Monitoring and Health Check module operates to oversee deployed models. It includes real-time performance tracking, user feedback integration, drift monitoring, and anomaly detection feedback analysis. This component ensures the system remains robust and responsive, triggering retraining events whenever performance degradation, data drift, or new vehicle models are detected.

[0152] Triggered retraining events are handled by the Retraining Trigger System, which analyzes causes such as performance degradation, data drift, or the inclusion of new vehicle models. It also integrates corrections from user feedback, allowing the system to learn from real-world performance gaps. This initiates a new cycle of model training, validation, and deployment — maintaining continuous learning and improvement across the lifecycle. Finally, real-world correction data from user interactions is cycled back into the data ingestion pipeline, reinforcing the feedback loop and ensuring the next iteration of models reflects actual operating conditions. This complete lifecycle — from ingestion to retraining — enables the platform to maintain high predictive precision and adaptive learning capabilities across changing fleet and user dynamics.

[0153] Figure 9 presents a detailed overview of the data flow, processing, analysis, and presentation layers that support the platform's dynamic dashboard system. This architecture enables seamless integration of raw telemetry data, predictive insights, and user feedback into an interactive, visually-driven decision support interface for fleet managers and vehicle owners. The architecture begins at the data acquisition layer, where real-time input is captured from multiple sensor sources, including the OBD-II interface, environmental sensors, cloud storage, accelerometer, GPS module, and edge device. These streams contain vehicle telemetry, environmental context, motion data, and geolocation coordinates. All incoming data is transmitted securely and routed through the PMIC (Power Management IC), which facilitates stable energy management, ensuring sensor uptime during live data capture. Once acquired,the data flows into the data ingestion and pre-processing layer, where a data ingestion module collects and standardizes the sensor streams. The data is then passed through the real-time data cleaning module to remove noise and erroneous values. Cleaned data is funneled into the feature engineering module, which extracts relevant attributes required for predictive and anomaly detection tasks. Processed features are routed simultaneously into downstream analytics and visualization components.

[0154] Anomaly detection occurs in a dedicated hybrid anomaly detection engine, where the system combines insights from multiple models to improve fault detection accuracy. This engine includes real-time vehicle diagnostics, edge-based anomaly detectors, and unsupervised models like Isolation Forest. Output from this engine is classified and then sent for evaluation through clustering models such as K-Means, and regressors including Support Vector Regression (SVR) and ARIMA, forming the core of the system's ensemble-based prediction layer. Predictive analytics modules, located downstream, consist of maintenance prediction modules that estimate component wear or failure, fault classification modules, and componentbased scoring engines. These work in conjunction with the anomaly detection engine to deliver actionable vehicle health forecasts and failure risk estimations. Their results are routed to the dashboard in real time. These insights populate various dashboard interface layers, including the Advanced Dashboard Renderer, which powers a user-friendly UI, and the Communication Layer, responsible for serving reports and alerts. The fleet management engine provides data to modules such as engine diagnostics and trip statistics analysis, while simultaneously supporting route-based risk assessments, driver behavior modeling, and vehicle-specific health ranking through the Al-powered behavioral risk module.

[0155] All visualizations and recommendations are fed into the dashboard UI, which is split into a visual analytics panel and a user interaction module. This module connects with user devices and supports features like custom alert setup, role-based access control (RB AC), and feedback loop management. User actions, such as alert acknowledgment or feedback, are captured and passed back into the data pipeline for model refinement. The system also incorporates realtime geo-visualization modules that track vehicle locations using edge trackers and GPS data, and present route-based risk overlays to enable predictive insights while the vehicle is in motion. These spatial data visualizations support decision-making for fleet routing and anomaly zone detection. Finally, predictive and anomaly insights, along with historical trends, are looped back into the model calibration system, ensuring that future predictions remainaccurate and context-aware. Alerts are pushed to users via configurable channels based on severity and risk scores, closing the loop in an intelligent, user-responsive dashboard experience.

[0156] For applications requiring low-latency processing, the invention supports edge computing, which shifts some of the data processing and anomaly detection tasks to local devices within the vehicle. By doing so, the system can provide real-time alerts and critical failure predictions without needing to rely on constant cloud connectivity. Edge computing also helps to optimize bandwidth usage by processing and filtering data locally, sending only relevant insights to the cloud for long-term storage or further analysis.

[0157] In environments where network connectivity is limited (e.g., remote areas), the edge computing capabilities ensure that the system continues to function effectively. This is particularly useful for fleet management and autonomous vehicle applications, where constant, uninterrupted operation is critical.

[0158] The system is designed for cloud-based deployment, which allows it to scale from a single vehicle to large fleets. The cloud infrastructure provides flexibility and ensures that predictive maintenance insights are available in real-time, regardless of the number of vehicles being monitored. Additionally, the system can be extended to incorporate edge computing for latency-sensitive applications, further enhancing its scalability and responsiveness.

[0159] Figure 10 showcases the vehicle-mounted telemetry device (202) comprising microcontroller / microprocessor (302), OBD-II interface (304) coupled to vehicle communication bus (306) acquiring engine parameters (308) and diagnostic parameters, positioning receiver (310), inertial sensor (312), wireless communication module (314), and a processing unit external to the vehicle (204), connected via a mobile data network.

[0160] The method underlying this invention for vehicle diagnostics and predictive maintenance combines ARIMA, SVR, and XGBoost models to provide high-accuracy failure predictions. The integration of loT plays a pivotal role by enabling real-time data collection from the OBD-II port, GNSS, and motion sensors. This holistic approach ensures comprehensive monitoring of the vehicle's condition. The device is designed with a compact, scalable architecture, ensuring it remains compatible with future vehicles and new protocols like CAN-FD. The modular hardware design allows for easy upgrades and expansions as new technologies emerge.

[0161] Although the current implementation primarily operates in cloud-connected or local batch modes, the system architecture has been designed with future readiness for real-time edge processing and streaming data support. Key provisions included are lightweight, edge-deployable model architectures (Transformer, VAE, GNN optimized), modular data ingestion pipelines capable of handling both streaming MQTT / Telematics and static CSV input, planned integration with loT-capable OBD-II devices with ML acceleration support (e.g., NXP i.MX processors) and fall-back to real-time mini -batch inference in case of network disruptions. By future-proofing the architecture, the system is positioned to transition seamlessly into full realtime, on-vehicle inference deployments as hardware capabilities mature. This ensures scalability across markets — from consumer vehicles using simple OBD-II dongles, to fleet and OEM deployments with advanced edge processing requirements.

[0162] Although the initial deployments rely primarily on cloud-based or server-based inference, the system architecture has been intentionally designed to be hardware-agnostic and edge-ready. Specifically lightweight variants of the Transformer, VAE, GNN, and Digital Twin models are being optimized for deployment on loT edge devices (e.g., embedded platforms with ARM cores, microcontrollers, or ML accelerators), planned hardware compatibility includes OBD-II devices with on-board computing capabilities (e.g., NXP i.MX 8M Plus, NVIDIA Jetson Nano, Qualcomm Snapdragon Automotive Solutions) and real-time anomaly detection, basic sensor failure prediction, and low-latency health monitoring can be performed locally without relying on constant cloud connectivity. Edge support ensures minimal latency for time-critical alerts, reduced dependency on external networks (especially in fleet vehicles operating in remote areas) and lower bandwidth consumption. This future capability allows the system to seamlessly scale into both consumer and fleet applications across diverse connectivity environments.

[0163] While the system currently operates using batch-processed telemetry data (CSV uploads or processed data streams), its ingestion pipeline and model orchestration are structured to support real-time streaming inputs. Architectural provisions already included MQTT / HTTP based ingestion modules ready for vehicle telematics streaming, online mini-batch preprocessing, including real-time normalization, feature extraction and anomaly scoring and windowed risktrend forecasting updated in near real-time as new data batches arrive. This enables future deployment where vehicles transmit data in small, incremental bursts (e.g., every minute or 5 minutes) rather than waiting for complete trip uploads. Real-time support enables on-the-fly risk scoring and notifications, dynamic driver behavior profiling and immediate anomaly and health degradation alerts. Positioning the system not just for reactive analytics but continuous vehicle monitoring and intervention.

[0164] To further enhance model training, evaluation, and performance robustness, the architecture is extensible to integrate Generative Adversarial Networks (GANs) for synthetic data generation in future releases. Planned capabilities include synthetic sensor data generation for rare failure modes (e.g., engine misfire under specific altitude and load conditions), augmentation of small or imbalanced datasets to prevent model overfitting and scenario simulation (e.g., sudden brake failure, extreme temperature degradation) to test model resilience. Using GANs will improve anomaly detection sensitivity for rare events and enable faster onboarding for new vehicle profiles even with limited real-world data. Although not part of the initial MVP, GAN-based augmentation modules are earmarked for future expansion once sufficient baseline operational data has been accumulated.

[0165] The system architecture is designed to support real-time inference and anomaly detection on embedded or loT edge devices through lightweight models. The data ingestion and processing pipelines are capable of real-time streaming analytics using MQTT / HTTP protocols for live vehicle monitoring. Optional modules for synthetic data generation using GANs may be integrated to enhance rare event prediction capabilities.

[0166] The vehicle health monitoring and predictive maintenance system presented in this invention is a solution for improving vehicle performance, reducing maintenance costs, and preventing unexpected breakdowns. By leveraging a combination of machine learning, loT technologies, and real-time data processing, the system provides accurate, actionable insights into vehicle health, benefiting a wide range of industries, from fleet management to autonomous vehicles.

[0167] Signal model, for each telemetry channel xA(k)_t sampled at frequency f_k, over windows of length W with stride S:

[0168] 1) Base forecast (per-channel): xA(k)_t = ARIMA_k(xA(k)_{t-W:t-l }).2) Residuals: sA(k) t = xA(k)_t - xA(k)_t.

[0169] 3) Residual learner (cross-channel): r_t = g([sA{(l..K)}{t-W:t}], context_t), where g is SVR / XGBoost (or Transformer) over residual sequences + context.

[0170] 4) Anomaly gate: Isolation-Forest score s_IF or VAE reconstruction loss s_VAE and / or GNN node-risk vector pooled to summary statistics.

[0171] 5) Calibration: Convert scores to calibrated probability p via isotonic or Platt scaling on held-out validation data.

[0172] 6) Maintenance index & policy: M t = a ||r_t|| + P p + ytrend(p _{t~A:t}) - S cooldown(t). Trigger part-specific recommendation when M t > r part for >N windows, with daily false-positive caps and debounce.

[0173] Mask-gated residual normalization (robust to missing sensors): r* = (m O r) / max(l, S m i), where m is a per-feature validity mask computed from coverage and plausibility rules.

[0174] Threshold selection to a target false-positive rate a: choose T to satisfy FPR val(r) ~ a (e.g., T is the (l~a) quantile of calibrated probabilities on negative validation samples).

[0175] Pseudocode:

[0176] for each window X_t:

[0177] m = validity _mask(X_t)

[0178] compute r_T, r_V, r_tw (optional), s gnn (optional)

[0179] z = features_from_mask_gated_residuals(r_T, r_V, r_tw, s gnn, m)

[0180] S' = wAT z + b or S' = f_theta(z)

[0181] p = sigmoid(S');

[0182] p hat = calibrator(p)

[0183] if data_insufficient(N<Nmin or sum(m) < k or missing model):

[0184] p hat = calibrator(scale(IF_score(X_t)))

[0185] if p hat >= tau: raise anomaly

[0186] In Eq. l-Eq.9, Xt denotes a vector of sensor readings at time t, x<t-i t-H denotes a future horizon of length H,Ax and ~x denote outputs of the temporal forecaster and autoencoder respectively, x{tw}denotes digital-twin predictions, rr, rv and r{tw} denote normalized residual norms from the temporal 35 forecaster, variational autoencoder and twin respectively, m is a per-featurevalidity mask with entries m_i E {0,1}, S(!snn!are per-node graph-neural -network scores,

[0187]

[0188] are their mean and maximum, F is the number ot teatures, miss rate is the fraction of unavailable features, z is the concatenated feature vector, w and b parameterize a linear fusion function, f{0} is a learned nonlinear fusion function, o is a logistic activation, g is a calibration mapping obtained on validation data,Ap is a calibrated anomaly probability, T is a decision threshold, and a is a target false-positive rate. FPR{vai}( ) denotes the false-positive rate measured on validation data at threshold r.

[0189] <

[0190]

[0191] & (Eq.8)>

[0192]

[0193] EXAMPLES

[0194] Example A — FPR control. On a held-out validation set with target a = 1%, the selected threshold yields FPR(val) ~ 1%; uncalibrated baselines exceed the target.

[0195] Example B — Missing sensors. With random masks up to k = 4 missing features, mask-gated fusion maintains true-positive rate at fixed FPR better than non -gated fusion. For instance, with K = 10 features and random masks that zero out up to 4 features, the mask-gated residual r* maintains a mean true positive rate within a small margin of the full-feature case at fixed falsepositive rate, whereas a non-gated design that divides by K instead of S m i exhibits a substantially larger degradation in true-positive rate under the same conditions.

[0196] Example C — Edge bandwidth (optional). With redge = T - 0.2, uplink volume reduces >30% while maintaining anomaly recall within 2% of full-uplink processing.

[0197] The system utilizes a high-performance processing unit, such as, but not limited to, the STM32F407VGT6 microcontroller (based on the ARM Cortex-M4 core) or an equivalent 32-bit RISC processor. The processing unit boasts a processing speed of approximately 166-170 MHz, 0.5-2 MB of flash memory, and 190-194 KB of RAM, along with a Floating Point Unit (FPU). This high-performance processor is merely one example of a component suitable for loT applications and real-time analytics; other microcontrollers or microprocessors with similar computational capabilities may be employed.

[0198] For wireless communication and location tracking, the system employs a cellular transceiver module, for example, the Quectel BG96 module or a similar Low-Power Wide-Area Network (LPWAN) modem. The selected module preferably supports standards such as LTE Cat-Ml / NBl for cellular connectivity and includes a built-in GNSS receiver (supporting GPS, GLONASS, Galileo, or BeiDou) with high positioning accuracy (e.g., 2-3 meters).

[0199] Vehicle telemetry is acquired via a vehicle bus interface component, such as the MCP2562FD CAN transceiver from Microchip or an equivalent automotive-grade interface IC. This componentprovides support for modem communication protocols, including CAN-FD and ISO 15765-4. Preferably, the chosen component is automotive-grade (e.g., capable of operating in temperatures from -40°C to +125°C) and integrates transceiver functionality to optimize PCB space and cost.

[0200] For motion tracking, the system integrates an inertial measurement unit (IMU), such as the BMI088 6-axis sensor by Bosch or a similar MEMS-based accelerometer and gyroscope. The sensor is preferably automotive-certified and capable of measuring high-g accelerations (e.g., ±16g to ±26g) to ensure high vibration stability and precise tracking of vehicle dynamics, including braking, cornering, and impact detection.

[0201] The power supply is managed by a wide-input voltage step-down (buck) converter, such as the LMR14030-Q1 or the TPS54332 from Texas Instruments. The regulator is selected to withstand input voltages exceeding typical vehicle transients (e.g., up to 40V or higher) while delivering a stable output (e.g., 2-4A) with high efficiency (90-94%). This ensures robust power management and reduces the thermal footprint compared to traditional linear regulators.

[0202] It should be understood that the specific hardware components, model numbers, and manufacturers described herein (e.g., STM32, Quectel, Bosch, Texas Instruments) are provided solely as illustrative examples of a preferred embodiment. These specifics are not intended to limit the scope of the invention. Those skilled in the art will recognize that equivalent components performing substantially the same function — such as alternative microcontrollers, cellular modems, or power management ICs — may be substituted without departing from the spirit and scope of the present disclosure.

[0203] Figure 11 illustrates an exemplary circuit-level block diagram (detailed schematic view) of a vehicle-mounted OBD-II telemetry device showing major electrical subsystems and their interconnections. In the illustrated embodiment, the device is configured to be removably coupled to a vehicle OBD-II port (e.g., JI 962) and to receive vehicle supply (VBAT) from the vehicle. The supply is conditioned by a power management subsystem that may include protection circuitry and one or more voltage conversion stages (e.g., a buck conversion stage and one or more regulation stages) to provide stable supply rails for logic and communication subsystems, and may optionally include an energy-storage element (e.g., backup battery) and charger circuitry for continuity of operation, and an LDO / regulation stage to generate logic rails. A processing unit (e.g., a microcontroller / microprocessor) is communicatively coupled to (i) a wireless communication subsystem including a cellular transceiver and / or positioningsubsystem (GNSS), coupled to a SIM interface and antennas, and (ii) a vehicle bus interface / transceiver subsystem coupled to the OBD connector for acquiring vehicle engine and diagnostic parameters through one or more supported vehicle protocols. The interconnections shown are exemplary; equivalent circuit topologies, equivalent components, and equivalent communication interfaces may be used without departing from the scope of the disclosure.

[0204] Figure 12 illustrates an exemplary functional architecture of the vehicle-mounted telemetry device, grouped into functional subsystems including an external interface (OBD connector, antennas, SIM and optional debug interface), an automotive-grade power management block configured to condition vehicle input and generate regulated rails, a main processing block including a processing unit, a wireless communication block including a cellular transceiver and / or positioning receiver (with optional level shifting between logic domains), a vehicle bus transceiver / interface block configured to couple the device to the vehicle communication bus via the OBD connector, and optional sensors and local storage (e.g., inertial sensing and nonvolatile memory). In operation, the processing unit coordinates acquisition of vehicle parameters via the vehicle bus interface, acquisition of location and motion signals via positioning and inertial subsystems (where present), time-alignment / formatting of telemetry, and transmission of telemetry through the wireless communication subsystem to an external processing unit. The architecture is illustrative and non-limiting; equivalent functional blocks may be substituted without departing from the scope of the invention.

[0205] The system of the present invention operates with standard OBD-II and commodity IMU / GNSS, deployable on embedded gateways or vehicle servers for fleets, warranty analytics, and preventive maintenance.

Claims

Claims:

1. A system (200) for real-time vehicle health monitoring and predictive maintenance, comprising:i. a vehicle-mounted telemetry device (202) comprising:a. a microcontroller or microprocessor (302);b. an OBD-II interface (304) coupled to a vehicle communication bus (306) and configured to acquire engine (308) and diagnostic parameters;c. a positioning receiver (310) configured to determine geographic location; d. an inertial sensor (312) configured to measure linear acceleration and angular velocity of the vehicle; ande. a wireless communication module (314) configured to connect to a mobile data network; andii. a processing unit (204) located external to the vehicle;wherein the vehicle-mounted telemetry device is a physical electronic apparatus configured to be removably coupled to an on-board diagnostics (OBD-II) port of a vehicle and to directly receive, during operation of the vehicle, electrical signals representing engine operating parameters and diagnostic status from a vehicle communication bus,wherein the microcontroller or microprocessor is configured to synchronize, at predefined sampling intervals, the engine and diagnostic parameters received through the OBD-II interface with geographic position signals from the positioning receiver and motion signals from the inertial sensor to generate time-aligned telemetry data corresponding to real -world vehicle operation,wherein the wireless communication module is configured to transmit the time-aligned telemetry data over a mobile data network to the processing unit external to the vehicle,wherein the processing unit comprises physical processors and a non-transitory memory storing executable instructions which, when executed by the processors, cause the processing unit to perform predictive analysis and anomaly detection on the telemetry data so as to derive component-level health indicators and maintenance requirement signals, andwherein the derived component-level health indicators are used to electronically generate and transmit machine-readable maintenance alerts and reports to a user device,thereby achieving a technical effect of continuous electronic monitoring of physical vehicle subsystems, early detection of deviations in vehicle operating behavior, and reduction of unexpected mechanical failures through processor-controlled analysis of sensor-derived vehicle telemetry.

2. The system as claimed in claim 1, wherein the telemetry devices for different vehicle models provide different subsets of sensor channels, and the processing unit is configured to accommodate such variation by mapping available sensor channels from each vehicle into a common feature representation that indicates presence or absence of predefined feature slots using validity indicators and applying learned fusion model to the common feature representation without retraining separate models for each distinct vehicle sensor schema.

3. The system as claimed in claim 1, wherein the microprocessor further comprises a memory storing firmware instructions that, when executed by the microcontroller or microprocessor, cause the apparatus to:i. periodically sample signals from the OBD-II interface, the positioning receiver, and the inertial sensor to generate telemetry data;ii. perform local pre-processing operations including formatting and basic validation of the telemetry data; andiii. transmit the telemetry data to a remote processing unit for hybrid predictive analytics, anomaly detection, and report generation, without performing the hybrid predictive analytics and anomaly detection models on the apparatus.

4. The system as claimed in claim 1, wherein the wireless communication module comprises a cellular modem with integrated Global Navigation Satellite System (GNSS) capability and is configured to transmit telemetry data to the remote processing unit using a publish- subscribe messaging protocol.

5. The system as claimed in claim 1, wherein the processing unit comprises a processor and a memory storing instructions that, when executed by the processor, cause the processing unit to:i. receive the telemetry data from the telemetry device;ii. pre-process the telemetry data by performing noise filtering, outlier removal,scaling, time alignment, and handling of missing values to produce cleaned time series data;iii. apply a hybrid predictive analytics model comprising an Autoregressive Integrated Moving Average (ARIMA) model, a Support Vector Regression (SVR) model, and an Extreme Gradient Boosting (XGBoost) model to the cleaned time-series data to generate predictions of vehicle operating variables;iv. compute residual signals based on differences between the predictions and corresponding observed telemetry values;v. provide features derived from the cleaned time-series data and the residual signals to an anomaly detection module comprising an Isolation Forest estimator configured to assign anomaly scores to the telemetry data;vi. determine, based on the anomaly scores and the predictions, a health state of a vehicle component and a predicted maintenance requirements; andvii. generate and transmit to a user device a report including the health state and the predicted maintenance requirements.

6. The system as claimed in claim 1, wherein the processing unit is configured to:i. receive, for each of time windows, multivariate telemetry comprising sensor measurements from a vehicle, the sensor measurements including a subset of OBD- II parameters and inertial measurements;ii. for each time window, apply at least two predictive models of different model families selected from temporal forecasting models and reconstruction-based models to the telemetry to obtain model outputs, the model families including a temporal forecaster configured to predict future or current sensor values from historical telemetry and a reconstruction-based model or digital-twin model configured to reconstruct sensor values from contemporaneous telemetry; iii. compute, for the time window, residual signals for the temporal forecaster and reconstruction-based model by comparing the model outputs to corresponding observed values;iv. compute, for each sensor channel, a validity indicator that encodes whether the corresponding measurement in the time window is present and within predefined plausibility bounds;v. compute residual-based features that are normalised by the number of valid channels, including feature obtained by aggregating residual magnitudes over onlythose sensor channels whose validity indicators denote valid data;vi. construct, for the time window, a feature vector comprising summary statistics of the residual signals over the window, summary statistics of reconstruction-error features over the window and statistic indicative of a proportion of valid measurements in the window;vii. provide the feature vector to the learned fusion model that outputs a scalar anomaly score or calibrated anomaly probability for the time window; andviii. select anomaly decision threshold for the learned fusion model using a validation dataset that contains predominantly non-fault telemetry such that a false-positive rate measured on the validation dataset does not exceed a predefined budget, and apply the anomaly decision threshold during deployment to generate anomaly alerts.

7. The system as claimed in claim 1, wherein the processing unit is further configured to: i. detect temporal patterns in the residual signals and the anomaly scores that persist over multiple time windows; andii. escalate a maintenance recommendation when such persistent patterns exceed thresholds associated with a likelihood of future failure.

8. The system as claimed in claim 1, wherein the processing unit is deployed in a cloud computing environment and is configured to manage telemetry and analytics for a plurality of vehicles, including fleet vehicles, while maintaining per-vehicle histories of telemetry, predictions, anomalies, and maintenance recommendations.

9. The system as claimed in claim 1, wherein the telemetry data further comprises driver behavior features, including braking intensity, acceleration variability, and cornering characteristics, and the processing unit is configured to incorporate the driver behavior features as inputs to the hybrid predictive analytics model and / or the anomaly detection module.

10. The system as claimed in claim 1, wherein the telemetry device and the processing unit communicate using a publish-subscribe messaging protocol over a mobile or wireless network, and data in transit is protected using authenticated, encrypted communication.

11. The system as claimed in claim 1, further comprising a user-facing application executing on a user device, the user-facing application being configured to:i. receive the report generated by the processing unit;ii. display summary health status, predicted maintenance events, and driver behavior insights; andiii. generate real-time alerts when anomaly scores or predicted failure probabilities exceed thresholds.

12. The system as claimed in claim 1, further comprising an edge processing node located within or near at least one of the vehicles, the edge processing node being configured to locally perform a subset of the pre-processing and anomaly detection steps to support lower-latency alerts in environments with intermittent or limited network connectivity, and to forward summarized results and raw data to the cloud-based processing unit for longterm analysis and model retraining.

13. The system as claimed in claim 1, wherein the system is housed in a compact enclosure configured to plug into an OBD-II port of the vehicle and is powered from the vehicle electrical system while being configured to operate within an automotive temperature and vibration range.

14. The system as claimed in claim 6, wherein the reconstruction-based model comprises a surrogate model of a vehicle subsystem, trained to approximate physical relationships between a plurality of sensor channels for that subsystem, and the reconstruction-error features include multi-sensor deviations from the surrogate model outputs.

15. The system as claimed in claim 6, wherein the learned fusion model comprises a linear model parameterized by a weight vector and a bias term applied to the feature vector, followed by a non-linear activation function that maps a result to an anomaly likelihood value, and wherein parameters of the learned fusion model are obtained by training on labelled historical data comprising examples of nominal vehicle operation and examples of known faults.

16. The system as claimed in claim 6, wherein the processing unit is configured to calibrate raw outputs of the learned fusion model using a calibration function trained on held-outvalidation data to produce calibrated anomaly probabilities, and to choose the anomaly decision threshold as a quantile of the calibrated anomaly probabilities on negative validation samples corresponding to a target false-positive rate.

17. The system as claimed in claim 6, wherein the processing unit is further configured to model correlations between sensor channels as a graph with nodes representing sensors and edges representing learned or engineered dependencies, to compute node-level or graphlevel anomaly scores using a graph-based model, and to include summary statistics of the anomaly scores as additional components of the feature vector.

18. A method for real-time vehicle health monitoring and predictive maintenance, comprising the following steps:i. receiving telemetry data transmitted over a wireless network from the telemetry device mounted in a vehicle, the telemetry data comprising OBD-II parameters, location information, and inertial measurements;ii. pre-processing the telemetry data by performing noise filtering, outlier removal, scaling, time alignment, and handling of missing values to produce cleaned time series data;iii. applying a hybrid predictive analytics model comprising an ARIMA model, an SVR model, and an XGBoost model to the cleaned time-series data to generate predictions of vehicle operating variables;iv. computing residual signals between the predictions and corresponding observed telemetry values;v. providing features derived from the cleaned time-series data and the residual signals to the anomaly detection module comprising an Isolation Forest estimator configured to assign anomaly scores;vi. determining, based on the anomaly scores and the predictions, a health state of the vehicle component and predicted maintenance requirements; andvii. generating a report including the health state and the predicted maintenance requirements and transmitting the report to a user device.

19. The method as claimed in claim 18, wherein the pre-processing in step ii comprises imputing missing values in the telemetry data using historical statistics, interpolation, or model-based estimates, and discarding samples that violate predefined physical plausibilityconstraints.

20. The method as claimed in claim 18, further comprising deriving from the telemetry data, driver behavior features including braking intensity, acceleration variability, or cornering events and using the driver behavior features as additional inputs for the hybrid predictive analytics models and the anomaly detection module when performing steps (iii) through (vi).

21. The method as claimed in claim 18, wherein the anomaly detection module is configured to:i. receive as input a feature vector comprising the residual signals, current telemetry values, and summary statistics over a time window; andii. output anomaly scores that indicate the degree of deviation of current operating behavior from learned normal behavior of the vehicle.

22. The method as claimed in claim 18, wherein determining the health state of the vehicle component in step vi comprises computing a health score as a function of the anomaly scores and residual signals and comparing the health score to thresholds mapped to discrete health categories.

23. The method as claimed in claim 18, wherein the report includes component-level health indicators, predicted time-to-service or mileage-to-service for the vehicle component and driver behavior scores and recommendations to reduce wear and improve safety.

24. The method as claimed in claim 18, further comprising:i. maintaining historical records of telemetry, predictions, and anomaly scores for each vehicle;ii. periodically retraining any of the ARIMA, SVR, and XGBoost models using the historical records to improve prediction accuracy for that vehicle or for a group of similar vehicles; andiii. updating parameters of the Isolation Forest estimator based on newly observed normal operating data to maintain alignment with changing operating conditions.

25. The method as claimed in claim 18, further comprising, prior to training the hybridpredictive analytics model and the anomaly detection module:i. identifying historical telemetry segments associated with rare failure events; ii. generating synthetic variants of such telemetry segments by perturbing or resampling the identified segments while preserving their characteristic failure patterns; andiii. augmenting a training dataset with the synthetic variants to increase an effective number of failure examples used for training.