Vehicle fault intelligent diagnosis and prediction method
By training the TinyLSTM model in the cloud and deploying it on the vehicle, and utilizing lateral federated learning and sensor temporal streams, the problems of traditional vehicle fault diagnosis, such as reliance on human experience, high equipment costs, high risk of misjudgment, and lack of predictability, are solved, thus achieving automated, real-time, and accurate diagnosis of vehicle faults.
Patent Information
- Application Number
- CN202510860141.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-10-24
AI Technical Summary
Traditional vehicle fault diagnosis relies on human experience, is susceptible to human factors, has high diagnostic equipment costs, a high risk of misdiagnosis, a complex diagnostic process, lacks fault predictability, and has significant differences in diagnostic standards among different car manufacturers.
The TinyLSTM model is trained using lateral federated learning. By training it in the cloud and deploying it on the vehicle, it enables automated real-time diagnosis of vehicle faults. It uses sensor time-series data for fault prediction, avoiding the need to rely on cloud resources for network connectivity.
It enables automated, real-time diagnosis of vehicle faults, reduces equipment costs, improves diagnostic accuracy and predictability, adapts to the diagnostic standards of different automakers, and meets the real-time requirements of vehicles during driving and operation.
Smart Images

Figure CN120833640A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of vehicle fault detection, in particular to an intelligent diagnosis and prediction method for vehicle faults. BACKGROUND
[0002] The traditional vehicle fault diagnosis database generation steps include: first, the framework of the ODX database is constructed according to the basic diagnosis specification, including the architecture of the system, the inheritance relationship, the detailed definition of the data parameters, such as the DTC fault code, the fault description, the hardware number, the software number, the data identifier, the ECU definition, the measured value (unit name, maximum value, minimum value), the test plan (step, condition, behavior, result) and the like;
[0003] Then, for each controller, the basic framework is customized according to its specific diagnosis specification. The unsupported services and parameters are removed, and the controller-specific content is added, and finally the complete vehicle ODX database (XML format) is integrated. Finally, combined with the safety algorithm of the vehicle / application program, the corresponding algorithm file is written and embedded into the ODX database of each control unit to ensure safety and consistency. Some key ECUs (such as engine control module, transmission controller) can have a built-in simplified version of the ODX diagnosis database for basic diagnostic functions. These data are usually stored in the embedded memory eMMC, Flash of the ECU, but not every ECU stores the complete ODX file to avoid redundancy. Now the vehicle can integrate the complete or modular ODX file in the central gateway, TBox, or independent diagnostic interface (such as OBD-II port related hardware). These modules usually have larger capacity storage media and support real-time access by diagnostic tools. The cloud (such as the vehicle manufacturer's back-end or third-party platform) usually stores global ODX diagnosis database, historical fault case library, manufacturer proprietary protocol, etc. During diagnosis, the tool can read the vehicle-side data first, and if it does not hit, it can request additional data from the cloud. With the improvement of vehicle network bandwidth (such as 5G-V2X), in the future, the trend is to develop a "light vehicle side + heavy cloud side" mode to reduce the storage cost of the vehicle side and ensure the timeliness of diagnosis.
[0004] Among them, the vehicle fault code generation step includes: ① Hardware level monitoring: ECU analog-to-digital converter ADC or digital input interface, real-time acquisition of sensor signals (such as temperature, pressure, speed, etc.); ② Software level monitoring: Through DEM diagnostic event manager, the signal is reasonably checked, consistency verified, and timing analyzed; ③ Fault determination: threshold overrun (such as sensor voltage <0.1V or >4.9V), timing update frequency anomaly (such as CAN bus signal loss >500ms), multiple sensor data conflict leading to logical contradiction (such as accelerator pedal position and throttle opening do not match); ④ Fault confirmation mechanism: according to the specification in the ODX diagnostic database, for example: Debounce anti-jitter time window (avoid transient interference false trigger, need to continuously abnormal for more than a set time, such as sensor low voltage for 5 seconds), Fault Counter counter (cumulative number of times, such as 3 consecutive driving cycles trigger to confirm fault), generate DTC fault code, status flag and fault log, etc.; ⑤ Store fault information to the target ECU eMMC or Flash; eMMC is used for high data storage (such as freeze frame, historical DTC log); Flash stores core DTC list and status flag, ensuring that power failure does not lose; ⑥ If Autosar architecture is used, the logic chain is: sensor data→BSW layer signal processing (acquire raw signals and convert them into physical quantities, such as voltage→temperature; periodically execute signal verification algorithm, such as 100ms task period)→DEM event trigger (unified management of all diagnostic event identifiers EventID, and association with DTC fault code)→DTC storage. Fault code reading process and interaction logic: ① DST fault diagnosis instrument connects to the vehicle's diagnostic port through OBD-II interface or other special interface, ensuring stable connection for data communication; ② DST sends specific UDS diagnostic command; ③ The target ECU will respond to the diagnostic command, read its own DTC and status information, and then return to the DST; If Autosar architecture is used, DCM diagnostic communication manager is responsible for the above logic processing; ④ DST receives response data and parses the specific meaning of the fault code; If more detailed information is needed, fault snapshot (such as including fault occurrence time, mileage, etc.) can be read through UDS diagnostic service; Fault information clearing mechanism: ① The diagnostic personnel manually send DTC clearing service instructions through the DST fault diagnosis instrument to clear the fault information of the target ECU; ② The ECU detects that the fault disappears and meets the conditions (such as 3 consecutive driving cycles without recurrence); ③ If Autosar architecture is used, DCM is responsible for controlling DTC clearing authority (security authentication is required).
[0005] Problems in the above process include:
[0006] ① It relies on manual experience and is easily affected by human factors. Its accuracy is difficult to guarantee and requires a high level of experience from the diagnostic personnel.
[0007] ② The cost of diagnostic equipment is high. DST-based fault diagnosis methods require professional diagnostic equipment and tools, which requires a large capital investment and fixed testing and diagnostic space and basic facilities. This is a considerable financial burden for small 4S repair shops or individual car owners.
[0008] ③ Risk of misdiagnosis. Although the DST diagnostic instrument has high accuracy, there is still a risk of misdiagnosis. The quality of the diagnostic equipment and the professional level of the operator will affect the accuracy of the diagnosis;
[0009] ④ The diagnostic process is complex. The DST fault diagnosis method requires quantitative analysis of diagnostic parameters. The operation is complex and requires operators to have high cultural quality and professional knowledge, which is a challenge for some non-professionals.
[0010] ⑤ Each automaker may define different fault codes and fault information based on the specific diagnostic specifications for each vehicle component. Only after a vehicle failure occurs does manual intervention, using a DST fault diagnostic instrument, acquire and analyze fault information. This approach lacks proactive foresight and is essentially a post-failure approach. Summary of the Invention
[0011] In order to solve the above technical problems, the present invention provides an intelligent vehicle fault diagnosis and prediction method, which can automatically and in real time predict vehicle faults.
[0012] The present invention solves the above problems through the following technical solutions:
[0013] A vehicle fault intelligent diagnosis and prediction method, comprising:
[0014] Step S10: Training a TinyLSTM model in the cloud based on horizontal federated learning (HFL) of sensor time series streams from multiple car companies;
[0015] Step S20: Deploy the trained TinyLSTM model to the vehicle.
[0016] Step S30: The TinyLSTM model on the vehicle side receives the vehicle sensor timing stream and outputs the intelligent diagnosis and prediction results of the vehicle fault.
[0017] The fault diagnosis and prediction method of the present invention is completely implemented on the vehicle side, without the need to connect to the Internet and borrow cloud resources, meeting the real-time requirements of fault diagnosis and prediction during vehicle driving or scene operation.
[0018] Furthermore, the step S10 specifically includes:
[0019] Step S11: Deploy the HFL central parameter server cluster and HFL client cluster hardware environment on the cloud. Each car company builds its own HFL client cluster and collects real-time sensor time series data for its target vehicles based on its own models and reports it to the cloud. Data reporting from the vehicle to the cloud generally uses 4G / 5G wireless cellular networks.
[0020] In step S12, the central parameter server cluster delegates the initial global model parameters to each client cluster. The client clusters use their local vehicle fault data to train the TinyLSTM model. The training results (gradients or model parameters) of each round are encrypted and uploaded to the central parameter server cluster. The central parameter server cluster delegates the global model parameters calculated by federated average gradient aggregation after each round of iterative training to each client cluster. Multiple rounds of iterative training are repeated until model training is complete. The connection between the HFL client cluster and the HFL central parameter server cluster can be a VPN based on the Internet or a DDN fiber-optic private network.
[0021] Furthermore, the step S12 specifically includes:
[0022] Step 1), the central parameter server initializes the global model parameters;
[0023] Step 2) The central parameter server randomly selects multiple clients and sends the current global model parameters in plain text to the selected clients;
[0024] Step 3) Determine whether the training has reached the preset number of iterations, or whether the model convergence performance meets the requirements. If so, end the training process; otherwise, proceed to the next step;
[0025] Step 4) The client uses its own local vehicle fault data to train the model;
[0026] Step 5) The client uses Nesterov to accelerate SGD and calculate the gradient to update its own model weight value;
[0027] Step 6) The client performs GradNorm normalization on the calculated gradient to prevent the model from diverging;
[0028] Step 7) The client adds a random mask to the normalized gradient to implement gradient obfuscation and prevent attackers from inferring the internal structure and parameters of the model and causing model attacks;
[0029] Step 8) The client uses the Paillier encryption scheme and the public key PK generated by the central parameter server to homomorphically encrypt the updated model weight value;
[0030] Step 9), the client sends the encrypted model weight value to the central parameter server, which cannot be decrypted in the middle;
[0031] Step 10), the central parameter server receives the ciphertext sent by all clients;
[0032] Step 11), the central parameter server uses the Shapley algorithm to monitor the contribution of each client node to the gradient prediction result in real time, and automatically isolates the participants with abnormal gradients;
[0033] Step 12), the central parameter server performs Paillier homomorphic addition operation to generate aggregated ciphertext;
[0034] Step 13), the central parameter server uses the private key SK generated by itself to decrypt the aggregated ciphertext to obtain the aggregated gradient, that is, the current global model parameter;
[0035] Step 14), return to step 2) and iterate until the model training end condition is reached.
[0036] Further, the TinyLSTM model comprises a TinyLSTM input layer, a core LSTM layer, a TinyLSTM feature enhancement layer, and a TinyLSTM output layer. The TinyLSTM model balances the calculation efficiency and feature extraction capability through modular design, wherein:
[0037] The TinyLSTM input layer is used to receive sensor data of a vehicle and perform sensor data preprocessing (including data normalization, missing value filling, noise filtering, etc.) before inputting the core LSTM layer.
[0038] The core LSTM layer combines the input gate and the forget gate of the traditional LSTM into an update gate to realize LSTM time step recursion, and is used to extract short-term / long-term time sequence dependence in the input data.
[0039] The TinyLSTM feature enhancement layer includes a lightweight self-attention mechanism, which is used to receive the output vector of the core LSTM layer in parallel. After calculation by the lightweight self-attention mechanism, 8 feature vectors are output, and then dequantization operation is performed. The output result of the core LSTM layer and the output result of the lightweight self-attention mechanism are input into a cross-layer residual connection layer, which performs feature fusion connection calculation and outputs the calculation result of the cross-layer residual connection.
[0040] The TinyLSTM output layer adopts a full connection network, and the input is a combined vector calculated through cross-layer residual connection. The first linear layer linearly transforms the input combined vector, and the feature dimension is reduced from 64 to 16. The first linear layer uses an activation function Leaky ReLU, and the slope of the negative input is alpha = 0.1 to alleviate the gradient disappearance. The last linear layer linearly transforms the vector calculated by the activation function Leaky ReLU again, and the feature dimension is reduced from 16 to 1. The last layer uses a Sigmoid function to constrain the output range, adapt to the probability interpretation requirement, and generate a vehicle fault diagnosis prediction result probability.
[0041] Further, the core LSTM layer calculation formula is as follows:
[0042] u t =σ(W u ·[h t-1 ,x t ]+b u );
[0043]
[0044] O t =σ(W o ·[h t-1 ,x t ]+b o );
[0045] h t =O t ⊙tanh(C t );
[0046] Wherein, u t is the update gate output, sigma is the sigmoid activation function, W u is the weight matrix of the update gate, h t-1 is the hidden state of the previous time step t-1, x t is the input at the current time t, and b u is the offset vector of the update gate. is the cell state of the candidate memory unit at the current time t, W C is the weight matrix of the candidate memory unit, and b C is the offset vector of the candidate memory unit; represents the element-wise multiplication Hadamard product, which is used for multiplication operation on corresponding position elements in the matrix or vector; tanh represents the hyperbolic tangent activation function; C t is the memory cell state at the current time step t, C t-1 is the memory cell state at the previous time step t-1; O t is the activation value of the output gate at time t, Wo is the weight matrix of the output gate, b o is the bias vector of the output gate; h t is the hidden state at time t; the core LSTM layer receives the input x t and the previous hidden state h t-1 at each time step, updates the cell state C t and outputs h t , realizes time step recursion.
[0047] Further, the training of the TinyLSTM model uses a cross-entropy loss function and a selected stochastic gradient descent method as an optimizer, specifically including:
[0048] A1, in a main for loop, using the loaded data set, the model is trained;
[0049] A2, the prediction result of the model and the actual label distribution are calculated by cross-entropy loss, and the difference is calculated;
[0050] A3, perform back propagation, calculate the gradient of the loss function to the model parameters;
[0051] A4, set the maximum threshold of the gradient norm (default L2 norm) to 1.0, and perform gradient clipping and parameter constraint on the original model tensor;
[0052] A5, embed a small for loop, judge the dimension of all parameters of the model, if any parameter dimension>1, limit the value of the model parameter data to [-2,2];
[0053] A6, update the model parameters according to the current calculated gradient value;
[0054] A7, calculate the value of the cumulative cross-entropy loss;
[0055] A8, when the main for loop ends, return the value of the average cross-entropy loss of the training.
[0056] Further, the TinyLSTM model verification and evaluation method is:
[0057] B1, the preloaded data set includes the input data set X and the predicted result true value data set y required for model verification and evaluation;
[0058] B2, temporarily disable the automatic differentiation function, that is, disable gradient calculation and history tracking torch.no_grad(), all operations involving tensors with requires_grad=True will not generate a computation graph or store gradients, and directly skip the intermediate state recording required for back propagation, thereby saving memory and accelerating calculation in the inference or verification stage;
[0059] B3, using a for loop, the model makes real-time inference prediction on the input dataset X, the prediction result pred and the true value data y are stored in the arrays y_pred and y_true respectively;
[0060] B4, after the for loop ends, the arrays y_pred and y_true are converted into NumPy arrays respectively;
[0061] B5, the true value in the NumPy array y_true and the predicted value in y_pred>0.5 are input into the f1_score function to calculate the index of evaluating the performance of the classification model, the f1 score takes into account the performance of accuracy and recall;
[0062] B6, the true value in the NumPy array y_true and the predicted value in y_pred are input into the roc_auc_score function to calculate the area score under the ROC receiver operating characteristic curve to measure the prediction performance of the binary classification model; ROC AUC is a threshold-independent index, its value range is [0.5, 1], the larger the value, the better the model performance;
[0063] B7, return the f1_score calculation score and roc_auc_score calculation score to comprehensively verify and evaluate the prediction performance of the model.
[0064] Further, the method for quantization deployment of the TinyLSTM model is:
[0065] C1, use the torch.quantization.quantize_dynamic function to realize dynamic quantization of the {nn.Linear} linear fully connected layer, save the quantized model quantized_model to adapt to the vehicle end deployment, the function of torch.quantization.quantize_dynamic is to convert the model weight from a floating-point number to a low-precision integer (usually INT8) method to reduce the model size and speed up the inference speed. This function is particularly suitable for CPU devices, and only the weight is quantized, while the activation remains FP32 precision, which helps to reduce the computational complexity and memory usage while maintaining high model accuracy;
[0066] C2, use the torch.jit.script function to convert the quantized model into a runnable script, the advantage of using the torch.jit.script function to convert is that it can reduce the overhead in the model execution process, because it eliminates the overhead of the Python interpreter;
[0067] C3, use the torch.jit.save function to save the converted script model to a cross-platform format tinylstm_quantized.pt. When saving the model using torch.jit.save, the overhead during model execution can be reduced because it eliminates the overhead of the Python interpreter. The saved model can be executed in an environment without Python dependencies, making the deployment of the model more flexible and efficient.
[0068] Further, the method for testing the TinyLSTM model is:
[0069] D1, real-time inference test: unsqueeze the test data set as sample_input; then temporarily disable the automatic differentiation function, i.e. disable gradient calculation and history tracking torch.no_grad();
[0070] D2, real-time inference of the input sample sample_input through the quantized model quantized_model to obtain the prediction result; quantized_model(sample_input).item() represents that the input sample sample_input is subjected to real-time inference through the quantized model quantized_model to obtain the prediction result; the item function can be used to obtain the prediction result of a single element;
[0071] D3, if the vehicle fault diagnosis prediction result probability is greater than 0.5, it means that the vehicle fault is wrong, otherwise it means normal.
[0072] Further, the step S20 specifically comprises: after the client cluster completes the TinyLSTM model training, the TinyLSTM model is distributed to the target vehicle for writing upgrade through the DOTA method.
[0073] The TinyLSTM model deployed on the vehicle end after quantization has a parameter size of 412KB, a real-time inference required cache of 89KB, a single inference delay of 3.7ms, and a model update DOTA differential parameter upgrade package of only 28KB, which can be completed within 5 seconds. The model fault real-time detection coverage rate is exemplified as follows: engine misfire (98.2%), battery over-temperature (95.6%), and brake pad wear (93.1%).
[0074] Compared with the prior art, the present application has the following advantages and beneficial effects:
[0075] (1) The present application meets the real-time requirement of vehicle fault diagnosis prediction when the vehicle is driving and working, and can diagnose and predict unknown vehicle faults in real time. Moreover, the whole process is automated and does not require human intervention.
[0076] (2) The application deploys a TinyLSTM deep learning model trained based on HFL transverse federal learning at the vehicle end, realizes real-time fault diagnosis prediction of the vehicle end (municipal sanitation vehicles, tractor vehicles, engineering special vehicles, etc. electric commercial vehicles) based on sensor time series flow (such as engine speed, vibration frequency, battery temperature, driving mileage, etc.), and can be used as an effective supplement to manual fault code diagnosis, and can diagnose and predict unknown vehicle faults in real time.
[0077] (3) If the vehicle cannot timely predict and diagnose possible faults before the fault code is generated, the vehicle driving operation risk is extremely high, which may cause dangerous working conditions, and this state may have a serious impact on personal safety, operation device system, and vehicle integrity. For example, the brush disc on the sanitation washing and sweeping vehicle cannot be stopped, the water spraying vehicle sprays water unnecessarily, the concrete mixing vehicle cannot normally load and unload materials, etc. The application meets the real-time requirement of vehicle fault diagnosis and prediction at the vehicle end during driving and engineering operation.
[0078] (4) Different vehicle enterprises each have complete sensor data (such as engine speed, vibration frequency, battery temperature, driving mileage, etc.) of the same type of vehicle, and for the operation scene, each functional vehicle type has specific sensor time series flow (such as the water level sensor of the clean water tank / polluted water tank of the sanitation washing and sweeping vehicle, the rear door full opening / tightening sensor, the total water valve opening / closing sensor, the support rod lifting sensor, the tank body return sensor, etc.), the characteristic space of these sensor time series flow is the same, but the data samples are not overlapped. Therefore, the transverse federal learning training TinyLSTM adopted by the application can jointly model under the premise that multiple data owners do not share original data in the vehicle fault diagnosis scene, and train by submitting the parameters or gradients of the model, so as to protect the data privacy while improving the performance of the model.
[0079] (5) The application trains a TinyLSTM deep learning model based on HFL horizontal federated learning technology by means of the software and hardware resources of each vehicle model project group of each OEM vehicle enterprise or vehicle enterprise group, and the cost budget is extremely low; TinyLSTM can process various types of electric commercial vehicle end sensor real-time stream data, support vehicle real-time fault diagnosis and prediction, meet the technical conditions of lightweight design, low delay design, multi-dimensional time series data adaptation, etc.; the TinyLSTM model parameter size after quantization deployed on the vehicle end is 412KB, the cache required for model real-time inference is 89KB, the single inference delay is 3.7ms, the model update DOTA differential parameter upgrade package is only 28KB, and the vehicle end deployment can be completed within 5 seconds; by simplifying the TinyLSTM model parameters and structure, adapting to the limited environment of vehicle end embedded hardware, efficient time series modeling is realized. The hierarchical structure and connection logic of the TinyLSTM model fully adapt to the real-time and reliability requirements of the vehicle end, and provide a landable AI infrastructure for the intelligent fault diagnosis and predictive maintenance of the next generation of electric commercial vehicles. BRIEF DESCRIPTION OF DRAWINGS
[0080] Figure 1 The figure is an AI model architecture diagram trained based on HFL horizontal federated learning in the application;
[0081] Figure 2 The figure is a core flowchart of the vehicle fault diagnosis and prediction model trained by HFL horizontal federated learning in the application;
[0082] Figure 3 The figure is a model parameter or gradient data transmission and decryption / encryption flowchart in the application;
[0083] Figure 4 The figure is a TinyLSTM model overall architecture diagram in the application;
[0084] Figure 5 The figure is a TinyLSTM input layer structure diagram in the application;
[0085] Figure 6 The figure is a core LSTM layer: double-gate simplified LSTM structure diagram in the application;
[0086] Figure 7 The figure is a core LSTM layer: two-stage double-gate simplified LSTM time step recursion structure diagram in the application;
[0087] Figure 8 The figure is a TinyLSTM feature enhancement layer: lightweight self-attention mechanism model composition structure diagram in the application;
[0088] Figure 9 The figure is a Scaled Dot-Product Attention model composition structure diagram in the application;
[0089] Figure 10 The computational graph for the inverse quantization operation in the present application;
[0090] Figure 11 The computational graph for the generation of the exponentiation table in the present application;
[0091] Figure 12 The flow chart for the calculation of the output probability value by simulating the Softmax function through the exponentiation table in the present application;
[0092] Figure 13 The computational graph for the cross-layer residual connection in the present application;
[0093] Figure 14 The full connection network structure diagram of the output layer: vehicle fault prediction in the present application. DETAILED DESCRIPTION
[0094] The present application will be further described in detail below in conjunction with embodiments, but the embodiments of the present application are not limited thereto. Before introducing the specific embodiments of the present application, first, the abbreviations involved in the present application are explained as follows:
[0095] 5G-V2X: 5G-Vehicle to Everything, vehicle connected with everything based on 5G network;
[0096] AUC: Area Under Curve, defined as the area surrounded by the coordinate axes under the ROC curve; Autosar: AUTomotive Open Systems ARchitecture, automotive open system architecture;
[0097] BSW: Basic Software Layer, basic software service;
[0098] DCM: Diagnostic Communication Manager, diagnostic communication manager;
[0099] DDN: Digital Data Network, digital data private network;
[0100] Debounce: Debounce;
[0101] DEM: Diagnostic Event Manager, diagnostic event manager;
[0102] DH: Diffie-Hellman, Diffie-Hellman key exchange algorithm;
[0103] DIM: Dimension, dimension
[0104] DOTA: Data Over The Air, data download over the air;
[0105] DST: Discovery Switching Troubleshooting, discovery, switching, troubleshooting;
[0106] DTC: Diagnostic Trouble Code, diagnostic trouble code;
[0107] DualGateLSTM: Dual-gate streamlined long short-term memory network;
[0108] ECU: Electronic Control Unit, electronic control unit;
[0109] Event ID: event identifier;
[0110] Fault Counter: fault counter;
[0111] FedAvg: Federated Averaging, federated averaging algorithm;
[0112] GradNorm: Gradient Normalization, gradient normalization;
[0113] HFL: Horizontal Federated Learning, horizontal federated learning;
[0114] L2 norm: Euclidean norm, one of the most commonly used norm types in vector space;
[0115] Leaky ReLU: Leaky linear rectifier function, an improved ReLU activation function;
[0116] LR: Learning Rate, learning rate;
[0117] LSTM: Long Short-Term Memory, long short-term memory network;
[0118] LUT: Lookup Table, a method that speeds up calculations by pre-calculating and storing function values;
[0119] NAG: Nesterov Accelerated Gradient, Nesterov accelerated gradient descent method;
[0120] OBD-II: the Second On-Board Diagnostics, the second generation of on-board diagnostic system;
[0121] ODX: Open Diagnostic data eXchange, an open diagnostic data format;
[0122] Paillier: a public-key cryptosystem that supports additive homomorphism;
[0123] ReLU: Rectified Linear Unit, a linear rectifier function, also known as a rectified linear unit, is a commonly used activation function in neural networks;
[0124] ROC: Receiver Operating Characteristic Curve, Receiver Operating Characteristic Curve, also known as Sensitivity Curve;
[0125] SGD: Stochastic Gradient Descent, Stochastic Gradient Descent;
[0126] Shapley: Shapley value calculation method;
[0127] Sigmoid: S-shaped function, also known as S-shaped growth curve, which is often used as an activation function in neural networks;
[0128] Softmax: Softmax, a mathematical function used for multi-phenolic problems, which can convert any real number vector into a probability distribution;
[0129] TBox: Telematics BOX, a car networking communication box;
[0130] TinyLSTM: Tiny Long Short-Term Memory, Tiny Long Short-Term Memory network;
[0131] Transformer: Transformer, a neural network architecture that uses self-attention mechanisms to capture relationships between elements in a sequence;
[0132] UDS: Unified Diagnostic Services, Unified Diagnostic Services.
[0133] Embodiment:
[0134] An intelligent diagnosis and prediction method for vehicle faults, comprising:
[0135] Step S10, training TinyLSTM model based on HFL horizontal federated learning of multiple OEM sensor time sequence streams in the cloud;
[0136] Step S20, deploying the trained TinyLSTM model to the vehicle end;
[0137] Step S30, the TinyLSTM model at the vehicle end receives the vehicle sensor time sequence stream and outputs the intelligent diagnosis and prediction result of vehicle failure.
[0138] The OEMs of new energy electric commercial vehicles such as sanitation vehicles, tractor vehicles, and engineering special vehicles generally develop multiple functional vehicle models at the same time, each of which has complete sensor data (such as engine speed, vibration frequency, battery temperature, driving mileage, etc.) of the same type of vehicle. In addition, for the working scene, each functional vehicle model has a specific sensor time sequence stream (such as the water level sensor of the clean water tank / polluted water tank of a sanitation washing and sweeping vehicle, the rear door full opening / contracting sensor, the total water valve opening and closing sensor, the support rod lifting sensor, the tank back position sensor, etc.). The feature spaces of these sensor time sequence streams are the same, but the data samples are not overlapped. Therefore, HFL horizontal federated learning based on multiple OEM electric commercial vehicle sensor time sequence streams in the cloud is considered to train a TinyLSTM model, and then the TinyLSTM model is deployed to the vehicle end for real-time vehicle failure diagnosis and prediction. This failure diagnosis and prediction method is completely implemented on the vehicle end without the need to use cloud resources, and meets the real-time requirements of electric commercial vehicles for failure diagnosis and prediction during driving or scene operation.
[0139] The TinyLSTM model architecture trained based on HFL horizontal federated learning is as shown in Figure 1
[0140] 1) Vehicle failure prediction is mainly for new energy electric commercial vehicles, including but not limited to sanitation washing and sweeping vehicles, sanitation fog cannon vehicles, hook arm garbage trucks, compression type garbage trucks, mine transport vehicles, engineering dump trucks, concrete mixing trucks, and other functional commercial vehicle models. These vehicles are ToG and ToB commercial vehicles, and the computing power of the vehicle end chips NPU, GPU, MPU, etc. is relatively small, so real-time failure diagnosis and prediction must be completely performed on the vehicle end, and a TinyLSTM small model trained and optimized through HFL horizontal federated learning must be used.
[0141] 2) Various electric commercial vehicles may come from different OEMs or vehicle model project groups under an OEM group, so the cloud needs to deploy HFL center parameter server clusters and HFL client cluster hardware environments;
[0142] 3) Each OEM builds its own HFL client cluster and combines its own produced vehicle models to perform real-time sensor time sequence data acquisition and reporting on the target vehicle.
[0143] 4) HFL center parameter server cluster, which downgrades (transfers in plaintext) the initial global parameters of the model and the global parameters calculated after each round of iteration training through federated average gradient aggregation to each OEM's own HFL client cluster. Each OEM's own HFL client cluster uses the vehicle data collected by itself to train the TinyLSTM model in the local cluster environment, and uploads (uses Paillier encrypted ciphertext) the training result (gradient or model parameter) of each round to the HFL center parameter server cluster through multiple rounds of iteration training;
[0144] 5) After the HFL client cluster completes the TinyLSTM model training, the model is distributed to the target vehicle for upgrade through the DOTA method.
[0145] The technical implementation principle of HFL horizontal federated learning training TinyLSTM model is as follows:
[0146] HFL horizontal federated learning: a distributed machine learning paradigm suitable for scenarios where data features are the same and sample subjects are different. In the vehicle fault diagnosis scenario, multiple car companies have the same sensor types (such as engine speed, frequency, etc.), but the data comes from different vehicle groups. HFL can be used to jointly model without sharing the original data of multiple data owners, and the model parameters or gradients are submitted for training, thereby improving the performance of the model while protecting data privacy;
[0147] Mathematical expression: Let there be K participants (multiple car companies, or multiple branches of a car company group, multiple vehicle model project groups), and the local data set of each participant k is D k = {(x i ,y i )}, where (x i ,y i ) represents the i-th training sample; the feature space X is the same, but the sample ID is different; let the loss function of the i-th local model training of each participant be f i (ω, x i ,y i ), where ω represents the global model parameter, which is obtained by minimizing the weighted loss function;
[0148] Then: where F k (ω) represents the local loss function of each participant;
[0149] The global optimization goal is: where N represents the total sample size;
[0150] The implementation steps of applying HFL transverse federal learning to train the TinyLSTM model for vehicle fault diagnosis prediction are as follows:
[0151] In the transverse federal learning training process, the federal average algorithm is used to solve the model training problem under distributed data privacy protection.
[0152] 1. Core idea of federal average algorithm
[0153] Client local training: multiple clients (each vehicle project or each vehicle project group under each vehicle enterprise) independently update the model parameters using local data.
[0154] Server aggregation parameters: the central parameter server collects the model updates uploaded by the clients, generates a global model through weighted averaging, and iteratively optimizes until convergence.
[0155] 2. Core workflow and key algorithm, combined Figure 2 as shown:
[0156] (1) The central parameter server initializes the global model.
[0157] (2) The central parameter server randomly selects multiple clients and sends the current global model to the selected clients in plaintext.
[0158] (3) Determine whether the training has reached the preset number of iterations or the model convergence performance meets the requirements.
[0159] (4) If the judgment is yes, the training process is completed; if the judgment is no, continue with the following steps.
[0160] (5) The client trains the model using its own local vehicle fault data.
[0161] (6) The client uses Nesterov accelerated SGD and calculates the gradient to update its own model weight value.
[0162] (7) The client performs GradNorm normalization processing on the calculated gradient to prevent model divergence.
[0163] (8) The client adds a random mask to the gradient to achieve gradient obfuscation, preventing attackers from inferring the internal structure and parameters of the model to cause model attacks.
[0164] (9) The client uses the Paillier encryption scheme to use the public key PK generated by the central parameter service to homomorphically encrypt the updated weight parameters.
[0165] (10) The client sends the encrypted weight parameters to the central parameter server, which cannot be decrypted in the middle.
[0166] (11) The central parameter server receives all the ciphertext sent by the clients;
[0167] (12) The central parameter server uses the Shapley algorithm to monitor the contribution of each client node to the gradient prediction result in real time, and automatically isolates the participants with abnormal gradients;
[0168] (13) The central parameter server performs Paillier homomorphic addition operation to generate aggregated ciphertext;
[0169] (14) The central parameter server uses its own generated private key SK to decrypt the aggregated ciphertext to obtain the aggregated gradient;
[0170] (15) Repeat step (2) to enter the loop iteration until the model training end condition is reached.
[0171] Specifically:
[0172] Model initialization: the central parameter server generates a global initial model ω global (0) and determines the hyperparameters (client selection ratio C, local training round E, learning rate η). In each round of communication, the central parameter server randomly selects K = C·M clients (M is the total number of clients). The central parameter server distributes the current global model ω global (t) to the selected clients.
[0173] Local training: each car enterprise (or each vehicle model project group under the car enterprise group) client trains the model using its own local vehicle fault data (such as battery cell failure, turbocharger stall, etc.) for E rounds (e.g., E = 3), adopts Nesterov accelerated SGD (learning rate η = 0.001, batch size batch_size = 64) and calculates the gradient Then normalize (prevent model divergence) and obfuscate (add random mask) the gradient .
[0174]
Local training update parameters
[0175] The client k performs E rounds of gradient descent locally, and each time uses a small batch of data batch_size to update the parameters:
[0176]
[0177] Where ω k (t) is the parameter of client k in the tth round, ω k (t+1) is the parameter of client k in the t+1th round, η is the learning rate, and F k (ωk (t) ) is the local loss function of client k at round t;
[0178]
Adopt Nesterov accelerated SGD random gradient descent
[0179] Nesterov accelerated gradient descent is a momentum acceleration algorithm, which accumulates the update direction of the previous times by introducing a momentum term, so as to be more stable when the gradient direction changes dramatically. The difference between Nesterov accelerated gradient and traditional momentum method is that Nesterov applies momentum update before calculating the current gradient, which can better predict the coming gradient change;
[0180] Mathematical representation of Nesterov accelerated random gradient descent:
[0181] Calculate the estimated point: y k = x k + γv k ;
[0182] Calculate the gradient g k at the estimated point y
[0183] Update momentum: v k+1 = γv k - ηg k ;
[0184] Update parameters: x k+1 = x k + v k+1 ;
[0185] Where x k represents the parameters of the current step k, v k represents the momentum term of the current step k, γ is the momentum factor (usually set between 0-1), η is the learning rate, y k represents the value of the estimated point of the current step k, g k represents the gradient at the estimated point y k , x k+1 represents the parameters of the next step of the current step k, v k+1 represents the momentum term of the next step of the current step k;
[0186]
Normalize the gradient g to prevent the model from diverging
[0187] Gradient normalization is a method used in multi-task learning to balance the gradients of different tasks. The purpose is to normalize the gradient of each task so that the gradients of different tasks can be compared and balanced on a unified scale, avoiding the negative impact of excessively large or small gradients on training.
[0188] Mathematical representation of gradient GradNorm normalization:
[0189] Where T represents the number of tasks, for the t-th task (t = 1, 2…T), g t represents the normalized value of the t-th task gradient;
[0190] represents the L2 norm (Euclidean norm) of the t-th task gradient, which measures the length of the gradient vector;
[0191] represents the sum of the L2 norms of all task gradients;
[0192]
Adding random masks to gradients to achieve gradient obfuscation
[0193] In horizontal federated learning, data is distributed across multiple clients, and secure aggregation techniques are used to protect data privacy and model security. Random mask technology can effectively prevent gradient leakage and protect users' personal information security;
[0194] The technical principle of gradient obfuscation is to use random masks to interfere with the calculation process of the gradient, making it difficult for attackers to infer the internal structure and parameters of the model through gradient information. Specifically, this technology adds random noise to the gradient to obfuscate the true gradient information, making gradient-based attack methods ineffective;
[0195] Implementation steps of gradient obfuscation:
[0196] Generate random masks: In horizontal federated learning, each client uses the DH (Diffie-Hellman) algorithm to generate public and private keys, and uses these keys to generate random masks that match the size and shape of the model parameter gradients;
[0197] Apply random masks: During the gradient aggregation process, each client adds its parameter gradient to the random mask, then sends the gradient with the mask to the central parameter server for aggregation. The central parameter server integrates all the gradients with masks and updates the model parameters using the average algorithm;
[0198] Protection mechanism: In this way, even if some clients drop out or data is stolen, attackers cannot recover the original data or infer the model parameters through gradient information, because the random mask will cancel out the true gradient information;
[0199] Parameter aggregation
[0200] The client sends the trained model parameters or gradients to the central parameter server after Paillier homomorphic encryption, uploads the gradient ω k (t+1) , ensures that it cannot be decrypted in the middle of the way, and the central parameter server uses the federated average weighted formula to aggregate these parameters to generate a new global model
[0201] where represents the total data volume of all participating clients;
[0202] The weight indicates that the client with large data volume has a greater impact on the global model;
[0203] Paillier encryption scheme
[0204] is a homomorphic encryption scheme based on public key cryptography, which allows direct calculation (such as addition and multiplication) on encrypted data, and the calculation process does not leak any information of the original text. The result of the calculation is still encrypted, and the user with the key decrypts the processed ciphertext data to get exactly the result of the processed original text;
[0205] The Paillier encryption scheme for parameter aggregation in horizontal federated learning uses the difficulty of discrete logarithms to ensure its security. In the encryption process, a pair of public and private keys is first generated, the public key is used to encrypt data, and the private key is used to decrypt data. When the plaintext needs to be added or multiplied, the plaintext can be encrypted using the public key to get the ciphertext, and then another plaintext can be encrypted using the public key, and the two ciphertexts can be combined to get the final ciphertext. In the decryption process, the private key is used to decrypt the ciphertext to get the addition or multiplication result of the plaintext;
[0206] Model parameter or gradient data transmission and encryption and decryption process, combined with Figure 3 as shown:
[0207] 1) The central parameter server generates a key pair: the central parameter server creates a public and private key pair (PK, SK) of the Paillier algorithm, where PK is the public key (used for encryption), and SK is the private key (used for decryption);
[0208] 2) Central parameter server public key distribution: the central parameter server securely distributes the public key PK to all clients (such as through digital certificates or secure channels);
[0209] 3) Client local training: the client trains the model on the local data set and calculates the gradient
[0210] 4) Gradient encryption: Client encrypts the gradient plaintext using the public key PK of the central parameter server, generating ciphertext
[0211] 5) Transmission of ciphertext: Client uploads the encrypted gradient C i to the central parameter server, which cannot be decrypted in the middle (because the private key SK is only held by the server);
[0212] 6) Homomorphic addition: The central parameter server receives all the ciphertexts {C1, C2, …, C n} sent by the clients, then performs the homomorphic addition operation of Paillier, generating aggregated ciphertext (Note: The plaintext space of Paillier is the modulus domain, which needs to be processed according to the specific parameter range):
[0213]
[0214] 7) Decryption of aggregated results: The central parameter server uses the private key SK to decrypt C agg , obtaining the aggregated gradient
[0215] 8) Plain text delivery: The central parameter server returns the aggregated gradient (or updated model parameters) to each client in plaintext form (No need for secondary encryption here, because the aggregated gradient does not contain the gradient sensitive information of individual clients, and plaintext transmission can reduce computational overhead);
[0216] 9) Client updates local model weights according to , completing this round of federated learning iteration.
[0217]
Model update
[0218] The central parameter server returns the aggregated model parameters ω global (t+1) to each client, and the client updates its own model weight values according to these parameters;
[0219]
Iterative training
[0220] Repeat the above steps until the preset number of iterations is reached or the model convergence performance meets the requirements. During the iteration process, the contribution of each node is monitored in real time (Shapley value calculation), and the participants with abnormal gradients (such as standard deviation > 3σ) are automatically isolated.
[0221] In the horizontal federated learning model, the Shapley value method is used to calculate the contribution of each client to the gradient prediction result, and its core idea is the fair distribution based on marginal contribution, which is mathematically expressed as:
[0222]
[0223] where N represents the set of the number of client training nodes in the horizontal federated learning, ф i represents the Shapley value contribution of the client training node i, S is the subset of other client training node elements except node i, f(S∪{i}) is the model prediction value after the subset S joins node i, f(S) is the model prediction value of the subset S, and |N|! is the product of the factorial of all client node elements in the set N, and |S|! is the product of the factorial of other client training node elements except node i.
[0224] The formula can be used to calculate the average marginal contribution of the client training node i to all possible combinations, ensuring that the contribution of each client training node is fairly distributed.
[0225] The TinyLSTM model architecture design and related implementation algorithms are as follows:
[0226] TinyLSTM is a lightweight time series model optimized for vehicle embedded hardware, and its core design goal is to perform real-time high-precision fault diagnosis under low power consumption (<1W), low memory occupation (<1MB), and low latency (<10ms).
[0227] (1) Data index input
[0228] Taking a new energy pure electric sanitation cleaning and sweeping vehicle as an example, the vehicle model is used for real-time fault diagnosis and prediction, and 37-dimensional data indicators are input (other pure electric commercial vehicles may have more or less input data indicator dimensions depending on the working function), and the specific data and units include: power system data (5): engine speed (RPM), battery energy consumption (kWh), motor temperature (℃), battery pack temperature (℃), motor torque (N·m); sampling frequency is 10-100Hz;
[0229] Electrical system data (3): battery voltage / inner resistance (mΩ), motor winding temperature (℃), controller CAN signal (Baud); sampling frequency is 50-200Hz;
[0230] Mechanical motion data (5): gearbox vibration frequency spectrum [displacement (mm), speed (mm / s), acceleration (mm / s 2 )], bearing acoustic signal (dB), brake pad wear thickness (mm); sampling frequency is 1-20kHz (vibration);
[0231] Environment perception data (5): ambient temperature (℃), ambient humidity (PPM), altitude pressure (hPa), road slope (%), GPS positioning (CEP); sampling frequency is 1-5Hz;
[0232] Driving behavior data (6): speed (m / s), acceleration G value (m / s 2 ), steering angular velocity (rad / s), frequency of emergency braking (t / h), driving mileage (km), tire pressure value (bar); sampling frequency is 10-50Hz;
[0233] Upper system data (13): low-pressure water pump injection amount (L / s), high-pressure water pump injection amount (L / s), cooling fan speed (RPM), brush motor speed (RPM), fan motor speed (RPM), fan cooling water pump speed (RPM), motor temperature (℃), fan temperature (℃), water tank liquid level value (cm), total water valve state value (KV), strut state value, tank return state value, rear door opening and closing state value; sampling frequency is 10-100Hz;
[0234] (2) TinyLSTM model hierarchical architecture
[0235] As shown in Figure 4 , TinyLSTM adopts a 4-layer heterogeneous architecture (TinyLSTM input layer, core LSTM layer, feature enhancement layer, TinyLSTM output layer), which balances the calculation efficiency and feature extraction capability through modular design. The description of each layer is shown in Table 1.
[0236] Table 1 TinyLSTM model hierarchical description table
[0237]
[0238] (3) Technical details and innovative design of each layer
[0239]
TinyLSTM input layer: sensor data preprocessing
[0240] Considering the differences in the working functions of new energy electric commercial vehicles, the amount of input data required for real-time fault diagnosis and prediction is mostly around dozens. The TinyLSTM input layer designed in the present application can receive up to 50-dimensional sensor data; the TinyLSTM input layer structure is shown in Figure 5 , and the TinyLSTM input layer realizes:
[0241] Dynamic normalization: for different sensor dimensions (such as temperature unit ℃, pressure unit kPa, acceleration unit m / s 2 , etc.), sliding window Z-Score standardization is adopted to calculate the mean μ and standard deviation σ in real time, formula: where window length 50 corresponds to 500 ms of data (10 ms sampling frequency) and ε represents the relative error;
[0242] Z-Score standardization, also known as standard score, standard deviation standardization, or Z-Transform, is a data normalization technique. It builds an abstract median based on certain quantiles of the original data, which is a value that is easier to use and understand than the original value. Z-Score standardization can generate data with the same scale for different variables, eliminating the difference in data magnitude between different variables, so as to better compare the differences between data.
[0243]
Core LSTM layer: double gate simplified structure
[0244] The core LSTM layer adopts a double gate simplified structure, as shown in Figure 6
[0245] Gate unit simplification: merge the input gate and forget gate of traditional LSTM into update gate, reduce 40% of parameter quantity, the calculation formula is as follows:
[0246] u t =σ(W u ·[h t-1 ,x t ]+b u )
[0247]
[0248] O t =σ(W o ·[h t-1 ,x t ]+b o )
[0249] h t =O t ⊙tanh(C t )
[0250] Where u t is the output of the update gate, σ is the sigmoid activation function, W u is the weight matrix of the update gate, h t-1 is the hidden state of the previous time step t-1, x t is the input at the current time t, and b u is the offset vector of the update gate; is the cell state of the candidate memory unit at the current time t, W C is the weight matrix of the candidate memory unit, and b C is the offset vector of the candidate memory cell; ⊙ denotes the element-wise multiplication (Hadamard product) for multiplication operation on corresponding position elements in matrix or vector; tanh denotes the hyperbolic tangent activation function; C t is the memory cell state of the current time step t, C t-1 is the memory cell state of the previous time step t-1; O t is the activation value of the output gate at time t, W o is the weight matrix of the output gate, b o is the offset vector of the output gate; h t is the hidden state at time t.
[0251]
Core LSTM layer: LSTM time step recursion based on two-stage dual gate reduction
[0252] As shown in Figure 7 , the time step recursion: the LSTM layer receives the input x t and the hidden state h t-1 at the previous time h t , updates the cell state C t , and outputs h t ;
[0253] (1) Initialize the 32-dimensional cell hidden state h(t) and the 32-dimensional candidate memory cell state c(t), and the time series length seq_len = 50;
[0254] (2) Input the Z-Score standardized 50-dimensional x_normalized(t), the initialized h(t) and c(t) to the one-stage dual gate reduction model DualGateLSTM. The h(t) and c(t) generated by the one-stage are input to the two-stage dual gate reduction model DualGateLSTM. The h(t) and c(t) generated by the two-stage are input to the dual gate reduction calculation process of the next time step t+1 together with the input x_normalized(t+1) of the next time step t+1, until the recursive calculation of 50 time steps is completed;
[0255] (3) Add the cell hidden state h(t) generated by the two-stage at each time step to the 32-dimensional array (lstm_outputs);
[0256] (4) After the array lstm_outputs summarizes all the hidden cell states h(t) of the 50 time steps, stack calculation is performed along the dimension 1, and finally the 32-dimensional lstm_out array is generated: lstm_out[seq_len = 50, batch_size = B, hidden_size = 32].
[0257] TinyLSTM feature enhancement layer: lightweight self-attention mechanism
[0258] The model structure of the lightweight self-attention mechanism of the TinyLSTM feature enhancement layer is as shown in Figure 8 The output vector of the parallel receiving core LSTM layer: lstm_out (32 dims) is calculated through the lightweight self-attention mechanism, and 8 feature vectors attn_out [seq_len = 50, batch_size = B, embed_dim = 32] are output;
[0259] Figure 8 The composition structure of the Scaled Dot-Product Attention model is as shown in Figure 9
[0260] Parameter compression technology: considering 8-bit fixed-point quantization (weight activation value using INT8 format, memory occupation reduced by 75%) and structured pruning (removing parameters with absolute value <0.01 in weight matrix, sparsity up to 60%);
[0261] Head dimension compression: the attention head dimension of the standard Transformer is 64, and the TinyLSTM compresses it to 8 dimensions, and the calculation amount is reduced to 1 / 8; wherein h = 8 represents the number of heads, and each head contains a separate scaled dot product attention;
[0262] Local attention window: only the attention weights of the 25 points before and after the current time step are calculated (not global), and the formula is as follows:
[0263]
[0264] Among them, Q t-25:t+25 is the information (i.e. search request) that needs to be paid attention to in the 25-point sequence before and after the current time step; K t-25:t+25 T represents the unique identifier of each token in the 25-point sequence before and after the current time step, which is used for similarity calculation with Query, and the role of K is to determine which information is most relevant to the current Query; V t-25:t+25 represents the actual content or feature of each token in the 25-point sequence before and after the current time step, which is used for weighted summation according to the similarity score to generate the final output; d k is the dimension of Key, which is used to scale the dot product to prevent gradient explosion or disappearance;
[0265] Linear transformation: Q, K, and V are obtained by linear transformation of the input matrix, as follows:
[0266] Q t-25:t+25 = XWt-25:t+25 Q , K t-25:t+25 = XW t-25:t+25 K , V t-25:t+25 = XW t-25:t+25 V
[0267] Similarity calculation: Similarity matrix is obtained by calculating the dot product of Q t-25:t+25 and K t-25:t+25 T (K t-25:t+25 transpose) To stabilize the training process, the dot product result is divided by
[0268] Weighted sum: The similarity matrix is normalized by Softmax to obtain the attention weight, and then multiplied by V t-25:t+25 to obtain the weighted output;
[0269] Hardware adaptation optimization: Use lookup table (LUT) to accelerate Softmax calculation to avoid exponential operation consuming MCU resources;
[0270] The basic principle of LUT is to precompute and store function values in an array or dictionary, and then quickly obtain the approximate value by looking up the table when calculating the function value. The Softmax calculation formula is as follows, and the calculation bottleneck is the exponential operation, which can be optimized by LUT lookup table method. The specific optimization methods include exponential operation replacement (convert floating point exponential calculation to fixed point lookup table), quantization and normalization (reduce dynamic range and reduce LUT size), parallel lookup and pipeline (hide delay through hardware parallelism);
[0271]
[0272] For a quantized model, the input of the model is generally an 8 / 16 / 64-bit data type original data data, a float type scale, and an int8 type zero_point. Taking int8 data type as an example, it can represent a range from -128 to 127, a total of 256 different data, and the float type data after dequantization is also 256. For small batch limited data, lookup table optimization is more appropriate;
[0273] The output of the model is an int8 type quantized value, but the output of the Softmax function is a float type dequantized value, so to realize the dequantization function, the dequantization operation is as shown in Figure 10
[0274] Implementation process of DequantizeTo32FP function:
[0275] (1) Convert the input T type data and int8 type data zero_point to float data type and calculate the difference between data and zero_point;
[0276] (2) Multiply the above difference by the scale value of float data type to complete the inverse quantization calculation and generate an output result of float type.
[0277] Then create a data table. The index of the table is int8 type data. The value of the table is the value after dequantization of the int8 type data and then exponential calculation:
[0278] The exponential operation table is generated as follows Figure 11 As shown:
[0279] (1) In the int8 data type range [-128, 127], the number size = 256, perform a for loop calculation;
[0280] (2) In each step of the for loop, the input key is dequantized by DequantizeTo32FP to generate the result deq;
[0281] (3) Using the result of the previous step as the power exponent, calculate the deq power of e and generate the result value;
[0282] (4) Value is written into the map exponential operation table exp_table with key as the key until the for loop ends.
[0283] Finally, replace all the exponential function operations in the Softmax function with exponential tables:
[0284] The Softmax function is simulated by the exponential operation table, and the output probability value is calculated as follows Figure 12 As shown:
[0285] (1) In the int8 data type range [-128, 127], the number size = 256, perform the first stage for loop calculation;
[0286] (2) In each step of the first stage for loop, according to the input key = input[i], the corresponding value is obtained from the exponential operation table exp_table;
[0287] (3) Assign value to the output array output[i];
[0288] (4) After the first stage for loop ends, sum all the values in the output array to generate a float type sum value sum;
[0289] (5) Enter the second stage for loop, and perform loop calculation according to size = 256;
[0290] (6) In each step of the second stage for loop, divide each value output[i] in the output array by the sum value sum, and write the generated result back to the array node output[i] until the second stage for loop ends.
[0291] The overall design principle is: pre-write the results of complex exponential operations into the exponential operation table; When needed later, directly obtain the pre-calculated results from the exponential operation table. Through this method, the exponential operation table simulates the Softmax function, calculates the output probability value with high precision, and avoids real-time complex exponential calculation when executing the Softmax function.
[0292]
TinyLSTM feature enhancement layer: cross-layer residual connection
[0293] Cross-layer residual connection: the output h of the LSTM layer t is directly connected to the output layer and spliced with the attention layer output to avoid gradient disappearance and retain the original time sequence features. The calculation formula is:
[0294] y t = f output ([h t ; Attention(h t-25:t+25 )])
[0295] The cross-layer residual connection is as shown in Figure 13 :
[0296] (1) The output results lstm_out[seq_len = 50, batch_size = B, hidden_size = 32] based on the two-stage double-door simplified LSTM time step recursion and the output results attn_out[seq_len = 50, batch_size = B, embed_dim = 32] of the lightweight self-attention mechanism are taken as inputs of the cross-layer residual connection layer;
[0297] (2) Perform feature fusion connection cat calculation on the dim = 2 dimension;
[0298] (3) Output the calculation result of the cross-layer residual connection:
[0299] combined[seq_len=50, batch_size=B, hidden_size=64].
[0300]
TinyLSTM output layer: fully connected network
[0301] For vehicle fault prediction, the width of the output layer of the fully connected network is set to a 64-16-1 unit structure, using a hybrid activation function, outputting an abnormal probability of 0-1. The fully connected network structure for vehicle fault prediction is as shown in Figure 14
[0302] (1) The input data of this layer is the output result combined vector of the previous step cross-layer residual connection;
[0303] (2) The first Linear Layer linear layer linearly transforms the input combined vector, reducing the feature dimension from 64 to 16;
[0304] (3) The first layer uses the activation function Leaky ReLU, with a negative input slope of a = 0.1 to alleviate gradient disappearance;
[0305] (4) The last Linear Layer linear layer linearly transforms the vector calculated by the activation function Leaky ReLU, reducing the feature dimension from 16 to 1;
[0306] (5) The last layer uses Sigmoid to constrain the output range, suitable for probability interpretation requirements, generating vehicle fault diagnosis prediction results.
[0307] TinyLSTM model training, validation evaluation, model quantization deployment and testing, performance verification, as follows:
[0308]
Setting training loss function and selecting optimizer
[0309] (1) The training process of the TinyLSTM model uses the cross-entropy loss function, which has the advantages of efficient gradient calculation, avoiding gradient disappearance, and fast convergence of the model, suitable for probability distribution matching tasks, especially in classification problems. Its mathematical properties (such as convexity and derivability) and probability interpretation (such as the relationship with KL divergence and maximum likelihood estimation) further enhance its practicality;
[0310] (2) The optimizer selects the stochastic gradient descent method, with a learning rate (learning rate) set to lr = 1e-3, using Nesterov momentum;
[0311]
TinyLSTM model training
[0312] (1) In a main for loop, use the loaded dataset to train the model;
[0313] (2) Calculate the cross-entropy loss between the model's predicted results and the actual label distribution to see how different they are;
[0314] (3) Perform backpropagation to calculate the gradient of the loss function with respect to the model parameters;
[0315] (4) Set the maximum threshold of the gradient norm (default L2 norm) to 1.0, and perform gradient clipping and parameter constraint on the original model tensor;
[0316] (5) Embed a small for loop to judge the dimensions of all parameters of the model. If the dimension of any parameter is >1, limit the value of the model parameter data to [-2, 2];
[0317] (6) Update the model parameters according to the current calculated gradient value;
[0318] (7) Calculate the cumulative loss value;
[0319] (8) When the main for loop ends, return the average loss value of the training.
[0320]
TinyLSTM model validation and evaluation
[0321] (1) The preloaded dataset includes the input dataset X and the predicted result true value dataset y required for model validation and evaluation;
[0322] (2) Temporarily disable the automatic differentiation function, i.e. disable gradient calculation and history tracking torch.no_grad(), all operations involving tensors with requires_grad=True will not generate a computation graph or store gradients, directly skipping the intermediate state recording required for backpropagation, thus saving memory and accelerating computation during inference or validation phase;
[0323] (3) Use a for loop, the model performs real-time inference and prediction on the input dataset X, and the predicted results pred and true value data y are stored in arrays y_pred and y_true respectively;
[0324] (4) After the for loop ends, convert the arrays y_pred and y_true to NumPy arrays;
[0325] (5) Input the true value in the NumPy array y_true and the predicted value y_pred>0.5 into the f1_score function to calculate the evaluation index of the classification model performance, the f1 score takes into account the accuracy and recall rate;
[0326] (6) The true value in NumPy array y_true and the predicted value in y_pred are input into the roc_auc_score function to calculate the area fraction under the ROC receiver operating characteristic curve to measure the prediction performance of the binary classification model; ROC AUC is a threshold-independent indicator, with a value range of [0.5, 1], and the larger the value, the better the model performance;
[0327] (7) Return the f1_score calculation score and the roc_auc_score calculation score to comprehensively verify and evaluate the prediction performance of the model.
[0328]
TinyLSTM model quantization deployment and testing
[0329] (1) Save the quantized model quantized_model to adapt to the vehicle end deployment, and use the torch.quantization.quantize_dynamic function to realize dynamic quantization of the {nn.Linear} linear fully connected layer. Its function is to convert the model weight from a floating-point number to a low-precision integer (usually INT8) method to reduce the model size and speed up the inference speed. This function is particularly suitable for CPU devices, and only quantizes the weight, while the activation remains FP32 precision, which helps to reduce computational complexity and memory usage while maintaining high model accuracy;
[0330] (2) Convert the quantized model into a runnable script using the torch.jit.script function. The advantage of using this function is that it can reduce the overhead during model execution, as it eliminates the overhead of the Python interpreter;
[0331] (3) Use the torch.jit.save function to save the model converted to a script as a cross-platform format tinylstm_quantized.pt. When saving the model using torch.jit.save, it can reduce the overhead during model execution, as it eliminates the overhead of the Python interpreter. The saved model can be executed in an environment without Python dependencies, making the deployment of the model more flexible and efficient;
[0332] (4) Real-time inference testing: unsqueeze the test dataset as sample_input; then temporarily disable the automatic differentiation function, i.e., disable gradient calculation and history tracking torch.no_grad();
[0333] (5) quantized_model(sample_input).item() represents that the input sample sample_input is inferred in real time by the quantized model quantized_model to obtain the prediction result; the item function can be used to obtain the prediction result of a single element;
[0334] (6) Print the prediction result. If the vehicle fault diagnosis prediction result probability > 0.5, it means that the vehicle fault is an error, otherwise it means normal.
[0335]
Performance verification
[0336] The TinyLSTM model deployed on the vehicle side after quantization has a parameter size of 412KB, the cache required for real-time inference of the model is 89KB, the single inference delay is 3.7ms, the model update DOTA difference parameter upgrade package is only 28KB, and the vehicle side deployment can be completed within 5 seconds. The model fault real-time detection coverage rate is exemplified as follows: engine misfire (98.2%), battery over-temperature (95.6%), and brake pad wear (93.1%).
[0337] Although the present application is described herein with reference to the explanatory embodiments of the present application, the above-described embodiments are only the preferred embodiments of the present application, and the embodiments of the present application are not limited to the above-described embodiments. It should be understood that those skilled in the art can design many other modifications and embodiments, which will fall within the scope and spirit of the principles disclosed in the present application.
Claims
1. A vehicle fault intelligent diagnosis prediction method, characterized in that, The method comprises the following steps: Step S10, training a TinyLSTM model based on horizontal federated learning HFL of time sequence streams of sensors of multiple vehicle enterprises in the cloud; Step S20, deploying the trained TinyLSTM model to a vehicle end; Step S30, the TinyLSTM model of the vehicle end receiving a time sequence stream of vehicle sensors and outputting intelligent diagnosis and prediction results of vehicle faults.
2. The method of claim 1, wherein, The step S10 specifically comprises: Step S11, deploying a central parameter server cluster and a client cluster hardware environment in the cloud, each vehicle enterprise constructing a respective client cluster, and collecting and reporting real-time sensor time sequence data of respective target vehicles to the cloud; Step S12, the central parameter server cluster downlinking initial global model parameters to each client cluster, the client cluster training a TinyLSTM model using its own local vehicle fault data, each round of training result being uploaded to the central parameter server cluster in an encrypted manner, and the central parameter server cluster downlinking global model parameters calculated through federated average gradient aggregation after each round of iterative training to each client cluster; repeating multiple rounds of iterative training until the model training is completed.
3. The method of claim 2, wherein, The step S12 specifically comprises: Step 1), initializing the global model parameters by the central parameter server; Step 2), the central parameter server randomly selecting multiple clients and downlinking the current global model parameters to the selected clients in a plaintext manner; Step 3), determining whether the training reaches a preset number of iterations or whether the model convergence performance meets the requirements, if yes, ending the training process; otherwise, proceeding to the next step; Step 4), the client training the model using its own local vehicle fault data; Step 5), the client updating its own model weight value by using Nesterov accelerated SGD and calculating the gradient; Step 6), the client normalizing the calculated gradient; Step 7), the client adding a random mask to the normalized gradient; Step 8), the client using a Paillier encryption scheme and using a public key PK generated by the central parameter server to homomorphically encrypt the updated model weight value; Step 9), the client sending the encrypted model weight value to the central parameter server; Step 10), the central parameter server receiving the ciphertext sent by all clients; Step 11), the central parameter server using a Shapley algorithm to monitor the contribution of each client node to the gradient prediction result in real time and automatically isolating the abnormal participants; Step 12), the central parameter server performing a Paillier homomorphic addition operation to generate aggregated ciphertext; Step 13), the central parameter server decrypting the aggregated ciphertext using a private key SK generated by itself to obtain an aggregated gradient, i.e., the current global model parameters; Step 14), returning to step 2) for cyclic iteration until the model training end condition is reached.
4. The method of claim 1, wherein, The TinyLSTM model comprises a TinyLSTM input layer, a core LSTM layer, a TinyLSTM feature enhancement layer and a TinyLSTM output layer, wherein: The TinyLSTM input layer is used for receiving sensor data of a vehicle and performing sensor data preprocessing before inputting into the core LSTM layer; The core LSTM layer combines the input gate and the forget gate of the traditional LSTM into an update gate, realizes LSTM time step recursion, and is used for extracting short-term / long-term time sequence dependence in the input data; The TinyLSTM feature enhancement layer includes a lightweight self-attention mechanism, is used for receiving output vectors of the core LSTM layer in parallel, outputs 8 feature vectors after calculation of the lightweight self-attention mechanism, and performs inverse quantization operation; the output result of the core LSTM layer and the output result of the lightweight self-attention mechanism are used as inputs of a cross-layer residual connection layer, the cross-layer residual connection layer performs feature fusion connection calculation, and outputs a calculation result of cross-layer residual connection; The TinyLSTM output layer adopts a fully connected network, the input is the calculation result combined vector of the cross-layer residual connection, the first linear layer performs linear transformation on the input combined vector, and the feature dimension is reduced from 64 to 16; the first linear layer uses an activation function Leaky ReLU, the slope α of negative input is 0.1, so as to relieve gradient disappearance; the last linear layer performs linear transformation on the vector calculated by the activation function Leaky ReLU again, and the feature dimension is reduced from 16 to 1; the last layer uses a Sigmoid function to constrain the output range, adapts to the demand of probability explanation, and generates a vehicle fault diagnosis prediction result probability.
5. The method of claim 4, wherein, The calculation formula of the core LSTM layer is as follows: u t = σ(W u · [h t-1 , x t ]+ b u ); O t = σ(W o · [h t-1 , x t ]+ b o ); h t =O t ⊙tanh(C t ); where u t is the update gate output, σ is the sigmoid activation function, W u is the weight matrix of the update gate, h t-1 is the hidden state of the previous time step t-1, x t is the input at the current time step t, b u is the bias vector of the update gate; is the candidate memory cell state at the current time step t, W C is the weight matrix of the candidate memory cell, b C is the bias vector of the candidate memory cell; ⊙ represents the element-wise multiplication Hadamard product, which is used to multiply the corresponding position elements in the matrix or vector; tanh represents the hyperbolic tangent activation function; C t is the memory cell state at the current time step t, C t-1 is the memory cell state at the previous time step t-1; O t is the activation value of the output gate at time t, W o is the weight matrix of the output gate, b o is the bias vector of the output gate; h t is the hidden state at time t; the core LSTM layer receives the input x t and the previous time step hidden state h t-1 at each time step, updates the cell state C t and outputs h t , realizing time step recursion.
6. The method of claim 4, wherein the method further comprises: The training of the TinyLSTM model adopts a cross-entropy loss function and a selected stochastic gradient descent method as an optimizer, and specifically includes the following steps: A1, in a main loop, using the loaded data set, training the model; A2, calculating the cross-entropy loss of the prediction result of the model and the actual label distribution; A3, performing back propagation to calculate the gradient of the loss function on the model parameters; A4, setting the maximum threshold 1.0 of the gradient norm, performing gradient clipping and parameter constraint on the original model tensor; A5, embedding a small loop, judging the dimension of all parameters of the model, if the dimension of any parameter is greater than 1, limiting the value of the model parameter data to [-2, 2]; A6, updating the model parameters according to the current calculated gradient value; A7, calculating the value of the accumulated cross-entropy loss; A8, when the main loop ends, returning the value of the average cross-entropy loss of the training.
7. The method of claim 4, wherein the method further comprises: The verification and evaluation method of the TinyLSTM model is as follows: B1, the preloaded data set includes an input data set X and a prediction result true value data set y required for model verification and evaluation; B2, temporarily disabling the automatic derivation function, that is, disabling gradient calculation and historical tracking torch.no_grad(), all operations involving tensors with requires_grad=True will not generate a calculation graph or store gradients, and directly skip the intermediate state recording required for back propagation; B3, using a for loop, the model makes real-time inference prediction on the input dataset X, the prediction result pred and the true value data y are stored in the arrays y_pred and y_true respectively; B4, after the for loop ends, the arrays y_pred and y_true are converted into NumPy arrays respectively; B5, the true value in the NumPy array y_true and the predicted value y_pred>0.5 are input into the f1_score function to calculate the index of evaluating the performance of the classification model, the f1 score considers the performance of accuracy and recall; B6, the true value in the NumPy array y_true and the predicted value in y_pred are input into the roc_auc_score function to calculate the area score under the ROC receiver operating characteristic curve to measure the prediction performance of the binary classification model; ROC AUC is a threshold-independent index, its value range is [0.5, 1], the larger the value, the better the model performance; B7, the f1_score calculation score and roc_auc_score calculation score are returned to comprehensively verify and evaluate the prediction performance of the model.
8. The method of claim 1, wherein, The method for quantization deployment of the TinyLSTM model is: C1, using the torch.quantization.quantize_dynamic function to realize dynamic quantization of the {nn.Linear} linear fully connected layer, and saving the quantized model quantized_model to adapt to the vehicle end deployment; C2, using the torch.jit.script function to convert the quantized model into a runnable script; C3, using the torch.jit.save function to save the model converted into a script as a cross-platform format tinylstm_quantized.pt.
9. The method of claim 8, wherein, The method for testing the TinyLSTM model is: D1, real-time inference test: decompress the test dataset as sample_input; then temporarily disable the automatic differentiation function, that is, disable gradient calculation and history tracking torch.no_grad(); D2, input the sample_input through the quantized model quantized_model to make real-time inference and get the prediction result; D3, if the vehicle fault diagnosis prediction result probability>0.5 indicates that the vehicle fault is wrong, otherwise it indicates normal.
10. The method of claim 1, wherein, The step S20 specifically comprises: after the client cluster completes the TinyLSTM model training, the TinyLSTM model is distributed to the target vehicle for writing upgrade through the DOTA method.
Citation Information
Cited By
Vehicle fault prediction method based on AUTOSAR
CN121560006A