Intelligent fault diagnosis method for complex equipment under digital multi-twinning assistance
Through the intelligent fault diagnosis method of complex equipment assisted by digital multi-twin, the problem that SoC estimation relies on the same distribution data and gradient conflicts in fault diagnosis of battery energy storage system is solved, and high-accurate fault diagnosis and traceability are achieved, improving the fault diagnosis capability of battery energy storage system.
Patent Information
- Application Number
- CN202510271873.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-01
AI Technical Summary
The prior art has the problem that SoC estimation depends on the same distribution of training and test data in the fault diagnosis of battery energy storage systems, and the fault diagnosis method based on deep learning consumes a lot of computing resources and cannot optimize gradient conflicts.
Using a complex equipment intelligent fault diagnosis method with digital multi-twin assist, data is processed through sliding window technology, a backbone network of multi-scale layers and feature fusion layers is built, task correlation is defined using cosine similarity, cosine annealing algorithm is used to adjust the learning rate, and a fault simulation and feature library based on digital twin models is built to realize real-time fault diagnosis and traceability.
It improves the accuracy of battery energy storage system fault diagnosis, from 55% to more than 92%, can accurately identify various complex faults such as battery overcharge, overdischarge, thermal runaway, and optimizes gradient conflict problems in multi-task learning.
Smart Images

Figure CN120233232A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of battery energy storage system fault diagnosis, and specifically to an intelligent fault diagnosis method for complex equipment assisted by digital multi-twins. Background Technique
[0002] As a key component of a battery energy storage system (BESS), a battery management system (BMS) aims to protect the battery from harmful and inefficient operations. In the BMS, state of charge (SoC) estimation and battery fault diagnosis are two of the most important functions. The SoC, defined as the ratio of available capacity to maximum capacity, is a fundamental state to ensure normal battery management. Battery fault diagnosis is very important for preventing battery faults from rapidly deteriorating from minor faults to out-of-control faults.
[0003] Normally, SoC estimation and battery fault diagnosis are studied separately. Deep learning-based SoC estimation methods can directly map sampled battery operating signals (such as voltage and current) to the SoC, thus eliminating the need for battery modeling or feature engineering. To address the problem that deep learning in estimating the SoC depends on training data and test data with the same distribution, Bian et al. proposed a deep transfer neural network based on multi-scale distribution adaptation (MDA) to generalize the domain adaptation ability of the deep estimator. Jiang et al. combined the mechanism of the battery with a long short-term memory (LSTM) network to realize the correlation between external measurements and the internal state of the battery, and thus adapted the SoC estimation under different training and operating conditions. Zhan et al. proposed a two-channel deep learning method, which has higher computational efficiency and smaller memory requirements compared with traditional single-channel methods.
[0004] Extensive research has been conducted on deep learning-based battery fault diagnosis methods. Zhang et al. proposed a method based on the Lebesgue time model to address the problem that traditional fault diagnosis methods based on LSTM under the Riemann sampling framework have high computational and training requirements. Compared with traditional fault state space models, it has lower complexity. Song et al. embedded the equivalent circuit model of the battery into the neural network structure to learn the voltage fault observer. This method not only has the interpretability of the physical model but also endows the model with the powerful nonlinear processing ability of the neural network. Liu et al. proposed an intelligent diagnosis method based on a feature-enhanced random configuration network and adversarial domain expansion for unbalanced battery fault data to address the problem of insufficient actual fault data in battery operation. This method balances the distribution of the sample domain, thus reducing model bias.
[0005] Gradient-based multi-objective optimization methods can be roughly divided into two categories: objective scalarization methods and adaptive gradient methods. The most commonly used scalarization method is simple linear scalarization, as shown in formula (1). However, due to the conflicting gradient vectors, this method may lead to mutual inhibition between tasks. In addition, scalarization methods do not guarantee the discovery of all Pareto solutions.
[0006]
[0007] Recognizing the limitations of scalarization methods, many adaptive gradient techniques have emerged. These methods aim to dynamically balance different tasks by adjusting the gradients during the optimization process. Among them, GradNorm is a representative method. Similar to BatchNorm, GradNorm calculates the gradient norm G W (t) for each task and uses the average gradient norm as the baseline for each training step t. Then, the gradient norm is normalized with respect to where θ represents the shared parameters of all tasks.
[0008] However, GradNorm has significant limitations. It requires calculating the gradients of each task with respect to the network's shared layers and storing the computational graph, which may consume a large amount of computational resources. In addition, GradNorm cannot optimize the gradient conflict problem existing between two tasks. To solve this problem, PCGrad was introduced. PCGrad uses the cosine similarity Φ(g i ,g j ) to quantify the similarity between task gradients. When a conflict is detected, this method projects the gradient of one task onto the normal plane of the conflicting gradient of another task. Although PCGrad successfully solves the conflict, it does not bring additional benefits to tasks without conflicts. Summary of the Invention
[0009] The purpose of the present invention is to provide a method for intelligent fault diagnosis of complex equipment assisted by digital multi-twins to solve the problems presented in the above background technology.
[0010] To achieve the above objective, the present invention provides the following technical solution. A method for intelligent fault diagnosis of complex equipment assisted by digital multi-twins includes the following steps:
[0011] S1. Data preprocessing: The original data is processed using the sliding window technique with a window size of 2048;
[0012] S2. Model construction: The backbone network of the model includes a multi-scale layer and a feature fusion layer;
[0013] S3. Adaptive gradient optimization: Use the cosine similarity φ(gi , g j ) = cos(g i , g j ) to define the correlation between two tasks;
[0014] S4, Adaptive learning rate: To ensure the convergence of the model, it is necessary to adaptively adjust the learning rate during the model training process. The cosine annealing algorithm is considered to gradually reduce the value of the learning rate η;
[0015] S5, Joint loss function: The scalarization method is to achieve parallel optimization in the multi-task learning framework;
[0016] S6, Fault simulation and feature library construction based on the digital twin model: In the digital twin model, simulate various possible fault scenarios, obtain the model output data under the corresponding fault scenarios, perform the same data preprocessing operations as in S1 on the output data, extract fault features, and construct a fault feature library. The fault feature library contains feature vectors corresponding to different fault types;
[0017] S7, Real-time fault diagnosis: Input the preprocessed real-time collected data into the trained fault diagnosis model. The fault diagnosis model determines whether the device has a fault and the type of the fault by comparing the real-time data features with the feature vectors in the fault feature library; if it is determined that the device has a fault, further use the digital twin model for fault tracing to determine the specific location and cause of the fault;
[0018] S8, Diagnosis result feedback and model update: Feedback the fault diagnosis result to the device operator and maintenance personnel, and update the digital twin model and the fault diagnosis model according to the actual maintenance situation and newly obtained data to improve the accuracy and adaptability of the diagnosis model.
[0019] Optionally, S1 includes: setting the first half of the data to 40 steps and the second half to 10 steps, so there are a total of 83602 samples. In addition, divide the data set into a training set, a test set, and a validation set according to a ratio of 7:2:1.
[0020] Optionally, the multi-scale layer includes: a multi-scale feature fusion block and a multi-scale channel fusion block.
[0021] Optionally, S2 includes: adding a branch from the i-th layer to the i + 1-th layer, and the sizes of the convolutional kernels on the branch are 23, 9, and 3 in sequence. Taking the second multi-scale layer to the third multi-scale layer as an example, let {R r , r = 1, 2} represent the output representation of the second layer, where each R r corresponds to the feature map of a specific scale r. In the third layer, the output representation is extended to {R' r, r = 1, 2, 3}, capturing information at more scales, and each output of the third layer represents R' r All are calculated through the following operations:
[0022] R' r = MaxPool(Conv r (Concat(R1, R2)))
[0023] where Concat(·) is the concatenation operation, Conv r (·) represents a one-dimensional convolution operation at scale r, and MaxPool(·) is a one-dimensional max pooling operation.
[0024] Optionally, the S2 includes: The last layer of the backbone network is a multi-feature fusion layer, that is, a linear transformation is performed on the extracted multi-scale features to generate a shared representation:
[0025]
[0026] where W is a learnable weight matrix, BN represents the batch normalization (BN) operation, b is a bias vector, R sh represents the finally integrated multi-scale feature representation, and σ represents the ReLU activation function.
[0027] Optionally, the S3 includes: where is the gradient of task i; for simplicity, φ(g i , g j ) is abbreviated as φ ij , when φ ij < 0, it is called that a conflict occurs; when a conflict appears, g i is replaced with the canonical value of g j when the conflict occurs:
[0028]
[0029] In order to enjoy the benefits brought by the algorithm under the condition that the gradients between tasks do not conflict, the gradients are linearly combined: g' i = αg i + βg j , where α and β are positive constants; the combinations of vectors are infinite; for convenience, α = 1 is set, and β is solved:
[0030]
[0031] where φ ij is the target direction of φ ij , and double exponential moving smoothing is used to calculate φ ij, The main advantage of DEMA is that it reduces the lag effect caused by traditional multi-EMA filtering, making the smoothing result closer to the changes in the original data. Compared with single-order or second-order EMA, DEMA can respond faster to short-term changes in data;
[0032]
[0033] where γ is the smoothing factor, 0 ≤ γ ≤ 1, k is the number of layers, and are the EMA of the data at step t and the EMA of the EMA respectively;
[0034] Let the gradients of task 1 and task 2 be and Let g = g1 + g2, then the deep learning model parameters are updated as follows:
[0035] θ + = θ - η·g′
[0036] where θ + is the model parameter after update, g′ is the target gradient; in addition, the smoothing factor γ is set to 0.5.
[0037] Optionally, S4 includes: The formula of the cosine annealing algorithm is as follows:
[0038]
[0039] where η max is the initial value of the learning rate, N cur is the current batch number used for training, N cur is the total number of training batches; under the action of the cosine function, the periodic decrease of the learning rate can enable the optimization process to gradually explore finer regions of the loss surface.
[0040] Optionally, S5 includes: For FC tasks and FL tasks, since the cross-entropy loss function is applicable to classification prediction tasks and can ensure robust learning of discrete class probabilities, the cross-entropy loss function is adopted.
[0041] Optionally, in the step of fault simulation and feature library construction based on the digital twin model, the fault simulation is achieved by perturbing the key parameters of the device in the digital twin model or modifying the structure of the model; the fault feature library is stored and managed using a database management system and has functions of fast query and update.
[0042] Optionally, in the real-time fault diagnosis step, the fault diagnosis model adopts deep learning algorithms such as convolutional neural network (CNN), recurrent neural network (RNN) or long short-term memory network (LSTM). When training the fault diagnosis model, the cross-validation method is used to improve the generalization ability of the model. The fault tracing is realized by analyzing the state changes and signal propagation paths before and after the fault occurs in the digital twin model.
[0043] Optionally, in the diagnosis result feedback and model update step, the diagnosis results are presented in a visual way, including the fault type, fault location, fault severity of the device and repair suggestions. The model update adopts the incremental learning method, gradually integrating the newly obtained data into the original model for training to reduce the time and computational resource consumption of model training.
[0044] Compared with the prior art, the present invention provides a complex device intelligent fault diagnosis method assisted by digital multi-twins, having the following beneficial effects:
[0045] 1. For the complex device intelligent fault diagnosis method assisted by digital multi-twins, the digital twin model constructed for the battery energy storage system comprehensively covers the electrochemical model, thermal management model and system circuit topology model of the battery pack. Taking a large-scale battery energy storage power station as an example, high-precision battery parameter measurement equipment is used to obtain characteristic data such as the capacity and internal resistance of battery monomers. Combining with the electrochemical principle, an accurate electrochemical model is constructed to accurately simulate the electrochemical reactions during the charging and discharging processes of the battery. The thermal management model is established based on the heat dissipation efficiency of the cooling fan, the flow characteristics of the coolant and the heat conduction law of the battery module, and can reflect the battery temperature changes in real time. The system circuit topology model is built according to the electrical connection relationship of the energy storage system. When simulating the battery aging fault, by adjusting the aging parameters of battery monomers in the electrochemical model, such as increasing the internal resistance and reducing the capacity, and combining the output data of the fault scenarios obtained from the thermal management model and the circuit topology model, it highly fits the actual fault state. Through actual tests, compared with traditional fault diagnosis means, the fault diagnosis accuracy based on the digital twin model assistance has been greatly improved from 55% to over 92%, and it can accurately identify various complex faults such as overcharging, over-discharging and thermal runaway of the battery.
[0046] 2. The complex device intelligent fault diagnosis method assisted by digital multi-twins effectively extracts the features of data in the time scale and space scale by integrating multi-scale information, and realizes the joint learning of the fault diagnosis and SoC estimation tasks of the battery energy storage system. On the basis of the traditional multi-task learning paradigm, an adaptive gradient optimization algorithm is introduced, which can effectively alleviate the gradient conflict problem in multi-task learning and optimize the performance of the multi-task model. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 Schematic diagram of the model mechanism of the present invention;
[0048] Figure 2 Schematic diagram of the experimental platform of the battery energy storage system of the present invention;
[0049] Figure 3 Schematic diagram of voltage and temperature data of the battery energy storage system of the present invention;
[0050] Figure 4 Schematic diagram of fault data of current and temperature sensors of the present invention;
[0051] Figure 5 Schematic diagram of the change of loss weight of the Uncertainty method of the present invention with the training process;
[0052] Figure 6 Schematic diagram of the change of loss weight of the DWA method of the present invention with the training process;
[0053] Figure 7 Schematic diagram of the SoC estimation result of the present invention. Detailed implementation manners
[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0055] As Figures 1-7 shown, the present invention provides a technical solution: a complex equipment intelligent fault diagnosis method assisted by digital multi-twin, including the following steps:
[0056] S1. Data preprocessing:
[0057] The original data is processed using the sliding window technique with a window size of 2048. The first half of the data is set to 40 steps and the second half is set to 10 steps, so there are a total of 83,602 samples. The reason for this is to consider the changes in charging and discharging. In addition, the voltage data is normalized, as Figure 3 shown. Subsequently, the data set is divided into a training set, a test set, and a validation set according to a ratio of 7:2:1. In addition, the initial value of the learning rate η is 0.01.
[0058] S2. Model construction:
[0059] The backbone network of the model of the present invention includes a multi-scale layer and a feature fusion layer. The multi-scale layer includes: a multi-scale feature fusion block and a multi-scale channel fusion block.
[0060] As Figure 1 , from the i-th layer to the (i + 1)-th layer, a branch is added, and the convolutional kernel sizes on the branch are 23, 9, and 3 in sequence. Taking the second multi-scale layer to the third multi-scale layer as an example, let {R r , r = 1, 2} represent the output representation of the second layer, where each R r corresponds to the feature map of a specific scale r. At the third layer, the output representation is extended to {R′ r , r = 1, 2, 3} to capture information of more scales. Each output representation R′ r at the third layer is calculated through the following operations:
[0061] R′ r = MaxPool(Conv r (Concat(R1, R2))) (12)
[0062] where Concat(·) is the concatenation operation, Conv r (·) represents the one-dimensional convolutional operation of scale r, and MaxPool(·) is the one-dimensional max pooling operation. The above process can be extended to deeper layers.
[0063] The last layer of the backbone network is the multi-feature fusion layer, that is, a linear transformation is performed on the extracted multi-scale features to generate a shared representation:
[0064] R sh = σ(BN(Concat(R1′, R2′, R3′)W + b)) (13)
[0065] where W is the learnable weight matrix, BN represents the batch normalization (BN) operation, b is the bias vector, and R sh represents the finally integrated multi-scale feature representation, and σ represents the ReLU activation function.
[0066] The core of the task representation head of the present invention is a linear transformation layer, which maps the shared feature space to the output space of a specific task. Mathematically, let R sh ∈ R d represent the shared feature representation extracted from the backbone network. The linear transformation of the i-th task is expressed as:
[0067] y i = W i R sh + b i (14)
[0068] where, represents the output of the i-th task, is the learnable weight matrix, is the bias vector.
[0069] S3. Adaptive gradient optimization:
[0070] Use the cosine similarity φ(g i , g j ) = cos(g i , g j ) to define the correlation between two tasks, where is the gradient of task i. For simplicity, abbreviate φ(g i , g j ) as φ ij . When φ ij < 0, it is called that a conflict occurs. When a conflict appears, replace g i with the normalized value of g j when the conflict occurs:
[0071]
[0072] To enjoy the benefits brought by the algorithm under the condition that the gradients between tasks do not conflict, it is natural to linearly combine the gradients: g
[0073] = αg i + βg i + βg j , where α and β are positive constants. However, the combinations of these vectors are infinite. For convenience, set α = 1 and solve for β:
[0074]
[0075] where φ ij is the target direction of φ ij . Use Double Exponential
[0076] Moving Average (DEMA) to calculate φ ij . The main advantage of DEMA is to reduce the lag effect caused by traditional multi-EMA filtering, making the smoothed result closer to the change of the original data. Compared with single-order or second-order EMA, DEMA can respond faster to the short-term changes of data.
[0077]
[0078]
[0079] where γ is the smoothing factor, 0 ≤ γ ≤ 1. k is the number of layers, and are the EMA of the data at step t and the EMA of the EMA respectively;
[0080] Let the gradients of Task 1 and Task 2 be and Let g = g1 + g2. Then the deep learning model parameter update is:
[0081] θ + = θ - η·g′ (14)
[0082] where θ + is the model parameter after update, and g′ is the target gradient. The target gradient is determined according to formulas (6)-(9). In addition, the smoothing factor γ is set to 0.5.
[0083] S4. Adaptive learning rate:
[0084] To ensure the convergence of the model, it is necessary to adaptively adjust the learning rate during the model training process. The cosine annealing algorithm is considered to gradually reduce the value of the learning rate η. The formula of the cosine annealing algorithm is as follows:
[0085]
[0086] where η max is the initial value of the learning rate, N cur is the current number of batches used for training, and N cur is the total number of training batches. Under the action of the cosine function, the periodic decrease of the learning rate can enable the optimization process to gradually explore finer regions of the loss surface. This method provides a smooth non-linear decay and avoids sudden changes in the learning rate, so as not to disrupt the stability of training or hinder convergence.
[0087] S5. Joint loss function: The scalarization method is a widely recognized method for implementing parallel optimization in the multi-task learning framework. According to formula (1), the joint loss function can be written as:
[0088]
[0089] To achieve the best performance in these tasks, task-specific loss functions are adopted and integrated into a unified joint loss function. Specifically, for the FC task and the FL task, since the cross-entropy loss function is suitable for classification prediction tasks and can ensure robust learning of discrete class probabilities, the cross-entropy loss function is adopted. For the SE task, the mean squared error (MSE) loss function is adopted because it can effectively minimize the variance of continuous regression predictions.
[0090] S6. Fault Simulation and Feature Library Construction Based on Digital Twin Model: In the digital twin model, various possible fault scenarios are simulated, and the model output data under the corresponding fault scenarios is obtained. After performing the same data preprocessing operations as in S1 on the output data, fault features are extracted to construct a fault feature library. The fault feature library contains feature vectors corresponding to different fault types. In the step of fault simulation and feature library construction based on the digital twin model, the fault simulation is achieved by perturbing the key parameters of the device or modifying the structure of the model in the digital twin model. The fault feature library is stored and managed using a database management system and has functions of fast query and update.
[0091] S7. Real-time Fault Diagnosis: The preprocessed real-time acquisition data is input into the trained fault diagnosis model. The fault diagnosis model determines whether the device has a fault and the type of the fault by comparing the real-time data features with the feature vectors in the fault feature library. If it is determined that the device has a fault, the digital twin model is further used for fault tracing to determine the specific location and cause of the fault. In the step of real-time fault diagnosis, the fault diagnosis model uses deep learning algorithms such as convolutional neural network (CNN), recurrent neural network (RNN), or long short-term memory network (LSTM). And when training the fault diagnosis model, a cross-validation method is adopted to improve the generalization ability of the model. The fault tracing is achieved by analyzing the state changes and signal propagation paths before and after the fault occurs in the digital twin model.
[0092] S8. Diagnosis Result Feedback and Model Update: The fault diagnosis results are fed back to the device operators and maintenance personnel, and the digital twin model and the fault diagnosis model are updated according to the actual maintenance situation and newly acquired data to improve the accuracy and adaptability of the diagnosis model. In the step of diagnosis result feedback and model update, the diagnosis results are presented in a visual way, including the fault type, fault location, fault severity, and maintenance suggestions of the device. The model update adopts an incremental learning method, gradually integrating the newly acquired data into the original model for training to reduce the time and computing resource consumption of model training.
[0093] Experimental analysis of this embodiment:
[0094] Algorithm Comparison: Six multi-objective optimization methods are selected. Among them, UW, STCH, and DWA belong to the objective quantization methods, while GradNorm, PCGrad, and GradVac belong to the adaptive gradient methods.
[0095] Such as Figure 6(As shown in (a), UW continuously amplifies the weights of losses, but this amplification effect is uneven and uncontrollable. It continuously amplifies the weights of the classification task, while the amplification factor for the weights of the SE task is relatively small. Eventually, the performance improvement of UW is only slightly better than its baseline, and the impact on the overall result is extremely limited. As Figure 6 (As shown in (b), the weighting scheme of DWA is relatively conservative, and the weights are close to 1. Eventually, the final weight values converge to around 1. This result can be regarded as the algorithm finally degenerating into a simple loss summation scheme. As can be seen from Table 1, the result of DWA is also only slightly improved compared to the result before improvement. STCH acts directly on the loss value, and its mechanism is more complex. However, this method does not show ideal results on the dataset and even leads to a significant performance decline.
[0096] On the other hand, GradNorm needs to calculate and store gradients for each task, which consumes a large amount of storage space. Unfortunately, the computing resources are not sufficient to train GradNorm on the complete dataset. Even when the data volume is reduced to 30% of the original volume, this method still cannot meet the required storage requirements. PCGrad achieved accuracies of 99.509% and 99.461% in the classification task respectively, and obtained an MSE of 0.043 and an MAE of 0.1314 in the SE task. GradVac achieved accuracies of 99.910% and 99.865% in the FC and FL tasks respectively, and obtained an MSE of 0.0268 and an MAE of 0.1415 in the SE task.
[0097] Compared with these methods, the proposed algorithm not only achieved an accuracy of 100% in the classification task, but also significantly reduced the MSE to 0.0023 and the MAE to 0.0370 in the SE task. These experimental results strongly demonstrate the superiority of the proposed method.
[0098] Table 1 Comparison of the performance of different algorithms
[0099]
[0100]
[0101] Perform SOC estimation on this embodiment:
[0102] The FUDS test mainly includes two stages: 1) the charging stage; 2) the continuous charge and discharge stage. The test results are reflected in the SoC curve. If the data set is randomly selected, it may lead to excessive selection of data in the charging stage, resulting in poor fitting of the discharge data. However, for the continuous charge and discharge stage, the data partitioning mechanism cannot distinguish whether the sample data is charging data or discharge data. Therefore, a smaller step value is adopted for the continuous charge and discharge stage during the data augmentation process. In addition, during the data partitioning stage, the SoC data is manually labeled and the data set is partitioned using stratified sampling.
[0103] First, as Figure 7 shown, except for the proposed method and the DWA method, all other methods show an offset phenomenon in both the charging and discharging stages of FUDS. This phenomenon proves the necessity of stratified sampling for charging and discharging data. Second, the Uncertainty method has a good fitting effect on the SoC curve in the charging stage, but a poor fitting effect in the continuous charge and discharge stage. A similar method is the Weight method, which indicates that simply increasing the loss weight of the task continuously cannot fundamentally solve the problem of imbalance between tasks. The DWA method has relatively close effects in both stages, and its adjustment of the loss weight is more precise, proving the infeasibility of manually adjusting the task loss weight. Finally, the PCGrad and GradVac methods have a poor fitting effect on the SoC curve, while the proposed method can well solve this problem.
[0104] The experimental platform of the present invention is as Figure 2 shown, and it consists of an electronic load, a DC power supply, and a battery pack integrated with a BMS. The battery unit is 100Ah lithium iron phosphate, and the operating voltage range is 2.5V to 3.65V. The battery module is composed of six series-connected battery packs, and each battery pack contains two parallel-connected batteries. Each battery pack includes a battery module controlled by an attiny84 microcontroller, and this microcontroller communicates with the central control unit through a communication bus. The control unit consists of an ArduinoMega2560 and an expansion board, which is responsible for adjusting the dynamic operating conditions and cutting off the power supply when a fault occurs in the BMS. The expansion board integrates an HC-05 Bluetooth module for application communication, and the data transmission frequency is 1Hz. The experimental operating conditions are derived from FUDS. The collected data is as Figures 3-6 shown.
[0105] The above has generally described the present invention in detail, but based on the present invention, some modifications or improvements can be made, which are obvious to those of ordinary skill in the technical field. Therefore, the modifications or improvements made without departing from the spirit of the present invention are within the protection scope of the present invention.
Claims
1. A complex equipment intelligent fault diagnosis method assisted by digital multi-twins, characterized by: The steps include: S1. Data preprocessing: The original data is processed using sliding window technology with a window size of 2048; S2. Model construction: The backbone network of the model includes a multi-scale layer and a feature fusion layer; S3, Adaptive gradient optimization: Using cosine similarity φ(g i ,g j )=cos(g i ,g j ) to define the correlation between two tasks; S4, Adaptive learning rate: In order to achieve model convergence, it is necessary to adaptively adjust the learning rate during model training. The cosine annealing algorithm is believed to be able to gradually reduce the value of the learning rate η; S5. Joint loss function: The scalarization method is to achieve parallel optimization in the multi-task learning framework; S6. Fault simulation and feature library construction based on digital twin model: In the digital twin model, various possible fault scenarios are simulated, and the model output data under the corresponding fault scenarios is obtained. After the output data is subjected to the same data preprocessing operation as S1, the fault features are extracted and a fault feature library is constructed. The fault feature library contains feature vectors corresponding to different fault types; S7, real-time fault diagnosis: the pre-processed real-time collected data is input into the trained fault diagnosis model, and the fault diagnosis model determines whether the device fails and the type of fault by comparing the real-time data features with the feature vectors in the fault feature library; if the device fails, the digital twin model is further used to trace the fault source and determine the specific location and cause of the fault; S8. Feedback of diagnostic results and model update: The fault diagnosis results are fed back to equipment operators and maintenance personnel, and the digital twin model and fault diagnosis model are updated according to the actual maintenance situation and newly acquired data to improve the accuracy and adaptability of the diagnostic model.
2. According to claim 1, a complex equipment intelligent fault diagnosis method assisted by digital multi-twins is characterized by: The S1 includes: the first half of the data is set to 40 steps, and the second half is set to 10 steps, so there are 83602 samples in total. In addition, the data set is divided into a training set, a test set, and a validation set in a ratio of 7:2:
1. The multi-scale layer includes: a multi-scale feature fusion block and a multi-scale channel fusion block.
3. According to claim 1, a complex equipment intelligent fault diagnosis method assisted by digital multi-twins is characterized by: The S2 includes: a branch is added from the i-th layer to the i+1-th layer, and the convolution kernel sizes on the branch are 23, 9 and 3 respectively. Taking the second multi-scale layer to the third multi-scale layer as an example, let {R r ,r=1,2} represents the output representation of the second layer, where each R r Corresponding to the feature map of a specific scale r, in the third layer, the output representation is expanded to {R′ r ,r=1,2,3}, capturing more scale information, each output of the third layer represents R′ r They are calculated by the following operations: R′ r =MaxPool(Conv r (Concat(R1,R2))) Concat(·) is a concatenation operation. r (·) represents a one-dimensional convolution operation with a scale of r, and MaxPool(·) is a one-dimensional maximum pooling operation; The S2 includes: the last layer of the backbone network is a multi-feature fusion layer, that is, a linear transformation is performed on the extracted multi-scale features to generate a shared representation: R sh =σ(BN(Concat(R1′,R2′,R3′)W+b)) Where W is the learnable weight matrix, BN represents the batch normalization (BN) operation, b is the bias vector, and R sh represents the final integrated multi-scale feature representation, and σ represents the ReLU activation function.
4. According to claim 1, a complex equipment intelligent fault diagnosis method assisted by digital multi-twins is characterized by: The S3 includes: is the gradient of task i; for simplicity, φ(g i ,g j ) is abbreviated as φ ij , when φ ij <0, it is called a conflict; when a conflict occurs, g i Replace with g when a conflict occurs j The standard value of: In order to enjoy the benefits of the algorithm without conflicting gradients between tasks, the gradients are linearly combined: g′ i =αg i +βg j , where α and β are positive constants; the combinations of vectors are infinite; for convenience, set α = 1 and solve for β: where φ′ ij Yes ij The target direction is calculated using double exponential moving smoothing. ij The main advantage of DEMA is that it reduces the lag effect caused by traditional multi-EMA filtering, making the smoothing result closer to the changes in the original data. Compared with single-order or second-order EMA, DEMA can respond to short-term changes in data more quickly; Among them, γ is the smoothing factor, 0≤γ≤1, k is the number of layers, and They are the EMA of the data and the EMA of the EMA at step t; Assume that the gradients of task 1 and task 2 are and Assume g = g1 + g2, then the deep learning model parameters are updated: i + =θ-η·g′ where θ + is the updated model parameter, g′ is the target gradient; in addition, the smoothing factor γ is set to 0.
5.
5. According to claim 1, a complex equipment intelligent fault diagnosis method assisted by digital multi-twins is characterized by: The S4 includes: the formula of the cosine annealing algorithm is as follows: Among them, η max is the initial value of the learning rate, N cur is the number of batches currently used for training, N cur is the total number of training batches; the periodic decrease of the learning rate under the action of the cosine function allows the optimization process to gradually explore finer areas of the loss surface.
6. According to claim 1, a complex equipment intelligent fault diagnosis method assisted by digital multi-twins is characterized by: The S5 includes: For the FC task and the FL task, the cross entropy loss function is adopted because the cross entropy loss function is suitable for classification prediction tasks and can ensure the robust learning of discrete class probabilities.
7. The method for intelligent fault diagnosis of complex equipment assisted by digital multi-twins according to claim 1 is characterized in that: In the fault simulation and feature library construction steps based on the digital twin model, the fault simulation is achieved by perturbing the key parameters of the equipment in the digital twin model or modifying the structure of the model; the fault feature library is stored and managed using a database management system and has fast query and update functions.
8. The method for intelligent fault diagnosis of complex equipment assisted by digital multi-twins according to claim 1 is characterized in that: In the real-time fault diagnosis step, the fault diagnosis model adopts a deep learning algorithm such as a convolutional neural network (CNN), a recurrent neural network (RNN) or a long short-term memory network (LSTM), and when training the fault diagnosis model, a cross-validation method is used to improve the generalization ability of the model; the fault tracing is achieved by analyzing the state changes and signal propagation paths before and after the fault occurs in the digital twin model.
9. The method for intelligent fault diagnosis of complex equipment assisted by digital multi-twins according to claim 1 is characterized in that: In the diagnostic result feedback and model update step, the diagnostic result is presented in a visual manner, including the equipment's fault type, fault location, fault severity, and maintenance recommendations; the model update uses an incremental learning method to gradually integrate the newly acquired data into the original model for training, so as to reduce the model training time and computing resource consumption.