Lithium battery diagnosis method and system based on multivariable and infrared point cloud
Through the model of Cyclometric geometric modal decomposition, CBAM and ECA attention mechanism optimization and PTP multimodal fusion method, the problem of insufficient multimodal data fusion in the existing technology is solved, efficient and accurate diagnosis of the healthy state of lithium batteries is achieved, and information utilization and model adaptability are improved.
Patent Information
- Application Number
- CN202510415922.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-18
AI Technical Summary
The existing lithium battery health monitoring technology has shortcomings in the efficient fusion of multimodal data, and lacks an effective cross-modal feature interaction mechanism, resulting in low information utilization and difficulty in fully reflecting the comprehensive health status of the battery.
The voltage, current, temperature, and voiceprint signals are modally decomposed by the CBAM attention mechanism optimization, and combined with the Tokenformer model optimized by the CBAM attention mechanism and the YOLOv9 model improved by the ECA attention mechanism, fault diagnosis is carried out through the PTP multimodal fusion method and the improved extreme learning machine (IELM), so that the deep fusion and efficient analysis of multimodal data can be achieved.
It significantly improves the diagnostic accuracy and reliability of the healthy state of lithium batteries, optimizes the computing efficiency, reduces the dependence on labeled data, and enhances the generalization ability and adaptability of the model.
Smart Images

Figure BDA0005343940630000021 
Figure BDA0005343940630000022 
Figure BDA0005343940630000031
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of lithium battery health prediction and monitoring, and particularly relates to a lithium battery diagnosis method and system based on multivariate and infrared point cloud. Background Art
[0002] In recent years, with the rapid development of new energy vehicles, energy storage systems, consumer electronics and other fields, as a core energy storage device, the health status monitoring and diagnosis technology of lithium batteries has received extensive attention. The industry's demand for real-time monitoring of lithium batteries is increasing day by day, including the accurate assessment of key indicators such as the charge and discharge performance, capacity attenuation, and internal resistance change of the battery. At present, lithium battery health monitoring technology has gradually developed from traditional single-parameter analysis to multi-dimensional and intelligent directions, combining advanced technologies such as sensor data, image recognition, infrared thermal imaging, and acoustic fingerprint analysis to improve the diagnostic accuracy and reliability. At the same time, with the progress of artificial intelligence and Internet of Things technologies, lithium battery health monitoring is evolving towards automation, real-time and predictive maintenance, providing important support for battery safety management and life optimization.
[0003] Existing lithium battery health monitoring technologies mainly rely on electrochemical parameter analysis and physical signal detection. Traditional methods evaluate the health status of the battery by collecting basic parameters such as voltage, current, and temperature of the battery, combined with ohmic internal resistance testing and capacity attenuation analysis. In addition, some advanced technologies use infrared thermal imaging to monitor the temperature distribution of the battery, use acoustic fingerprint recognition technology to capture abnormal sounds inside the battery, and combine image processing algorithms to detect battery appearance defects (such as electrolyte leakage, electrode deformation, etc.).
[0004] Although certain progress has been made in the existing technologies, there are still obvious deficiencies in the efficient fusion of multi-modal data. Current monitoring methods usually process different modal data (such as electrical signals, infrared images, acoustic fingerprints, etc.) independently, lacking an effective cross-modal feature interaction mechanism, resulting in low information utilization rate and difficulty in comprehensively reflecting the comprehensive health status of the battery. Summary of the Invention
[0005] Object of the Invention: The object of the present invention is to provide a lithium battery diagnosis method based on multivariate and infrared point cloud that realizes multi-modal coupling, improves information utilization rate and battery health detection accuracy; on the other hand, to provide a lithium battery diagnosis system based on multivariate and infrared point cloud.
[0006] Technical Solution: The lithium battery diagnosis method described in the present invention includes the following steps:
[0007] (1) Use sensors to collect the voltage, current, temperature, and voiceprint information of lithium batteries respectively, and use a monitoring camera to obtain infrared three-dimensional point cloud imaging and image information, so as to comprehensively collect the multi-dimensional state information of lithium batteries, provide a complete data basis for subsequent analysis, and ensure the comprehensiveness and accuracy of monitoring;
[0008] (2) Adopt the symplectic geometric mode decomposition (SGMD) to perform mode decomposition on the voltage, current, temperature, and voiceprint signals to obtain the intrinsic mode functions (IMFs), and combine the original signals with the IMFs into a two-dimensional input matrix, which can more comprehensively extract the features in the signals, enhance the completeness of feature expression, greatly enhance the generalization ability and robustness of the model, and is suitable for the processing and analysis of complex signals;
[0009] (3) Use the CBAM attention mechanism to optimize the Tokenformer model, use this model to process the two-dimensional input matrix, and weighted output the prediction results, which can efficiently process multi-modal input data, achieve fast convergence and efficient training, and improve the accuracy and efficiency of lithium battery health state prediction;
[0010] (4) Use the ECA attention mechanism to improve the YOLOv9 model, input the video image and infrared three-dimensional point cloud into this model and extract image features, significantly improve the detection ability of small targets, optimize feature fusion, reduce redundant information, and the improved YOLOv9 model has significant improvements in detection accuracy and speed, and can significantly improve the efficiency of fault diagnosis;
[0011] (5) Adopt the PTP multi-modal fusion method to fuse the prediction results of Tokenformer and the image features of YOLOv9 to generate a two-dimensional fusion matrix. Through the explicit interaction of high-order moments and low-rank tensor approximation, the features in multi-modal data are efficiently deeply fused to generate a richer feature representation, improve the accuracy and reliability of fault diagnosis, optimize the calculation efficiency, reduce the dependence on a large amount of labeled data, enhance the generalization ability and adaptability of the model, so as to perform well in lithium battery health monitoring and diagnosis, and significantly improve the overall performance of the system;
[0012] (6) Use the improved extreme learning machine (IELM) to perform fault diagnosis on the fusion matrix and output the lithium battery safety monitoring results, which combines the advantages of robustness and kernel methods, effectively reduces the influence of noise and outliers, realizes high-precision fault diagnosis and battery state evaluation, and improves the reliability of the entire monitoring system.
[0013] Preferably, the use of SGMD mode to decompose and combine variables into a two-dimensional input matrix described in step 2 includes:
[0014] (21) Let a given discrete vibration signal of a variable be \(x = x_1, x_2, \ldots, x\) n, the one-dimensional signal is reconstructed into a multi-dimensional signal through the Takers embedding theorem, so as to generate a trajectory matrix X to reconstruct the given signal;
[0015] (22) Perform a flat geometric matrix transformation to construct a covariance symmetric matrix A = X T X, and construct a Hamiltonian matrix M from the matrix A:
[0016]
[0017] Let W = M 2 , and perform a symplectic geometric similarity transformation to obtain an orthogonal symplectic matrix Q:
[0018]
[0019] Among them, the matrix B is an upper triangular matrix, and b y = 0 (i > j + 1);
[0020] (23) Let the matrix P be a Householder matrix, let the matrix G = [P 0; 0 P], use the matrix G to replace the matrix Q, and use the Schmidt orthogonalization to transform the upper triangular matrix B into the matrix W; the eigenvalues λ1, λ2,..., λ d of the matrix B are obtained from the matrix decomposition, and the eigenvalues of the matrix A are obtained by descending order as Denote H i (i = 1, 2,..., d) as the eigenvectors corresponding to the eigenvalues of the matrix A;
[0021] (24) Through the transformation coefficient matrix obtain the reconstructed initial single-component matrix Z i = H i S i (i = 1, 2,…, d), and then obtain the reconstructed matrix Z = Z1 + Z2 +... + Z d , and perform diagonal averaging on the matrix Z i . Let the elements of the matrix Z i be Z ij , and define:
[0022]
[0023] Then perform diagonal averaging according to the following formula:
[0024]
[0025] Convert the reconstructed matrix Z into d groups of reconstructed independent components Y i (y1, y2,…, y n ), and represent the original time series Y = Y1 + Y2 + Y3 +... + Y d, combine Y i with the components having a relatively high similarity to obtain the symplectic geometry component SGC1;
[0026] (27) Delete the components that make up SGC1 from the matrix Y to obtain the residual signal g 1 , and calculate the normalized mean square error NMSE by the following formula:
[0027]
[0028] where h is the number of iterations, and g h represents the residual component;
[0029] Repeat the above iterative process until the NMSE is less than the preset threshold, then the iteration stops to obtain the IMF component:
[0030]
[0031] where N is the number of obtained SGC components, and g (N+1) (n) is the residual component, and the obtained mode is IIMF;
[0032] Repeat the above process to perform modal decomposition on the variable voltage, temperature, and voiceprint respectively. The obtained intrinsic mode functions are VIMF, CIMF, and SIMF respectively;
[0033] Pairwise combine the original signal x with N intrinsic mode functions x 1 imf , x 2 imf ,..., x N imf respectively to form a two-dimensional input matrix X.
[0034] Through iterative optimization and recombination of similar components, ensure the accuracy and completeness of the decomposition results, and combine the original signal with the IMF components to construct a two-dimensional input matrix, which not only completely retains the time-frequency characteristics of the signal, but also significantly enhances the feature expression ability, provides high-information input features for subsequent multi-modal analysis, and effectively improves the model's ability to extract the state characteristics of lithium batteries under complex working conditions.
[0035] Preferably, the processing process of the Tokenformer model described in step 3 includes:
[0036] (31) Initialize the model parameters, where the parameters include the dimension d1 of the input token, the dimension d2 of the output token, and the number n of parameter tokens. The parameter tokens are divided into Key parameter tokens KP and Value parameter tokens VP, and their dimensions are n×d1 and n×d2 respectively;
[0037] (32) Calculate the Pattention layer, through the formula Calculate the interaction between the input token and the parameter token, where: X is the input token, with a shape of , T is the sequence length, and d1 is the input dimension. k p is the Key parameter token, with a shape of V P is the Valu parameter token, with a shape of Θ is a modified softmax operation for stable optimization;
[0038] (33) In the Pattention layer, calculate the attention score S:
[0039]
[0040] where is the similarity matrix between the input token and the Key parameter token, τ is the scaling factor, defaulting to f is a non-linear function, and here the GeLU function is used;
[0041] (34) Introduce the CBAM attention mechanism in the Pattention layer to refine the input feature map in two stages, and calculate the output M c (F):
[0042] M C M(F) = σ(MLP(AvgPool(F)) + MLP(MaxPool(F)))
[0043] where F is the input feature map, AvgPool and MaxPool respectively represent the global average pooling and max pooling operations, MLP represents the multi-layer perceptron, and σ represents the Sigmoid activation function Q;
[0044] Calculate the output M s (F):
[0045] M S M(F) = σ(f 7×7 ([AvgPool(F); MaxPool(F)]))
[0046] where f 7×7 represents a 7×7 convolution operation, and [Avgpool(F); MaxPool(F)] represents concatenating the average pooling and max pooling results along the channel axis;
[0047] (35) The model is extended by adding a new key-value parameter token without changing the input or output dimensions. The extended parameter token set is:
[0048]
[0049] where [old, new] represents the concatenation operation along the token dimension, and is the scaled parameter set. The forward pass of the scaled model is defined as
[0050] (36) The output layer of the model is designed to have the same number of neurons as the number of classes C. The softmax activation function is used to convert the original output of the model into a probability distribution. The softmax function converts the original output z i of each class into a probability p i :
[0051]
[0052] (37) For each sample, the cross-entropy loss is defined as:
[0053]
[0054] where y i is the one-hot encoding of the true label. If the sample belongs to the i-th class, then y i = 1, otherwise y i = 0;
[0055] During the training process, the average loss of a batch is calculated to stabilize and accelerate the training process:
[0056]
[0057] where N is the number of samples in the batch, is the cross-entropy loss of the k-th sample;
[0058] The prediction result is output after weighting the current, voltage, temperature, and voiceprint.
[0059] A dynamic interaction relationship between input tokens and model parameters is established through a parameterized attention mechanism (Key / Value parameter tokens), and an improved softmax operation is adopted to ensure the stability of the training process. Secondly, the CBAM module is innovatively integrated into the attention layer, and through the dual refinement mechanism of channel attention (global average / max pooling + MLP) and spatial attention (convolutional feature fusion), intelligent focusing on multi-dimensional time series features such as voltage and current is achieved. Finally, through an extensible parameter token architecture and cross-entropy loss optimization, while maintaining the flexibility of the model structure, the training efficiency is ensured. This design enables the model to adaptively extract key features in each modality signal, significantly improving the discrimination of feature expressions. The differential weighting processing of different modalities such as current and voltage further strengthens the contribution of important features, and finally realizes the high-precision prediction of the health state of lithium batteries, providing an effective time series feature analysis tool for battery fault diagnosis under complex working conditions.
[0060] Preferably, the extraction of image features by the YOLOv9 model described in step 4 includes:
[0061] (41) Introduce a multi-scale ECA attention mechanism in the Backbone layer, and through one-dimensional convolution, channel weight allocation is realized to generate the final output feature map;
[0062] (42) Use a combination of cross-entropy loss and generalized dice loss to optimize model training, and its calculation formula is:
[0063]
[0064] Among them, the first term is the cross-entropy loss function, the second term is the generalized dice loss function, and p, q, N, and c respectively represent the sign function, predicted probability, total number of pixels in the image, and number of label categories.
[0065] By improving the YOLOv9 model, the ability to extract lithium battery image features is significantly improved. A multi-scale ECA attention mechanism is innovatively introduced in the Backbone layer, and one-dimensional convolution is used to achieve adaptive channel weighting, effectively enhancing the representation ability of key features (such as battery hot spots, structural abnormalities, etc.). At the same time, a composite loss function of cross-entropy loss and generalized dice loss is adopted, which not only ensures classification accuracy but also optimizes pixel-level feature matching, enabling the model to achieve more accurate positioning and recognition of tiny defects and temperature anomalies in the infrared three-dimensional point cloud while maintaining a high detection speed, providing high-quality visual feature representations for subsequent multi-modal fusion.
[0066] Preferably, the realization of channel weight allocation through one-dimensional convolution to generate the final output feature map includes:
[0067] Perform global average pooling on the input feature map;
[0068] Perform one-dimensional convolution operation with a convolution kernel of size k, and generate the weights w of each channel through the Sigmoid activation function ω = σ(C1Dk(y)).
[0069] Multiply the obtained weights with the original input feature map channel by channel to generate the final output feature map.
[0070] This design first replaces the traditional fully connected layer with lightweight one-dimensional convolution, while reducing the computational complexity and maintaining the channel interaction ability; secondly, it uses the Sigmoid function to implement soft attention weighting in the range of 0-1, effectively highlighting important channel features; finally, the combination of global average pooling and local convolution captures both global statistical characteristics and local channel correlations, and the finally generated feature map has stronger representation ability, which is especially suitable for the enhanced extraction of weak thermal anomaly features in lithium battery infrared images.
[0071] Preferably, the PTP multimodal fusion method described in step 5 includes:[[]]
[0072] (51) Merge the feature sets into a joint compact representation z by explicitly interacting with the PTP block using higher-order moments, and concatenate a set of M feature vectors into a long feature vector
[0073] (52) Use the P-th tensor product of the concatenated feature vectors z12···M to obtain a P-th order polynomial feature tensor
[0074] In the formula is the tensor product operator, add the constant term 1, Z P can represent all possible P-th order polynomial expansions;
[0075] The influence of the P-polynomial interaction between features is completely determined by the pooling weight tensor as:[[]]
[0076]
[0077] where z h represents the h-th element of the h-dimensional fusion vector z, i p represents the higher-order term of the p-th order modality, and the number of parameters of W h grows exponentially with the polynomial order p;
[0078] (53) Use low-rank TNs to effectively approximate W h , assuming that W h allows the rank-RCP format, then z h becomes:[[]]
[0079]
[0080] Assume is meaningful for all p ∈ [P], thus is the set of fusion parameters to be estimated. If W h allows the TR format, then:
[0081]
[0082] where the third-order kernel tensor is the fusion parameter, defined by r P+1 = r1 are the TR-ranks. For all p ∈ [P], assume a shared
[0083] (54) The two-dimensional matrix combining the output prediction result with the image
[0084] This design breaks through the limitations of traditional feature stitching, realizes the non-linear deep fusion of cross-modal features, not only retains the interaction information between temporal features such as voltage and temperature and the spatial features of infrared images, but also optimizes the calculation efficiency through the parameter sharing mechanism, providing a high-information fusion feature representation for the accurate diagnosis of the health state of lithium batteries.
[0085] Preferably, the IELM in step 6 for fault diagnosis of the fusion matrix includes:
[0086] (61) Let the data set Φ = ((x i ; t i ), x i ∈ R m , t i ∈ R p ), where x N = [x i , x i1 ,..., x i2 ,..., x im T is the network input, and t i = [t i1 , t i2 ,..., t ip T is the network output; the network structure of ELM is described by the following formula:
[0087]
[0088] where G(α l , β l , x) is the output of the l-th hidden layer node; construct the minimized error as ‖H - Z‖. By solving the least squares of H = T, we can obtain γ = (H T H) -1 H T Z, where H is the output of the hidden layer node and Z is the label;
[0089] (62) Let the training samples be X = [x1, x2,..., x N1 , which contains N1 samples from C-class targets. The number of hidden layer nodes of RKELM is N h . Then, first perform prototype clustering on X and divide it into N h non-overlapping training sample clusters;
[0090] Use the K-means++ algorithm to cluster the training samples X, where the number of clustering clusters is N h . The obtained clustering centers are y j as the prototype vectors of the j-th cluster, j = 1, 2,..., N h ;
[0091] Take the clustering center Y p as the hidden layer nodes of RKELM and calculate the kernel matrix K0:
[0092]
[0093] where is the kernel function;
[0094] (63) According to the empirical error minimization criterion, calculate the output layer weight matrix
[0095]
[0096] where ξ is the regularization parameter and I is the N h -dimensional identity matrix; in the test stage, let x n be the test sample, then its class vector can be calculated as t n = K(x n , Y p )β, where the class corresponding to the maximum value in t n is the class of the sample x n ;
[0097] (64) Introduce the output matrix of the ELM network under Robust estimation as:
[0098]
[0099] where P(V) = diag(P1(V i ), P2(Vi ),...,P n (V i )),
[0100] The clustering structure of the samples is characterized by the prototype vector set Y p and the distances between the vectors in Y p are relatively large. Taking Y p as the support vectors of RKELM to represent the feature space of the target;
[0101] (65) Minimize the objective function:
[0102]
[0103] where N is the total number of test samples, f(x i ) is the prediction result of the model for the i-th test sample x i , y i is the true label of the i-th test sample, and I(f(x i ≠y i ) is an indicator function. When the prediction result f(x i ) of the model for the i-th sample is not equal to the true label y i , the value of this function is 1, indicating that a diagnostic error has occurred; otherwise, the value of this function is 0.
[0104] By using the K-means++ algorithm to perform prototype clustering on the training samples and selecting representative cluster centers as hidden layer nodes, the coverage ability of the model for the feature space is significantly improved; by calculating the output weights through kernel function mapping and regularized least squares optimization, the generalization performance of the model is ensured; by innovatively introducing a Robust estimation mechanism, the interference of noise and outliers to the diagnostic results is effectively suppressed. This algorithm realizes the accurate discrimination of fault types by minimizing the error diagnosis indicator function. Its unique prototype vector selection strategy not only reduces the computational complexity but also ensures the completeness of feature representation, enabling the system to significantly improve the recognition accuracy of early lithium battery faults while maintaining real-time response capabilities.
[0105] The lithium battery diagnostic system described in the present invention includes:
[0106] A multi-source data acquisition module for collecting voltage, current, temperature, and acoustic fingerprint information of the lithium battery through multiple sensors, and obtaining infrared three-dimensional point cloud imaging and image information through a monitoring camera;
[0107] A signal modal decomposition module, which is used to perform modal decomposition on voltage, current, temperature, and voiceprint signals by using symplectic geometric modal decomposition (SGMD) to obtain intrinsic mode functions (IMFs), and combine the original signals and IMFs into a two-dimensional input matrix;
[0108] A time-series feature extraction module, which is used to process the two-dimensional input matrix by using a Tokenformer model optimized by the CBAM attention mechanism and weighted-output a prediction result;
[0109] A visual feature extraction module, which is used to input infrared three-dimensional point clouds and images into a YOLOv9 model improved by the ECA attention mechanism to extract image features;
[0110] A multi-modal fusion module, which is used to fuse the prediction result of Tokenformer and the image features of YOLOv9 by using the PTP fusion method to generate a two-dimensional fusion matrix;
[0111] A fault diagnosis module, which is used to perform fault diagnosis on the fusion matrix by using an improved extreme learning machine (IELM) and output the lithium battery safety monitoring result.
[0112] A computer-readable storage medium, on which a computer program is stored. The computer program, when executed by a processor, implements the lithium battery diagnosis method based on multi-variables and infrared point clouds.
[0113] A computer device, which includes a memory and a processor. The memory stores a computer program that can be loaded and executed by the processor to implement the lithium battery diagnosis method based on multi-variables and infrared point clouds.
[0114] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages: 1. Through a multi-modal coupling structure, the accuracy and reliability of fault diagnosis are improved, the computing efficiency is optimized, the dependence on a large amount of labeled data is reduced, and the generalization ability and adaptability of the model are enhanced; 2. The CBAM and ECA dual attention mechanisms achieve dynamic focusing and collaborative analysis of cross-modal key features; 3. Based on the PTP fusion method, high-order coupling and robust diagnosis of electrical signals and visual features are realized, comprehensively improving the accuracy and reliability of the monitoring system. Description of the Drawings
[0115] Figure 1 It is a schematic flow chart of the present invention;
[0116] Figure 2 It is a schematic diagram of multi-modal coupling input of the present invention;
[0117] Figure 3 It is a structural diagram of the Tokenformer model of the present invention;
[0118] Figure 4 This is the YOLOv9 model structure diagram of the present invention. DETAILED DESCRIPTION
[0119] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.
[0120] The present invention provides a lithium battery diagnosis method based on multivariable and infrared point cloud, such as Figure 1 As shown, the specific steps include:
[0121] Step 1: Use high-precision voltage and current sensors to collect dynamic electrical parameters during the charging and discharging process in real time; place temperature sensors at key locations of the battery to precisely monitor changes in its thermal state; use highly sensitive microphone equipment to capture and analyze battery soundprint signals to explore the evolution of its internal structure; at the same time, use advanced infrared thermal imaging technology and three-dimensional point cloud scanning equipment to obtain temperature distribution and deformation information on the battery surface; in addition, use a high-definition camera system to monitor the battery appearance in real time to capture any potential visual abnormalities.
[0122] Through this series of means, real-time, multi-dimensional data collection of lithium battery electrochemical performance, thermal state, acoustic characteristics, infrared three-dimensional point cloud images and appearance morphology information is achieved, thereby more comprehensively reflecting the health status of lithium batteries, providing strong support for early warning and precise positioning of faults, and laying a rich and accurate information foundation for subsequent data fusion and comprehensive analysis.
[0123] Step 2: Use symplectic geometric mode decomposition to perform mode decomposition on voltage, current, temperature and voiceprint signals, which can adaptively decompose complex signals into multiple intrinsic mode functions (IMFs) and retain the physical characteristics of the signals; reconstruct the one-dimensional signal into a multi-dimensional trajectory matrix through the Takens embedding theorem, and combine symplectic geometric matrix transformation and eigenvalue decomposition to extract the key features of the signal; finally, as Figure 2 As shown in the figure, the original signal and the decomposed IMF are combined to form a two-dimensional input matrix, which provides rich feature information for the subsequent deep learning model. While improving the accuracy and efficiency of signal analysis, it also enhances the robustness and generalization ability of the model. The specific steps are as follows:
[0124] (21) Let the given discrete vibration signal of a variable be x = x1, x2, …, x n , by applying the Takers embedding theorem, one-dimensional data can be reconstructed into a multidimensional signal, thereby generating a trajectory matrix X to reconstruct the given signal:
[0125]
[0126] where ε and d are the delay time and the embedding dimension respectively, m = n - (d - 1)ε. In most cases, the delay time ε is set to 1, and the embedding dimension d is determined by the following formula:
[0127]
[0128] where n is the data length, f s is the sampling frequency, and f max is the maximum peak frequency in the power spectral density of the signal x;
[0129] (22) Perform a flat geometric matrix transformation. After multidimensionally reconstructing the original signal and obtaining the trajectory matrix, construct the covariance symmetric matrix A = X T X, and construct the Hamiltonian matrix M from the matrix A:
[0130]
[0131] Let W = M 2 , which can satisfy the conditions of symplectic geometric similarity transformation. Given that the matrix w is a Hamiltonian matrix, the rate orthogonal symplectic matrix Q is obtained from A = X T X:
[0132]
[0133] where the matrix B is an upper triangular matrix, b y = 0 (i > j + 1). Since the matrix Q is an orthogonal symplectic matrix and it has the properties of a symplectic matrix, the phase space structure of the Hamiltonian matrix w is maintained during the matrix transformation;
[0134] (23) For the matrix Q and B in the formula Q T WQ is unknown. Perform a symplectic geometric similarity transformation on the matrix w to obtain the matrix Q. Let the matrix P be a Householder matrix, and let the matrix G = [P 0; 0 P]. The matrix G is a Householder matrix and can be used to replace the matrix Q
[52] , and use the defined Schmidt orthogonalization method to transform the upper triangular matrix B to the matrix W;
[0135] (24) Indirectly obtain the eigenvalues of the matrix A by calculating the eigenvalues of the matrix B. The eigenvalues λ1, λ2,..., λ d are the eigenvalues of the matrix B, and the eigenvalues of the matrix A obtained by descending order are:
[0136]
[0137] Denote H i (i = 1, 2,..., d) as the eigenvectors corresponding to the eigenvalues of the matrix A;
[0138] (25) Obtain the transform coefficient matrix Obtain the initial reconstructed single-component matrix Z through the transform coefficient matrix S i = H i S i (i = 1, 2, …, d), the reconstructed matrix Z can be expressed as Z = Z1 + Z2 + … + Z d ;
[0139] (26) Perform diagonal averaging. The dimension of the initial component matrix obtained by the above transformation is m×d. Diagonal average the matrix Z i Perform diagonal averaging. Assume the elements of the matrix Z i are z ij , and define:
[0140]
[0141] Then perform diagonal averaging according to the following formula:
[0142]
[0143] Convert the reconstructed matrix Z into d groups of reconstructed independent components Y of length n according to the above formula i (y1, y2, …, y n ), and represent the original time series Y = Y1 + Y2 + Y3 + … + Y d ;
[0144] (27) Perform similar component recombination. Conduct similarity analysis on the reconstructed time series Y, and merge the similar components. Merge the components with high similarity to Y i with Y1 to obtain the symplectic geometric component SGC1. Delete the components that make up SGC 1 from the matrix Y to obtain the residual signal g 1 , and then calculate the normalized mean square error (NMSE):
[0145]
[0146] where h is the number of iterations, and g h represents the residual component;
[0147] (28) Repeat the above iterative process until the NMSE is less than the given threshold, then stop the iteration to obtain the final result:
[0148]
[0149] In the formula, N is the number of obtained SGC components, and g (N+1) (n) is the residual component. Let N = 5, where N is the number of obtained SGC components, and g (N+1)(n) is the residual component, and the obtained mode is IIMF; repeat the above process to perform modal decomposition on the variable voltage, temperature, and voiceprint respectively. The obtained intrinsic mode functions are VIMF, CIMF, and SIMF respectively;
[0150] (29) Modal coupling input structure, which combines the original signal x with N intrinsic mode functions x 1 imf , x 2 imf ,..., x N imf where the intrinsic mode functions are obtained by decomposing the original signal as described above, N = 4. In the designed signal input structure, the original signal and the decomposed intrinsic mode functions are paired and combined respectively to form a two-dimensional input matrix X.
[0151] Step 3: Process the two-dimensional input matrix using the Tokenformer model optimized by the CBAM attention mechanism, and weighted output the prediction result, as Figure 3 shown. The specific steps are as follows:
[0152] (31) Initialize the model parameters, where the parameters include the dimension d1 of the input token, the dimension d2 of the output token, and the number n of parameter tokens. The parameter tokens are divided into Key parameter token KP and Value parameter token VP, and their dimensions are n×d1 and n×d2 respectively;
[0153] (32) Calculate the Pattention layer, and calculate the interaction between the input token and the parameter token through the following formula:
[0154]
[0155] where: X is the input token, with a shape of , T is the sequence length, d1 is the input dimension, k p is the Key parameter token, with a shape of V P is the Value parameter token, with a shape of , Θ is an improved softmax operation for stable optimization;
[0156] (33) Then calculate the attention score. In the Pattention layer, calculate the attention score S through the following formula:
[0157]
[0158] where is the similarity matrix of the input token and the Key parameter token, and τ is the scaling factor, which defaults to f is a non-linear function, and the GeLU function is used here;
[0159] (34) Introduce the CBAM attention mechanism in the Pattention layer of Tokenformer to refine the input feature map in two stages, enhance the pertinence of feature extraction, and the output M of the channel attention module can be calculated by the following formula c (F):
[0160] M C (F) = σ(MLP(AvgPool(F)) + MLP(MaxPool(F)))
[0161] where F is the input feature map, AvgPool and MaxPool represent global average pooling and max pooling operations respectively, MLP represents a multi-layer perceptron, and σ represents the Sigmoid activation function Q;
[0162] The output M of the spatial attention module s can be calculated by the following formula:
[0163] M S (F) = σ(f 7×7 ([AvgPool(F); MaxPool(F)]))
[0164] where f 7×7 represents a 7×7 convolution operation, and [Avgpool(F); MaxPool(F)] represents concatenating the average pooling and max pooling results along the channel axis;
[0165] (35) Expand the model by adding new key-value parameter tokens without changing the input or output dimensions. The expanded parameter token set is:
[0166]
[0167] where [old, new] represents the concatenation operation along the token dimension, and is the scaled parameter set. The forward pass of the scaled model is defined as:
[0168]
[0169] (36) Design the output layer of the model to have the same number of neurons as the number of classes C, and use the softmax activation function to convert the original output of the model into a probability distribution. The softmax function converts the original output z of each class i into a probability pi :
[0170]
[0171] (37) For each sample, the cross-entropy loss is defined as:
[0172]
[0173] where y i is the one-hot encoding of the true label. If the sample belongs to the i-th class, then y i = 1, otherwise y i = 0;
[0174] During the training process, the average loss of a batch is calculated to stabilize and accelerate the training process:
[0175]
[0176] where N is the number of samples in the batch, is the cross-entropy loss of the k-th sample;
[0177] (38) It is allowed to integrate any number of parameters without changing the input or output dimensions, making it converge faster and accelerating the overall scaling process by increasing the weights. Based on the expert's experience and in-depth understanding of the system characteristics, the weights of current, voltage, temperature, and voiceprint are assigned as 0.2, 0.2, 0.2, and 0.4 respectively, and the predicted results are output after weighted processing.
[0178] Step 4: Input the video image and infrared three-dimensional point cloud data into the YOLOv9 model improved by the ECA attention mechanism for processing, efficiently extract key features and optimize feature fusion. The ECA mechanism efficiently extracts the interdependence between channels through one-dimensional convolution, adaptively assigns channel weights, significantly improves the feature extraction efficiency and small target detection ability. As Figure 4 shown, the improved YOLOv9 model has improved in both detection accuracy and speed, while enhancing the generalization ability, optimizing the loss function, supporting multi-modal data fusion, reducing the dependence on labeled data, and improving the overall efficiency of the model. The specific implementation steps are as follows:
[0179] (41) Introduce an efficient multi-scale ECA attention mechanism at the end of the Backbone layer of the YOLOv9 model to improve the performance of the model in detecting small targets in lithium battery device images. By using one-dimensional convolution, local cross-channel interaction is effectively realized, thereby extracting the interdependence between channels. The specific process is as follows:
[0180] (411) Perform global average pooling on the input feature map:
[0181] (412) Perform a one-dimensional convolution operation with a convolution kernel of size k, and generate the weights w for each channel through the Sigmoid activation function. The formula is as follows:
[0182] ω = σ(C1Dk(y))
[0183] (413) Multiply the obtained weights with the original input feature map channel by channel to generate the final output feature map.
[0184] (42) In the YOLOv9 model, the loss function during the training process is the sum of the cross-entropy loss and the generalized dice loss. Its calculation formula is as follows:
[0185]
[0186] Among them, the first term is the cross-entropy loss function, and the second term is the generalized dice loss function. p, q, N, and c represent the sign function (taking values of 0 or 1), the predicted probability, the total number of pixels in the image, and the number of label categories respectively. The smaller the value of the cross-entropy loss function, the smaller the difference between the probability distribution output by the model and the true label distribution, that is, the better the performance of the model; the generalized dice loss function is calculated at the pixel level. By first performing pixel classification and then calculating the dice loss pixel by pixel, the detailed prediction of the model is optimized.
[0187] Step 5: Use the PTP multi-modal fusion method to fuse the prediction results of the ITokenformer model for the lithium battery state data and the image training results output by the YOLOv9 model. Through the explicit interaction of high-order moments, tensor product operations, and low-rank tensor approximation, the features in the multi-modal data can be efficiently fused in depth to generate a richer feature representation. The specific steps are as follows:
[0188] (51) The PTP block effectively combines the feature sets into a joint compact representation z by using the explicit interaction of high-order moments. Concatenate a set of M feature vectors into a long feature vector:
[0189]
[0190] (52) Use the P-order tensor product of the concatenated feature vector z12···M to obtain a P-order polynomial feature tensor:
[0191]
[0192] In the formula is the tensor product operator, where the constant term "1" is added, and Z PIt can represent all possible P-order polynomial expansions, and the influence of the P-polynomial interaction between features is completely determined by the pooling weight tensor is defined as:
[0193]
[0194] where z h represents the h-th element of the h-dimensional fusion vector z, and i p represents the high-order term of the p-th order modality. The number of parameters of W h grows exponentially with the polynomial order p. Low-rank TNs are used to effectively approximate W h ;
[0195] (53) Assume that W h allows the rank-RCP format, then becomes:
[0196]
[0197] Assume that is meaningful for all p ∈ [P]. Therefore, is the set of fusion parameters to be estimated. If W h allows the TR format, then we can get:
[0198]
[0199] where the third-order core tensor is the fusion parameter, and is defined by r P+1 = r1 as the TR-ranks. For all p ∈ [P], it is reasonably assumed that there is a shared
[0200] (54) The output prediction result is combined with the image into a two-dimensional matrix
[0201]
[0202] Step 6: Use the improved extreme learning machine (IELM) to perform fault diagnosis on the multimodal fusion results of ITokenformer and the YOLOv9 model. The fast convergence, nonlinear mapping ability, and Softmax output of IELM enable it to capture complex feature relationships, improve the diagnosis accuracy, and be able to provide real-time feedback on the prediction results to ensure the safe operation of lithium batteries. The specific steps are as follows:
[0203] (61) Assume there is a dataset Φ = ((x i ; t i ), x i ∈ R m , ti ∈R p ) N , where x i =[x i1 , x i2 ,..., x im T is the network input, and t i =[t i1 , t i2 ,…, t ip T is the network output. The network structure of ELM is described by the following formula:
[0204]
[0205] where G(α l , β l , x) is the output of the l-th hidden layer node; constructing the minimum error as ‖H - Z‖, by solving the least squares of H = T, we can get γ=(H T H) -1 H T Z, where: H is the output of the hidden layer nodes, and Z is the label;
[0206] (62) When the number of hidden layer nodes in the ELM network is small, there may be a problem that some support vectors are highly correlated. This will not only lead to a waste of hidden layer nodes but also make the recognition results unstable. Let the training samples be X = [x1, x2,…, x N1 , which contains N1 samples from C-class targets, and the number of hidden layer nodes of RKELM is N h . Then, first perform prototype clustering on X and divide it into N h non-overlapping training sample clusters. This clustering method can effectively reduce the redundancy between hidden layer nodes and improve the robustness of the model at the same time;
[0207] (63) Use the K-means++ algorithm to cluster the training samples X, where the number of clustering clusters is N h , and the obtained clustering centers are y j as the prototype vectors of the j-th cluster, j = 1, 2,…, N h ;
[0208] (64) Use the clustering centers Y p as the hidden layer nodes of RKELM and calculate the kernel matrix K0:
[0209]
[0210] where, is the kernel function;
[0211] (65) Calculate the output layer weight matrix according to the principle of minimizing empirical error Minimize the following objective function:
[0212]
[0213] where ξ is the regularization parameter and I is the N h dimensional identity matrix; in the test phase, let x n be the test sample, then its class vector can be calculated as:
[0214] t n = K(x n , Y p )β
[0215] In the formula, the class corresponding to the maximum value in t n is the class of the sample x n
[0216] (66) To further improve the robustness of the model, introduce the output matrix of the ELM network under Robust estimation. The introduced output matrix of the ELM network under Robust estimation is:
[0217]
[0218] In the formula, P(V) = diag(P1(V i ), P2(V i ),..., P n (V i ))
[0219] The clustering structure of the sample is characterized by the prototype vector set Y p , and the distance between pairwise vectors in Y p is relatively large. Using Y p as the support vectors of RKELM can more comprehensively represent the feature space of the target and improve the target recognition performance.
[0220] (67) To build a model that can accurately predict the battery health status and achieve this goal, it is necessary to minimize the number of errors in the diagnosis process of the model;
[0221] Define the following objective function:
[0222]
[0223] In the formula, N is the total number of test samples, f(x i ) is the prediction result of the model for the i-th test sample x i , and y i is the true label of the i-th test sample, I(f(xi ≠y i ) is an indicator function. When the prediction result f(x i ) of the model for the i-th sample is i not equal to the true label y, the value of this function is 1, indicating that a diagnostic error has occurred; otherwise, the value of this function is 0.
[0224] By minimizing the above objective function, it is ensured that the model minimizes the number of errors as much as possible during the diagnosis process, thereby improving the accuracy and reliability of the model.
[0225] A lithium battery diagnosis system corresponding to the above lithium battery diagnosis method includes the following modules:
[0226] A multi-source data acquisition module, which is used to collect voltage, current, temperature and acoustic fingerprint information of the lithium battery through multiple sensors, and obtain infrared three-dimensional point cloud imaging and image information through a monitoring camera;
[0227] A signal modal decomposition module, which is used to perform modal decomposition on voltage, current, temperature and acoustic fingerprint signals by using symplectic geometric modal decomposition (SGMD) to obtain intrinsic mode functions (IMFs), and combine the original signals with the IMFs into a two-dimensional input matrix;
[0228] A time series feature extraction module, which is used to process the two-dimensional input matrix by using the Tokenformer model optimized by the CBAM attention mechanism and output the prediction result with weights;
[0229] A visual feature extraction module, which is used to input the infrared three-dimensional point cloud and image into the YOLOv9 model improved by the ECA attention mechanism to extract image features;
[0230] A multi-modal fusion module, which is used to fuse the prediction result of the Tokenformer and the image features of the YOLOv9 by using the PTP fusion method to generate a two-dimensional fusion matrix;
[0231] A fault diagnosis module, which is used to perform fault diagnosis on the fusion matrix by using the improved extreme learning machine (IELM) and output the lithium battery safety monitoring result.
[0232] The embodiment of the present invention also discloses a computer-readable storage medium.
[0233] Specifically, the computer-readable storage medium is used to store a computer program. When the computer program is executed by a processor, the methods in the above method embodiments are implemented. Those skilled in the art can understand that to implement all or part of the processes in the above method embodiments of the present application, it can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above types of memories.
[0234] The embodiment of the present invention also discloses a computer device.
[0235] Specifically, the computer device can be a desktop computer, a laptop computer, a palm computer, a cloud server, and other computer devices. The computer device can include, but is not limited to, a processor and a memory. Among them, the processor and the memory can be connected through a bus or other means. Among them, the processor can be a central processing unit (CPU). The processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, graphics processing units (GPUs), embedded neural network processors (NPUs), or other dedicated deep learning coprocessors, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or a combination of the above types of chips.
[0236] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the methods in the above embodiments of the present application. The processor executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory, that is, to implement the methods in the above method embodiments. The memory may include a program storage area and a data storage area. Among them, the program storage area can store control units and application programs required for at least one function; the data storage area can store data created by the processor and the like. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include memories remotely provided relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
Claims
1. A lithium battery diagnosis method based on multi-variables and infrared point cloud, characterized in that, It includes the following steps: (1) Use sensors to collect the voltage, current, temperature, and voiceprint information of the lithium battery respectively, and use a monitoring camera to obtain infrared three-dimensional point cloud imaging and image information; (2) Adopt the symplectic geometric mode decomposition (SGMD) to perform mode decomposition on the voltage, current, temperature, and voiceprint signals to obtain the intrinsic mode functions (IMFs), and combine the original signals with the IMFs into a two-dimensional input matrix; (3) Use the CBAM attention mechanism to optimize the Tokenformer model, and use this model to process the two-dimensional input matrix and weighted output the prediction results; (4) Use the ECA attention mechanism to improve the YOLOv9 model, input the video image and infrared three-dimensional point cloud into this model and extract image features; (5) Adopt the PTP multi-modal fusion method to fuse the prediction results of Tokenformer and the image features of YOLOv9 to generate a two-dimensional fusion matrix; (6) Use the improved extreme learning machine (IELM) to perform fault diagnosis on the fusion matrix and output the lithium battery safety monitoring results.
2. The lithium battery diagnosis method according to claim 1, wherein The SGMD mode decomposition and combination into a two-dimensional input matrix described in step 2 include: (21) Let the variable discrete vibration signal be \(x = x_1, x_2, \ldots, x\) n , and according to the Takens embedding theorem, the one-dimensional signal is reconstructed into a multi-dimensional signal, and the trajectory matrix \(X\) is generated to reconstruct the given signal; (22) Perform a plane geometric matrix transformation to construct a covariance symmetric matrix A = X T X, and construct a Hamiltonian matrix M from matrix A: Let \(W = M\) 2 , perform a symplectic geometric similarity transformation to obtain the orthogonal symplectic matrix \(Q\): Among them, matrix B is an upper triangular matrix, and b y = 0 (i > j + 1); (23) Let the matrix \(P\) be a Householder matrix, and let the matrix \(G=\begin{bmatrix}P&0\\0&P\end{bmatrix}\). Use the matrix \(G\) to replace the matrix \(Q\), and transform the upper triangular matrix \(B\) into the matrix \(W\) by using the Schmidt orthogonalization; the eigenvalues \(\lambda_1,\lambda_2,\cdots,\lambda\) of the matrix \(B\) are obtained from the matrix decomposition; d , and the eigenvalues of the matrix \(A\) obtained by arranging them in descending order are Denote \(H\) i \((i = 1,2,\cdots,d)\) as the eigenvectors corresponding to the eigenvalues of the matrix \(A\); (24) Obtain the initial single-component matrix \(Z\) of reconstruction by transforming the coefficient matrix i = \(H\) i \(S\) i (i = 1, 2, …, d), and obtain the reconstruction matrix \(Z = Z_1+Z_2+\cdots+Z\) (i = 1, 2, …, d), and obtain the reconstruction matrix \(Z = Z_1+Z_2+\cdots+Z\) d , and perform diagonal averaging on the matrix \(Z\) i . Assume that the elements of the matrix \(Z\) i are \(Z\) ij , and define: Then perform diagonal averaging according to the following formula: Convert the reconstructed matrix Z into d groups of reconstructed independent components Y with a length of n i (y1, y2, …, y n ),, represent the original time series Y = Y1 + Y2 + Y3 + … + Y with the sum of d groups of reconstructed components d , and merge Y i with the components with higher similarity to obtain the symplectic geometric component SGC1; (25)Delete the components that make up SGC1 in matrix Y to obtain the residual signal g 1 , and calculate the normalized mean square error NMSE by the following formula: where h is the number of iterations and g h represents the residual component; Repeat the above iterative process until the NMSE is less than the preset threshold, then stop the iteration to obtain the IMF components: where N is the number of obtained SGC components, and g (N+1) (n) is the residual component, and the obtained mode is IIMF; Repeat the above process to perform mode decomposition on the variables voltage, temperature, and voiceprint respectively. The obtained intrinsic mode functions are VIMF, CIMF, and SIMF respectively; (26) Pair the original signal x with N intrinsic mode functions x 1 imf , x 2 imf ,..., x N imf respectively to form a two-dimensional input matrix X.
3. The lithium battery diagnosis method according to claim 1, characterized in that, The processing process of the Tokenformer model described in step 3 includes: (31) Initialize the model parameters, where the parameters include the dimension d1 of the input token, the dimension d2 of the output token, and the number n of parameter tokens. The parameter tokens are divided into Key parameter tokens (KP) and Value parameter tokens (VP), and their dimensions are n×d1 and n×d2 respectively; (32) Calculate the Pattention layer, through the formula Calculate the interaction between the input token and the parameter token, where: X is the input token, with a shape of R T×d1 , T is the sequence length, and d1 is the input dimension. k p is the Key parameter token, with a shape of V P is the Valu parameter token, with a shape of Θ is a modified softmax operation for stable optimization; (33) In the Pattention layer, calculate the attention score S: Among them is the similarity matrix of the input token and the Key parameter token, and τ is the scaling factor, defaulting to f is the GeLU function; (34) Introduce the CBAM attention mechanism in the Pattention layer to refine the input feature map in two stages and calculate the output M of the channel attention module c (F): M C (F) = σ(MLP(AvgPool(F)) + MLP(MaxPool(F))) where F is the input feature map, AvgPool and MaxPool respectively represent global average pooling and max pooling operations, MLP represents a multi-layer perceptron, and σ represents the Sigmoid activation function Q; Calculate the output M of the spatial attention module s (F): M S (F) = σ(f 7×7 ([AvgPool(F); MaxPool(F)])) where f 7×7 represents a 7×7 convolution operation, and [Avgpool(F); MaxPool(F)] represents concatenating the average pooling and max pooling results along the channel axis; (35) Expand the model by adding new key-value parameter tokens. The expanded parameter token set is: where [old, new] denotes the concatenation operation along the token dimension, and is the scaled set of parameters, and the forward pass of the scaling model is defined as (36) Design the output layer of the model to have the same number of neurons as the number of classes C, and use the softmax activation function to convert the original output of the model into a probability distribution. The softmax function converts the original output z of each class i into a probability p i : (37) For each sample, the cross-entropy loss is defined as: Among them, y i is the one-hot encoding of the true label. If the sample belongs to the i-th class, then y i = 1, otherwise y i = 0; During the training process, calculate the average loss of a batch to stabilize and accelerate the training process: where N is the number of samples in the batch, is the cross-entropy loss of the k-th sample; Output the prediction results after weighted processing of the current, voltage, temperature, and voiceprint.
4. The lithium battery diagnosis method according to claim 1, wherein The extraction of image features by the YOLOv9 model described in step 4 includes: (41) Introduce a multi-scale ECA attention mechanism in the Backbone layer, and achieve channel weight allocation through one-dimensional convolution to generate the final output feature map; (42) Adopt the combination of cross-entropy loss and generalized dice loss to optimize the model training, and its calculation formula is: where the first term is the cross-entropy loss function, the second term is the generalized dice loss function, and p, q, N, and c respectively represent the sign function, prediction probability, total number of pixels in the image, and number of label categories.
5. The lithium battery diagnosis method according to claim 4, wherein The channel weight allocation through one-dimensional convolution to generate the final output feature map includes: Performing global average pooling on the input feature map; Performing one-dimensional convolution operation with a convolution kernel of size k, and generating the weight w of each channel through the Sigmoid activation function ω = σ(C1Dk(y)); Multiplying the obtained weights with the original input feature map channel by channel to generate the final output feature map.
6. The lithium battery diagnosis method according to claim 1, wherein The PTP multimodal fusion method described in step 5 includes: (51) Merge the feature sets into a joint compact representation z by explicitly interacting the PTP blocks using higher-order moments, concatenate a set of M feature vectors into a long feature vector (52) Using the P - th tensor product of the connected eigenvectors z12···M, a P - th order polynomial feature tensor is obtained where is the tensor product operator, adding a constant term 1, Z P can represent all possible P-order polynomial expansions; The influence of the P-polynomial interaction between features is completely determined by the pooling weight tensor defined as: where, z h represents the h-th element of the h-dimensional fusion vector z, and i p represents the high-order term of the p-th order mode. The number of parameters of W h grows exponentially with the polynomial order p; (53) Use low-rank TNs to effectively approximate W h , assuming that W h allows the rank-RCP format, then z h becomes: Hypothesis is meaningful for all p ∈ [P], so is the set of fusion parameters to be estimated. If W h allows the TR format, then: Among them, the third-order core tensor is a fusion parameter, defined by r P+1 = r1 as the TR-ranks. For all p ∈ [P], assume a shared (54)Two-dimensional matrix combining output prediction results and images 7. The lithium battery diagnosis method according to claim 1, characterized in that, The IELM for fault diagnosis of the fusion matrix described in step 6 includes: (61) Let the data set Φ = ((x i ;t i ),x i ∈R m ,t i ∈R p ) N , where x i =[x i1 ,x i2 ,…,x im ] T is the network input, t i =[t i1 ,t i2 ,…,t ip ] T is the network output; the network structure of ELM is described by the following formula: where G(α l , β l , x) is the output of the l-th hidden layer node; constructing the minimized error as ‖H - Z‖, by solving the least squares of H = T, we can obtain γ = (H T H) -1 H T Z, where H is the output of the hidden layer nodes and Z is the label; (62) Let the training samples be X = [x1, x2, …, x N1 , which contains N1 samples from C classes of targets, and the number of hidden layer nodes of RKELM is N h . Then, X is first subjected to prototype clustering and divided into N h non-overlapping training sample clusters; Cluster the training samples X using the K-means++ algorithm, where the number of clustering clusters is N h , and the obtained clustering centers are y j which is the prototype vector of the j-th cluster, j = 1, 2, …, N h ; Take the clustering center Y p as the hidden layer nodes of RKELM, and calculate the kernel matrix K0: wherein is a kernel function; (62) Calculate the output layer weight matrix according to the empirical error minimization criterion where ξ is the regularization parameter and I is the N h -dimensional identity matrix; in the test phase, let x n be the test sample, then its class vector can be calculated as t n = K(x n , Y p )β, where the class corresponding to the maximum value in t n is the class of the sample x n . (63) Introducing the output matrix of the ELM network under Robust estimation as: where P(V) = diag(P1(V i ), P2(V i ),..., P n (V i )) The clustering structure of the samples is characterized by a set of prototype vectors Y p and the distances between pairs of vectors within Y p are relatively large. Taking Y p as the support vectors of RKELM to represent the feature space of the target; (64) Minimizing the objective function: where N is the total number of test samples, f(x i ) is the prediction result of the model for the i-th test sample x i , y i is the true label of the i-th test sample, and I(f(x i ≠y i ) is an indicator function. When the prediction result f(x i ) of the model for the i-th sample is not equal to the true label y i , the value of this function is 1, indicating that a diagnostic error has occurred; otherwise, the value of this function is 0.
8. A lithium battery diagnosis system based on multivariable and infrared point cloud, characterized in that, Including: A multi-source data acquisition module, which is used to collect voltage, current, temperature and voiceprint information of the lithium battery through multiple sensors, and obtain infrared three-dimensional point cloud imaging and image information through a monitoring camera; A signal modal decomposition module, which is used to perform modal decomposition on voltage, current, temperature and voiceprint signals by using symplectic geometric modal decomposition (SGMD) to obtain intrinsic mode functions (IMFs), and combining the original signals with the IMFs into a two-dimensional input matrix; A time-series feature extraction module, which is used to process the two-dimensional input matrix by using the Tokenformer model optimized by the CBAM attention mechanism and weighted output the prediction result; A visual feature extraction module, which is used to input the infrared three-dimensional point cloud and image into the YOLOv9 model improved by the ECA attention mechanism to extract image features; A multimodal fusion module, which is used to fuse the prediction result of the Tokenformer and the image features of the YOLOv9 by using the PTP fusion method to generate a two-dimensional fusion matrix; A fault diagnosis module, which is used to perform fault diagnosis on the fusion matrix by using the improved extreme learning machine (IELM) and output the lithium battery safety monitoring result.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the lithium battery diagnosis method based on multivariate and infrared point cloud according to any one of claims 1 to 7.
10. A computer device, characterized in that, Including a memory and a processor, and the memory stores a lithium battery diagnosis method based on multivariate and infrared point cloud according to any one of claims 1 to 7 that can be loaded and executed by the processor.
Citation Information
Cited By
Power battery thermal runaway voiceprint early warning method and system
CN120521709A
Fault diagnosis method for pumped storage unit
CN121278491A
A pumped storage unit fault diagnosis method
CN121278491B
Physically guided deep learning signal modulation identification method and system
CN122339908A
Physical guided deep learning signal modulation recognition method and system
CN122339908B