Monitoring device fault diagnosis method based on multi-modal data feature fusion

CN122885605APending Publication Date: 2026-10-09JIANGSU HONGLUN INFORMATION TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610924686.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-25
Publication Date
2026-10-09

AI Technical Summary

Technical Problem

然而,现有技术存在明显缺陷:首先,严重忽略了蕴含高价值的运维文本日志,多源异构数据间存在巨大的“模态鸿沟”;其次,传统融合方法难以量化时序波形与视觉影像之间深层次的非线性物理耦合关系,无法实现跨模态语义的有效对齐;最后,工业现场传感器极易受电磁或链路干扰导致单边数据失真,现有诊断系统缺乏对跨模态数据一致性的置信度核查机制,极易引发系统的频繁误报

Benefits of technology

1、本发明通过构建设备知识树并结合双向长短期记忆网络与条件随机场架构,精准提取复杂工业实体,进而利用双向编码器表示模型将其转化为高维稠密文本特征向量。该向量作为跨模态融合的语义锚点,为后续物理信号的特征对齐提供了坚实的行业逻辑与认知基准。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122885605A_ABST
    Figure CN122885605A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of electric digital data processing, in particular to a monitoring equipment fault diagnosis method based on multi-modal data feature fusion, which comprises the following steps: constructing a knowledge tree according to a monitoring equipment, performing named entity recognition on operation and maintenance logs based on the knowledge tree by using a bidirectional long short-term memory network and a conditional random field architecture, extracting industrial entities, and encoding and outputting dense text feature vectors; taking the dense text feature vectors as semantic anchor points, projecting time sequence waveforms and infrared thermal image features into a shared embedding space by using a time sequence and a visual encoder, optimizing a mapping matrix by using a particle swarm algorithm combined with information noise contrast estimation loss function, and outputting a cross-modal feature matrix; inputting the cross-modal feature matrix into a mutual information neural network estimation model, estimating a lower bound of a nonlinear semantic correlation measure as a confidence score by using a variational representation; and intercepting fault alarms or matching an abnormal event library to determine a physical fault type. The application can accurately intercept false fault alarms through semantic consistency checking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electronic digital data processing technology, specifically to a fault diagnosis method for monitoring equipment based on multimodal data feature fusion. Background Technology

[0002] With the development of the Industrial Internet of Things and smart manufacturing, the operational reliability of core monitoring equipment is of paramount importance. When equipment experiences performance degradation or failure, it is usually accompanied by synchronous gradual or abrupt changes in multi-physical field characteristics such as electricity and heat, as well as unstructured operation and maintenance logs.

[0003] Currently, equipment fault diagnosis often employs single-modal analysis or simple splicing and fusion of multimodal data. However, existing technologies have significant drawbacks: First, they severely neglect valuable operational log text, resulting in a huge "modal gap" between heterogeneous data from multiple sources; second, traditional fusion methods struggle to quantify the deep nonlinear physical coupling between time-series waveforms and visual images, failing to achieve effective alignment of cross-modal semantics; finally, industrial field sensors are highly susceptible to electromagnetic or link interference, leading to unilateral data distortion, and existing diagnostic systems lack confidence verification mechanisms for cross-modal data consistency, easily causing frequent false alarms.

[0004] Therefore, overcoming the modal gap between multi-source heterogeneous data, accurately measuring the deep nonlinear correlation between cross-modal features, and effectively identifying and eliminating false faults caused by monitoring link interference, thereby reducing the false alarm rate in complex operating conditions, are technical problems that urgently need to be solved in this field.

[0005] To address this, a fault diagnosis method for monitoring equipment based on multimodal data feature fusion is proposed. Summary of the Invention

[0006] The purpose of this invention is to provide a fault diagnosis method for monitoring equipment based on multimodal data feature fusion, which can accurately intercept false fault alarms through cross-modal semantic consistency verification.

[0007] To achieve the above objectives, the present invention provides the following technical solution: A fault diagnosis method for monitoring equipment based on multimodal data feature fusion includes: A knowledge tree is constructed based on the monitoring equipment. Based on the knowledge tree, a bidirectional long short-term memory network and a conditional random field architecture are used to perform named entity recognition on the operation and maintenance logs. The industrial entities corresponding to the knowledge tree are extracted and input into the bidirectional encoder representation model for vectorization encoding, and dense text feature vectors are output. Using dense text feature vectors as semantic anchors, the temporal waveform and infrared thermal image features collected within the same macroscopic time window are projected into the shared embedding space by the temporal encoder and visual encoder respectively. The mapping matrix is ​​optimized by using particle swarm optimization combined with information noise contrast estimation loss function, so that the temporal waveform and infrared thermal image features gather towards the semantic anchor in the shared embedding space, and the cross-modal feature matrix is ​​output. The cross-modal feature matrix is ​​input into the mutual information neural network estimation model, and the lower bound of the nonlinear semantic correlation measure between the time series waveform and the infrared thermal image features is estimated using variational representation, which serves as the confidence score for cross-modal semantic consistency. When the confidence score is lower than a preset safety threshold, it is determined that there is external interference in the monitoring link and the fault alarm is blocked; When the confidence score is not lower than the preset safety threshold, the cross-modal feature matrix is ​​matched with the abnormal event database to determine the physical fault type of the monitoring equipment.

[0008] Preferably, constructing the knowledge tree based on the monitoring equipment includes: extracting equipment components, fault phenomena, and maintenance measures of the monitoring equipment as core categories; establishing hierarchical mapping relationships between the core categories; establishing a tree-like logical topology structure between the equipment components, the fault phenomena, and the maintenance measures; and generating the knowledge tree.

[0009] Preferably, extracting industrial entities includes: capturing the contextual semantic features of the operation and maintenance logs in both positive and negative directions using the bidirectional long short-term memory network; By using the global state transition probability matrix of the conditional random field architecture to logically constrain the contextual semantic features, the optimal annotation sequence that conforms to the knowledge tree topology is output, thus obtaining the industrial entity.

[0010] Preferably, the output of the dense text feature vector includes: performing word embedding encoding, position embedding encoding, and sentence fragment embedding encoding on the industrial entities respectively to obtain corresponding word embedding vectors, position embedding vectors, and sentence fragment embedding vectors; superimposing the word embedding vectors, the position embedding vectors, and the sentence fragment embedding vectors to generate a composite input matrix; inputting the composite input matrix into the transformer encoder of the bidirectional encoder representation model for nonlinear feature mapping and semantic weight aggregation, and outputting the dense text feature vector.

[0011] Preferably, the output of the cross-modal feature matrix includes: extracting features from the time-series waveform using the time-series encoder to generate a time-series feature vector; and extracting features from the infrared thermal image using the visual encoder to generate a visual feature vector. Based on a dual-scale sliding time window mechanism, the temporal feature vectors and visual feature vectors that are within the same macroscopic time window and have aligned timestamps are paired as positive samples and projected onto the shared embedding space respectively. The information noise contrast estimation loss function is constructed, with feature vector pairs under the same fault state as positive samples and feature vector pairs under different fault states as negative samples. By minimizing the information noise contrast estimation loss function, positive samples move closer to the semantic anchor in the shared embedding space, while negative samples repel each other in the shared embedding space. The mapping matrix is ​​updated and the cross-modal feature matrix is ​​concatenated and output.

[0012] Preferably, the mapping matrix is ​​optimized by using a particle swarm optimization algorithm combined with information-noise contrast estimation of the loss function, including: A particle swarm is constructed by initializing multiple particles. The position vector of each particle is mapped to the temperature hyperparameter and learning rate of the information-noise contrast estimation loss function. The convergence error value of the information noise contrast estimation loss function is set as the fitness evaluation function of the particle; Calculate the current fitness value of each particle one by one, record the individual optimal position in the historical trajectory of a single particle, and record the global optimal position of the entire particle swarm. Based on the inertia weight, acceleration factor, random number, and the individual optimal position and the global optimal position, the velocity vector and position vector of each particle are iteratively calculated and updated; The iteration terminates when the fitness value at the global optimal position meets the preset convergence accuracy condition. The absolute optimal parameters at the global optimal position at this time are extracted and parsed. The absolute optimal parameters are then substituted into the information noise contrast estimation loss function to complete the optimization iteration of the mapping matrix.

[0013] Preferably, the mutual information neural network estimation model includes a statistical evaluation network; calculating the confidence score includes: extracting aligned time-series waveform features and infrared thermal image features from the cross-modal feature matrix; and inputting the time-series waveform features and infrared thermal image features into the statistical evaluation network to obtain the network output value; Calculate the first expected value of the network output value under the joint probability distribution of the temporal waveform features and the infrared thermal image features; calculate the second expected value of the network output value under the product of the independent marginal probability distributions of the temporal waveform features and the infrared thermal image features; calculate the difference between the natural logarithm of the first expected value and the second expected value according to the Donsk-Varazan variational representation; solve for the supremum of the difference through parameter iteration optimization, and use the supremum as the lower bound of the nonlinear semantic association measure; use the value of the lower bound of the nonlinear semantic association measure as the confidence score of the cross-modal semantic consistency.

[0014] Preferably, the physical fault types include: high-energy discharge, low-energy discharge, partial discharge, high-temperature overheating, medium-temperature overheating, and low-temperature overheating.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention accurately extracts complex industrial entities by constructing a device knowledge tree and combining it with a bidirectional long short-term memory network and a conditional random field architecture. Then, it uses a bidirectional encoder representation model to transform these entities into high-dimensional dense text feature vectors. These vectors serve as semantic anchors for cross-modal fusion, providing a solid industry logic and cognitive benchmark for subsequent feature alignment of physical signals.

[0016] 2. This invention introduces a particle swarm optimization algorithm to globally adaptively optimize the hyperparameters of the information-noise contrast estimation loss function, driving temporal waveforms and infrared thermal imaging features to highly aggregate towards semantic anchors in a shared embedding space. This mechanism achieves statistical alignment of temporal and visual features in the shared embedding space through contrastive learning, making the representations of the two modalities under the same fault state close to each other in the vector space. This establishes a cross-modal statistical correlation and improves the representation accuracy of the multimodal feature matrix.

[0017] 3. This invention constructs a highly robust cross-modal semantic consistency verification defense line, reducing the false alarm rate of the system under complex operating conditions. By introducing a mutual information neural network estimation model, variational representation is used to rigorously quantify the nonlinear correlation measure between temporal and visual features at the mathematical level. Based on confidence scoring, it can accurately identify single-modal logical conflicts caused by electromagnetic interference or link distortion, precisely intercept false fault alarms, and endow the decision engine with powerful anti-false alarm and anti-spoofing capabilities. Attached Figure Description

[0018] Figure 1 This is a flowchart of the fault diagnosis method for monitoring equipment based on multimodal data feature fusion proposed in this invention; Figure 2 This is a flowchart illustrating the fault diagnosis method for monitoring equipment based on multimodal data feature fusion proposed in this invention. Figure 3 This is a structural schematic diagram of the physical fault type proposed in this invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Example 1 Please see Figures 1 to 3 This invention provides a fault diagnosis method for monitoring equipment based on multimodal data feature fusion, the technical solution of which is as follows: Fault diagnosis methods for monitoring equipment based on multimodal data feature fusion, such as Figure 1 - Figure 2 As shown, it includes: A knowledge tree is constructed based on the monitoring equipment. Based on the knowledge tree, a bidirectional long short-term memory network and a conditional random field architecture are used to perform named entity recognition on the operation and maintenance logs. The industrial entities corresponding to the knowledge tree are extracted and input into the bidirectional encoder representation model for vectorization encoding, and dense text feature vectors are output. Using dense text feature vectors as semantic anchors, the temporal waveform and infrared thermal image features collected within the same macroscopic time window are projected into the shared embedding space by the temporal encoder and visual encoder respectively. The mapping matrix is ​​optimized by using particle swarm optimization combined with information noise contrast estimation loss function, so that the temporal waveform and infrared thermal image features gather towards the semantic anchor in the shared embedding space, and the cross-modal feature matrix is ​​output. The cross-modal feature matrix is ​​input into the mutual information neural network estimation model, and the lower bound of the nonlinear semantic correlation measure between the time series waveform and the infrared thermal image features is estimated using variational representation, which serves as the confidence score for cross-modal semantic consistency. When the confidence score is lower than a preset safety threshold, it is determined that there is external interference in the monitoring link and the fault alarm is blocked; When the confidence score is not lower than the preset safety threshold, the cross-modal feature matrix is ​​matched with the abnormal event database to determine the physical fault type of the monitoring equipment.

[0021] Furthermore, constructing the knowledge tree based on the monitoring equipment includes: extracting equipment components, fault phenomena, and maintenance measures of the monitoring equipment as core categories; establishing hierarchical mapping relationships between the core categories; establishing a tree-like logical topology structure between the equipment components, the fault phenomena, and the maintenance measures; and generating the knowledge tree.

[0022] Specifically, the process of extracting equipment components, fault phenomena, and maintenance measures from the monitoring equipment, as a core category, is as follows: Historical text data from the operation and maintenance management system corresponding to the monitoring equipment is acquired. This historical text data includes historical fault work orders and equipment maintenance manuals. A natural language processing toolkit based on dependency parsing is used to extract and classify entities from the historical text data. Specifically, a part-of-speech tagger is used to identify nouns and verbs in the text, and nouns are extracted as candidate nodes. Dependency parsing is used to construct dependency arcs between words. When a candidate noun has a subject-predicate relationship with the core verb and its hypernym belongs to the equipment structural vocabulary, the noun is classified into the equipment component node set. When a candidate noun has a verb-object relationship or a modifying relationship with the main degenerate verb, the noun is classified into the fault phenomenon node set. When a candidate noun exists within a predefined maintenance verb modification structure, the noun is categorized into the maintenance measure node set. This constitutes the core category.

[0023] The process of establishing a tree-like logical topology structure between the equipment components, the fault phenomena, and the maintenance measures, and generating the knowledge tree, is as follows: The logical framework of the knowledge tree is constructed using a web ontology language, and a directed graph structure is established. , where the set of nodes edge set This indicates the causal or hierarchical relationship path between different core category nodes; The computer-quantified execution process is as follows: A knowledge tree is constructed using a graph data structure; directed connections between nodes in the knowledge tree are extracted; and a preset category state transition matrix corresponding to the knowledge tree is constructed. For any two nodes, classify... and If the knowledge tree contains a... point to If a reasonable physical causal relationship exists (e.g., from "equipment component" to "fault phenomenon"), then the matrix elements will be... Assign the value 1 to the state connectivity scalar; if no reasonable physical association exists, then set the matrix elements to 1. The value is assigned to the minimum penalty constant. .

[0024] The preset category state transition matrix This is directly superimposed onto the global state transition probability matrix of the subsequent conditional random field architecture. This is achieved through the minimal penalty constant. The masking constraint forces the blocking of entity transfer paths that are illogical in terms of physical common sense, ensuring that the optimal output annotation sequence is strictly controlled by the topology of the knowledge tree.

[0025] This invention constructs a device knowledge tree and transforms it into a state transition matrix, which is then directly superimposed onto the conditional random field architecture as a priori constraint. This mechanism uses a minimal penalty constant to forcibly shield unreasonable entity transition paths that violate industry physics common sense, achieving a deep integration of industrial expert experience and deep learning models, thereby improving the logical accuracy and reliability of named entity recognition in complex operation and maintenance texts.

[0026] Further, extracting industrial entities includes: using the bidirectional long short-term memory network to capture the contextual semantic features of the operation and maintenance logs in both positive and negative directions; By using the global state transition probability matrix of the conditional random field architecture to logically constrain the contextual semantic features, the optimal annotation sequence that conforms to the knowledge tree topology is output, thus obtaining the industrial entity.

[0027] Specifically, the quantization process of capturing the contextual semantic features of the operation and maintenance logs using the bidirectional long short-term memory network in both positive and negative directions is as follows: The operation and maintenance logs are segmented and serialized to map a basic word vector sequence. A standard bidirectional long short-term memory network model is used to extract features from the basic word vector sequence layer by layer, and the output is an emission probability matrix that reflects the probability that each word belongs to each target entity label in a specific context.

[0028] The process of using the global state transition probability matrix of the conditional random field architecture to logically constrain the contextual semantic features and outputting the optimal annotation sequence that conforms to the knowledge tree topology is as follows: In the sequence scoring and decoding stage of the conditional random field architecture, the preset category state transition matrix (i.e., the matrix containing the state connectivity scalar and the minimum penalty constant) generated based on the knowledge tree is directly used as the initial weight of the global state transition probability matrix. By combining the emission probability matrix and the global state transition probability matrix, the global score of all possible labeled paths of the current input statement is calculated; The Viterbi dynamic programming algorithm is used to find the label sequence with the maximum global score, which is then considered the optimal label sequence. During this optimization process, due to the effect of the minimum penalty constant, the global score of any entity state transition path that violates the logical association of the knowledge tree will be set to zero. Thus, the Viterbi algorithm automatically prunes and removes them from the search tree, thereby ensuring that the final parsed sequence of industrial entities strictly conforms to physical common sense and industry logic.

[0029] This invention deeply integrates the knowledge tree transition matrix and the conditional random field architecture, combined with the Viterbi algorithm for sequence decoding. By introducing a minimal penalty constant, it automatically prunes and eliminates entity transition paths that violate physical common sense during optimization. This mechanism ensures that the extracted industrial entities strictly conform to industry logic, eliminates common-sense extraction errors, and improves the accuracy and reliability of named entity recognition in complex operation and maintenance logs.

[0030] Further, outputting the dense text feature vector includes: performing word embedding encoding, position embedding encoding, and sentence fragment embedding encoding on the industrial entities respectively to obtain corresponding word embedding vectors, position embedding vectors, and sentence fragment embedding vectors; superimposing the word embedding vectors, the position embedding vectors, and the sentence fragment embedding vectors to generate a composite input matrix; inputting the composite input matrix into the transformer encoder of the bidirectional encoder representation model for nonlinear feature mapping and semantic weight aggregation, and outputting the dense text feature vector.

[0031] Specifically, the process of performing word embedding encoding, position embedding encoding, and sentence fragment embedding encoding on the industrial entities, and generating a composite input matrix, is implemented using computer quantization as follows: A bidirectional encoder representation model was pre-trained on a large-scale industrial power corpus. This model features a 12-layer transformer encoder, 768-dimensional hidden layers, and 12 multi-head self-attention heads. Full-domain fine-tuning was performed using a professional power corpus consisting of 50,000 text paragraphs. The training epochs during fine-tuning were set to 10, the batch size to 32, and the learning rate to [missing information]. The fine-tuned bidirectional encoder representation model is specifically designed to extract and output the dense text feature vector with a fixed 768-dimensional spatial dimension.

[0032] The bidirectional encoder represents the model's built-in preset industrial vocabulary, converting the input industrial entity sequence into a single-hotspot code, and then mapping it to a matrix with a dimension equal to the sequence length multiplied by a fixed dimension through text embedding matrix multiplication. The word embedding matrix; based on the absolute position sequence number of each character in the industrial entity in the input text sequence, index out the characters with the same dimension (sequence length multiplied by a fixed dimension) in the preset position encoding matrix. The position embedding matrix is ​​used; since the input maintenance log is a single-sentence text structure, a sentence fragment embedding matrix is ​​constructed with the exact same shape as the word embedding matrix and all elements are 0; the word embedding matrix, position embedding matrix, and sentence fragment embedding matrix are superimposed by matrix element addition at corresponding positions to generate a composite input matrix. The composite input matrix is ​​then input into the transformer encoder of the bidirectional encoder representation model for nonlinear feature mapping and semantic weight aggregation, and the execution process is as follows: The composite input matrix is ​​input into the converter encoder, which is composed of stacked multi-layer converters. The multi-head self-attention mechanism inside the converter encoder is used to calculate the feature association weights between any two characters in the input sequence in parallel, thereby capturing long-range semantic dependencies. The attention weight matrix output by the multi-head self-attention mechanism performs a fully connected linear transformation on the input features, and then performs nonlinear feature mapping using a concatenated activation function to achieve dynamic aggregation of semantic weights. Finally, it outputs a dense text feature vector that can represent the deep industry semantic information of the industrial entity with high fidelity.

[0033] This invention achieves semantic weight aggregation by explicitly defining the matrix quantization and superposition method of word, position, and fragment embedding, and introducing a multi-head self-attention mechanism. This mechanism can break free from the constraints of traditional word frequency statistics, automatically capturing long-range semantic dependencies and deep causal relationships within industrial entities and their contexts. This results in dense text feature vectors with extremely high-dimensional linear separability and industry semantic accuracy, laying a solid foundation for the efficient alignment and accurate fusion of subsequent cross-modal features through discrete feature digitization.

[0034] Further, outputting a cross-modal feature matrix includes: using the time encoder to extract features from the time-series waveform to generate a time-series feature vector; and using the visual encoder to extract features from the infrared thermal image to generate a visual feature vector. Based on a dual-scale sliding time window mechanism, the temporal feature vectors and visual feature vectors that are within the same macroscopic time window and have aligned timestamps are paired as positive samples and projected onto the shared embedding space respectively. The information noise contrast estimation loss function is constructed, with feature vector pairs under the same fault state as positive samples and feature vector pairs under different fault states as negative samples. By minimizing the information noise contrast estimation loss function, positive samples move closer to the semantic anchor in the shared embedding space, while negative samples repel each other in the shared embedding space. The mapping matrix is ​​updated and the cross-modal feature matrix is ​​concatenated and output.

[0035] Specifically, the network architecture execution process for feature extraction using the temporal encoder and the visual encoder is as follows: The temporal encoder employs a dilated convolutional network with three layers of one-dimensional convolutions. The kernel size of the first, second, and third layers is set to 3, and their dilation rates are set to 1, 2, and 4, respectively. By cascading dilated convolutions, the receptive field of the temporal series is expanded. Finally, a global average pooling layer outputs a one-dimensional temporal feature vector with a dimension of 256. The visual encoder adopts a standard Swin-Transformer-Base architecture. The input is an infrared thermal image. Spatial morphological features are extracted through a moving window self-attention mechanism. Finally, a fully connected layer outputs a one-dimensional visual feature vector with a dimension of 1024.

[0036] The positive sample pairing and projection based on the dual-scale sliding time window mechanism is executed as follows: A macroscopic time window is set. The macroscopic time window serves as the primary observation period. The time span is consistent with the single data update cycle of the online monitoring device for dissolved gases in transformer oil; the time-series waveform is truncated to the current macroscopic time window. The gas concentration variation trend characteristics within the region are output by the time encoder as a time-series feature vector characterizing the slow-change chemical accumulation state; within the same macroscopic time window Internally, microscopic sliding time windows are opened in parallel. The infrared thermal image sequence transmitted back by the inspection robot is dynamically received; each of the aforementioned microscopic sliding time windows is extracted. The temperature rise gradient of the internal infrared thermogram, and within the macroscopic time window. When closed, the infrared thermal image frame corresponding to the maximum temperature rise gradient within the entire window is extracted, and the visual feature vector is output by the visual encoder to characterize the extreme thermal field anomaly; the same macroscopic time window is used. The corresponding temporal feature vector and the visual feature vector are paired across scales for causal alignment.

[0037] The duration of the macroscopic time window is set to 15 minutes, consistent with the single chromatographic analysis cycle of the online dissolved gas monitoring device in transformer oil. The width of the microscopic sliding time window is set to 1 minute, and the sliding step size is set to 30 seconds. Within the 15-minute macroscopic time window, up to 29 microscopic sliding windows are adaptively extracted using a sliding segmentation algorithm, and the temperature rise gradient value of the infrared thermogram is extracted within each microscopic window. When the macroscopic time window closes, the maximum value among the 29 temperature rise gradient values ​​within the entire window is retrieved, and the infrared thermogram frame corresponding to this maximum value is extracted and output as a visual feature vector via the visual encoder.

[0038] The mathematical quantization process of constructing the information-noise contrast estimation loss function and updating the mapping matrix by minimizing the loss function is as follows: Let the number of samples collected within the same macroscopic time window at the same timestamp be . In this embodiment, the fixed spatial dimension of the shared embedded space... The dimension is set to 768, which is completely consistent with the hidden layer dimension of the dense text feature vector output by the bidirectional encoder representation model, thereby achieving direct dimensional alignment of the text semantic anchor space, temporal latent space and visual latent space.

[0039] Accordingly, the time-series mapping matrix is ​​defined as follows: Its matrix shape is The visual mapping matrix is ​​defined as follows: Its matrix shape is .

[0040] Let the dimension be The time-series feature vector of dimension is obtained through the time-series mapping matrix. Perform a linear projection transformation to output the temporal projection vector after entering the shared embedding space. (Its dimension is 768); let the dimension be... The visual feature vector of dimension is passed through the visual mapping matrix. Perform a linear projection transformation to output the visual projection vector after entering the shared embedding space. (Its dimension is 768); In order to enable the dense text feature vector, which also has a 768-dimensional original hidden layer, to seamlessly serve as a centripetal constraint in the latent space, in this embodiment, the dense text feature vector does not require additional spatial projection transformation and is directly used as a semantic anchor point in the shared embedding space, denoted as... (Its dimensions are 768), where This indicates the index number of the feature pair of the current sample. ; To promote the unification of temporal and visual features towards the semantic anchor in the shared embedding space. The information noise contrast estimation loss function is aggregated from time-to-text contrast loss components. Visual-text contrast loss component Together they are constructed, and their overall mathematical formula is defined as follows: ; ; ; In the formula, Represents the computation of vectors with vector Dot product cosine similarity between them; The absolute optimal temperature hyperparameter is determined through iterative optimization using the aforementioned particle swarm optimization algorithm. The index number for traversing text anchor points within the current feature batch; The calculated total loss function value; During the model training phase, the total loss function value is minimized in the reverse direction by using the gradient descent algorithm in conjunction with the absolutely optimal learning rate determined by the aforementioned particle swarm optimization algorithm. The matrix weights of the temporal mapping matrix and the visual mapping matrix are dynamically updated through the error backpropagation mechanism.

[0041] When the total loss function value When convergence reaches the preset error accuracy, lock the updated temporal mapping matrix and the visual mapping matrix; then, transfer the trained temporal projection vectors... With visual projection vector The matrix is ​​concatenated column by column according to the feature dimensions, and the final output is the cross-modal feature matrix that can simultaneously and faithfully map the three-dimensional correlation information of electricity, heat and literature.

[0042] This invention constructs a bidirectional text-guided information-noise contrast estimation loss function formula, explicitly using textual semantic anchors as centripetal constraints for temporal and visual mapping. This mechanism, leveraging cosine similarity quantification and adaptive temperature hyperparameters, powerfully narrows the distance between heterogeneous features representing the same fault state in the latent space, fundamentally overcoming the modal gap problem of inconsistent mathematical expressions among electrical, thermal, and textual data sources, and significantly broadening the collaborative representation capability of feature matrices for complex and weak defects.

[0043] Furthermore, the mapping matrix is ​​optimized by combining the particle swarm optimization algorithm with information-noise contrast estimation of the loss function, including: A particle swarm is constructed by initializing multiple particles. The position vector of each particle is mapped to the temperature hyperparameter and learning rate of the information-noise contrast estimation loss function. The convergence error value of the information noise contrast estimation loss function is set as the fitness evaluation function of the particle; Calculate the current fitness value of each particle one by one, record the individual optimal position in the historical trajectory of a single particle, and record the global optimal position of the entire particle swarm. Based on the inertia weight, acceleration factor, random number, and the individual optimal position and the global optimal position, the velocity vector and position vector of each particle are iteratively calculated and updated; The iteration terminates when the fitness value at the global optimal position meets the preset convergence accuracy condition. The absolute optimal parameters at the global optimal position at this time are extracted and parsed. The absolute optimal parameters are then substituted into the information noise contrast estimation loss function to complete the optimization iteration of the mapping matrix.

[0044] Specifically, the computer execution process of setting the convergence error value of the information noise contrast estimation loss function as the particle fitness evaluation function is as follows: A separate validation dataset is pre-defined from the multimodal dataset. The temperature hyperparameters and learning rate derived from the current particle's position vector are substituted into the information-noise contrast estimation loss function. The weights of the current temporal mapping matrix and visual mapping matrix remain unchanged. Forward propagation is performed on the validation dataset, and the total loss function value on the validation set is output. This total loss function value is established as the fitness value of the particle. The smaller the total loss function value, the higher the cross-modal feature aggregation degree under the guidance of this set of hyperparameters.

[0045] The quantitative mathematical execution process of iteratively calculating and updating the velocity vector and position vector of each particle based on inertia weight, acceleration factor, random number, and the individual optimal position and the global optimal position is as follows: The position vector of each particle is set as a two-dimensional space vector. ,in Indicates the first The particle in the first Temperature hyperparameter at the next iteration Indicates the first The particle in the first The learning rate at each iteration; setting the velocity vector of each particle as a two-dimensional velocity vector. ; The velocity and position components of each dimension are iteratively updated alternately using the following formula: ; ; In the formula, Indicates component dimension index and The set of values ​​is ; and Indicates a closed interval Independently generated random numbers; Indicates the first The optimal position of each particle in its history Dimensional components; The position representing the global optimal position in the entire history of the particle swarm. Dimensional components; To prevent the optimization process from getting lost in local optima, the inertial weights are... The update is performed using a linear differential decreasing mechanism, with the first acceleration factor... With the second acceleration factor The update is performed using a linear adjustment mechanism, and the specific calculation formulas are as follows: ; ; ; In the formula, This represents the preset maximum value of the inertia weight. This represents the preset minimum inertia weight; This represents the initial value of the first acceleration factor. This represents the final value of the first acceleration factor; This represents the initial value of the second acceleration factor. This represents the final value of the second acceleration factor; This indicates the preset maximum number of iterations.

[0046] The iteration terminates when the fitness value at the global optimum position satisfies a preset convergence accuracy condition, wherein the preset convergence accuracy condition refers to the current iteration number. Reaching the maximum number of iterations .

[0047] The search range for the temperature hyperparameter is set to [0.01, 1.0], and the search range for the learning rate is set to [1×10⁻⁻⁶]. 5 [1×10⁻²]. The first dimension of particle position is mapped to the temperature hyperparameter search range using linear mapping, and the second dimension is mapped to the learning rate search range. The particle swarm size is set to 30 particles. The maximum value of the inertia weight is... = 0.9, minimum inertia weight = 0.4, initial value of the first acceleration factor Final value Initial value of the second acceleration factor Final value Maximum iteration algebra =100.

[0048] This invention constructs a bidirectional text-guided information-noise contrast estimation loss function and utilizes an improved particle swarm optimization algorithm with linearly decreasing inertial weights and linearly adjusted acceleration factors for hyperparameter adaptive global optimization. This mechanism eliminates the blindness of manual empirical parameter tuning, achieves fast and accurate convergence of temperature hyperparameters and learning rate in a non-convex high-dimensional space, and enhances the centripetal aggregation and representation accuracy of cross-modal feature matrices in the shared embedding space.

[0049] Furthermore, the mutual information neural network estimation model includes a statistical evaluation network; calculating the confidence score includes: extracting aligned temporal waveform features and infrared thermal image features from the cross-modal feature matrix; and inputting the temporal waveform features and infrared thermal image features into the statistical evaluation network to obtain the network output value. Calculate the first expected value of the network output value under the joint probability distribution of the temporal waveform features and the infrared thermal image features; calculate the second expected value of the network output value under the product of the independent marginal probability distributions of the temporal waveform features and the infrared thermal image features; calculate the difference between the natural logarithm of the first expected value and the second expected value according to the Donsk-Varazan variational representation; solve for the supremum of the difference through parameter iteration optimization, and use the supremum as the lower bound of the nonlinear semantic association measure; use the value of the lower bound of the nonlinear semantic association measure as the confidence score of the cross-modal semantic consistency.

[0050] Specifically, the mutual information neural network estimation model includes a statistical evaluation network, and its network architecture and mathematical quantification process are as follows: The statistical evaluation network adopts a multilayer perceptron architecture, denoted as , with parameter space . And the network weight is Continuously differentiable functions To ensure that the statistical evaluation network can accurately approximate the lower bound of the Donsk-Warahan variant of mutual information, while taking into account the stability of training and the smooth flow of gradients, this embodiment imposes the following quantitative constraints on the network topology parameters of the multilayer perceptron: The input dimension of the input layer is fixed at 1536 dimensions. This dimension is naturally constructed by lossless feature concatenation of the temporal projection vector (768 dimensions) aligned in the 768-dimensional shared embedding space and the visual projection vector (768 dimensions).

[0051] Two fully connected hidden layers are configured. The first hidden layer contains 512 hidden units, followed by a LeakyReLU activation function (with a negative slope coefficient set to 0.01) and a batch normalization layer. The second hidden layer contains 128 hidden units, also configured with the LeakyReLU activation function, to achieve compression and interactive mapping of nonlinear high-dimensional semantic features layer by layer.

[0052] The output layer contains one output unit, which does not have any external nonlinear activation function and directly outputs a real scalar value to characterize the boundary evaluation score of the joint distribution sample or the marginal distribution sample under the current network weights.

[0053] The computer discrete empirical sampling calculation process for calculating the first expected value and the second expected value is as follows: In each training iteration, the sample batch size extracted from the cross-modal feature matrix is ​​[size missing]. Batch of feature data; For joint probability distribution sampling, the time-series waveform features that are synchronously paired at the same timestamp are directly extracted. With infrared thermal imaging features As joint sample pairs; For sampling the product of independent marginal probability distributions, the temporal waveform characteristics are preserved. The order remains unchanged, and the sequence indexes of the infrared thermal imaging features within the feature data batch are randomly shuffled to obtain mismatch features. To form edge sample pairs; where, For each joint sample pair, the independent traversal index number is... For each edge sample pair, the independent traversal index number is... This refers to the mismatch index number generated by mapping after randomly shuffling the visual feature sequences within the current batch; and The range of values ​​is all ; Based on the Donsk-Varazan variational representation, calculate the difference of the natural logarithms, and its corresponding discrete empirical loss function. The formula is defined as: ; In the formula, the discrete mean expression of the first term represents the first expected value under the joint probability distribution, and the discrete mean expression within the logarithm of the second term represents the second expected value under the product of independent marginal probability distributions. This represents the network prediction scalar value output after inputting the input sample into the statistical evaluation network; Represents an exponential function with the natural constant as its base; The process of finding the supremum of the difference through parameter iteration optimization is as follows: An adaptive moment estimation gradient ascent algorithm is employed to maximize the discrete empirical loss function. The value of is the sole optimization objective in the parameter space. Iteratively update the network weights along the positive gradient direction. When the gradient ascent converges, extract the discrete empirical loss function at that convergence moment. The maximum calculated value is used as the lower bound of the nonlinear semantic association measure that approximates the true physical mutual information value.

[0054] This invention transforms abstract mutual information theory into a discrete empirical computation process that can be precisely quantified and executed by a computer by introducing statistical evaluation networks and variational representations. Through the pairing and random shuffling of feature batches, it cleverly achieves unbiased estimation of joint and marginal distributions. This mechanism does not require a pre-set complex probability distribution prior model; it directly optimizes and approximates the strict mathematical lower bound of cross-modal semantic association using gradient algorithms, providing a robust and highly confident anti-spoofing metric.

[0055] Furthermore, such as Figure 3As shown, the physical fault types include: high-energy discharge, low-energy discharge, partial discharge, high-temperature overheating, medium-temperature overheating, and low-temperature overheating.

[0056] Specifically, the abnormal event database pre-stores feature mapping rules that correspond one-to-one with the physical fault types, and the computer matching and execution process is as follows: The volume fraction of the characteristic gas is parsed from the cross-modal feature matrix, the component proportion is calculated and the three ratios are encoded to generate a characteristic gas ratio combination matrix; when the cross-modal feature matrix is ​​matched with the abnormal event database, the feature mapping rule is retrieved according to the characteristic gas ratio combination matrix. When the characteristic gas ratio combination matrix shows that both the acetylene volume fraction and the hydrogen gas integral are higher than the first preset concentration limit, and the infrared thermal image features show that the temperature rise gradient exceeds the limit, the current state of the monitoring device is matched to the specific classification of the high-energy discharge or the low-energy discharge; the temperature rise gradient exceeding the limit refers to the maximum value of the infrared thermal image temperature rise gradient exceeding the preset temperature rise gradient threshold (set to 5℃ / min in this embodiment).

[0057] When the characteristic gas ratio combination matrix shows a significant surge in the hydrogen gas integral and methane volume fraction exceeding the third preset concentration limit, and the acetylene volume fraction is below the safety judgment lower limit, and the infrared thermal image features show that they are in the normal temperature range without a significant temperature rise gradient, the current state of the monitoring device is matched to the specific classification of the partial discharge. When the characteristic gas ratio combination matrix shows that the volume fractions of methane and ethylene are both higher than the second preset concentration limit, and the volume fraction of acetylene is lower than the safety judgment lower limit, the current state of the monitoring device is matched to the specific classification of high temperature overheating, medium temperature overheating, and low temperature overheating.

[0058] The abnormal event database pre-stores feature mapping rules that correspond one-to-one with the physical fault types. Based on GB / T 7252 "Guidelines for Dissolved Gas Analysis and Judgment in Transformer Oil" and IEC 60599 standard, this embodiment makes the following quantitative limitations on the core absolute concentration thresholds in the feature mapping rules: The first preset concentration limit (acetylene warning value) is set as follows: ; The second preset concentration limit (ethylene warning value) is set as follows: ; The third preset concentration limit (methane warning value) is set as follows: ; The safety judgment lower limit (acetylene baseline value) is set as follows: .

[0059] This invention establishes a rigorous causal correspondence between abstract physical fault types and precise numerical boundaries in cross-modal feature matrices by deeply embedding standardized feature mapping rules into an anomaly event database. This mechanism overcomes the limitations of single-modal threshold criteria, enabling precise qualitative and graded classification of extremely subtle overheating and abnormal discharge defects, thus endowing the diagnostic system with strong physical interpretability and engineering application value.

[0060] Example 2 This embodiment takes the actual operation monitoring and fault diagnosis scenario of a 220kV main transformer in a heavy-load traction substation as an example to illustrate in detail the specific workflow of the above-mentioned fault diagnosis method for monitoring equipment based on multi-modal data feature fusion: The 220kV main transformer in the traction substation is subjected to frequent train load impacts and high-order harmonic interference for a long time. During a certain operating cycle, the insulation paper inside the high-voltage bushing of phase A of the transformer showed early slight deterioration, and at the same time, the site was exposed to strong traction electromagnetic interference and solar radiation (background temperature interference).

[0061] The system first acquires historical text data from the operation and maintenance management system of the traction substation. This data includes historical defect records such as "Poor grounding of the A-phase bushing end screen," "Bushing overheating," and "Abnormal insulating oil chromatography." The system then constructs a knowledge tree of the equipment using a graph data structure, connecting nodes such as "A-phase bushing" (equipment component), "insulation degradation" (fault phenomenon), and "oil sample testing" (maintenance measures) in a directed manner, and generating values ​​assigned to either a state connectivity scalar of 1 or a minimal penalty constant. A preset category state transition matrix is ​​then used. Subsequently, a bidirectional long short-term memory network and a conditional random field architecture are employed to scan the current operation and maintenance logs. The mask constraint automatically eliminated identification errors that violated physical common sense, such as "sleeve-software crash," and accurately extracted the industrial entity "A-phase sleeve anomaly." Furthermore, this entity was input into a bidirectional encoder representation model, and through feature mapping and semantic weight aggregation via a multi-head self-attention mechanism, a high-dimensional dense text feature vector was output, which served as the "semantic anchor point" for this diagnosis.

[0062] At the same timestamp, the monitoring system simultaneously acquires two physical signals: (1) Timing waveforms: The characteristic gas time series uploaded in real time by the online monitoring device for dissolved gases in transformer oil, and the grounding current of the bushing end screen. The system uses a multi-scale one-dimensional hollow convolutional network (time encoder) to extract the timing feature vector containing long-range evolution trends.

[0063] (2) Infrared thermal image: Infrared image of the A-phase sleeve transmitted back by the inspection robot. The system uses a Swin-Transformer network (visual encoder) to segment local thermal plumes that are not affected by the sunlight background through a moving window attention mechanism, generating visual feature vectors. Entering the cross-modal alignment stage, the system uses a particle swarm optimization algorithm with a linear decreasing inertial weight mechanism to globally adaptively optimize the temperature hyperparameter and learning rate of the information noise contrast estimation loss function. Driven by the optimal parameters, the information noise contrast estimation loss function is minimized, forcing the above "temporal feature vector" and "visual feature vector" to uniformly align with the "semantic anchor point (A-phase sleeve anomaly)" generated in step one in the shared embedding space. After iteratively updating the mapping matrix, the finally spliced ​​output high-fidelity cross-modal feature matrix is ​​output.

[0064] Because the secondary communication cable of the traction substation has been subjected to long-term interference from high-frequency harmonic radiation from the traction power supply, a sudden digital signal distortion occurs in the communication link of the online dissolved gas monitoring device in the transformer oil when it finishes a long-period sampling (e.g., a single analysis cycle of 15 minutes) and transmits a data message to the backend. This causes an abnormal "data spike" (false surge value) to appear in the acetylene data channel received by the backend during the current sampling cycle. Traditional single-mode threshold alarm systems, which only perform limit-breaking logic judgments for a single value, would immediately trigger a critical "internal arc discharge" alarm.

[0065] The present invention inputs the aligned cross-modal feature matrix into the statistical evaluation network of the mutual information neural network estimation model. The system performs sample pairing and random shuffling within the feature batch, and calculates the supremum of the difference between the logarithm of the joint probability expectation and the logarithm of the independent marginal probability expectation based on the Donsk-Varazan variational representation.

[0066] The calculation results show that although the communication link mutation causes a "pseudo-acetylene surge" in the time series characteristics of the current period, the infrared thermal image characteristics transmitted back by the infrared visual encoder in the same period show that they are "within the normal temperature range and without obvious temperature rise gradient." The two exhibit a serious cross-modal logical conflict in terms of physical causality. The lower bound (confidence score) of the nonlinear semantic association measure obtained by the mutual information neural network estimation model is only 0.35, far below the preset safety threshold (e.g., set to 0.85).

[0067] Based on this, the system determined that the sudden change in the current time series data did not correspond to any substantial physical defect evolution, but rather to unilateral digital distortion caused by external harmonic interference encountered by the monitoring communication link. The system then decisively intercepted the false fault alarm, effectively preventing unplanned power outages and human error at the substation.

[0068] After the interference period, as the internal degradation of the bushing progressed, the time-series waveform (the acetylene and hydrogen exhibited a real, physically continuous, slight increase) and the visual infrared thermographic image (the bushing as a whole showed a real temperature rise gradient exceeding the limit) were highly coupled in the latent space. After being calculated again by the mutual information neural network estimation model, the confidence score rose to 0.94, which was higher than the preset safety threshold of 0.85, and the consistency check was successfully passed.

[0069] The system allows the data to pass and matches the characteristic gas ratio combination matrix in the cross-modal characteristic matrix with the abnormal event database. Based on the feature mapping rules, it is found that the acetylene volume fraction and hydrogen gas integral are both consistently and stably higher than the first preset concentration limit, and the infrared temperature rise gradient exceeds the limit. The system ultimately accurately determines the physical fault type of the monitoring equipment as high-energy discharge and triggers the back-end operation and maintenance system to generate a maintenance work order.

[0070] The preset safety threshold is not a blindly set fixed constant; its specific implementation process is as follows: Historical cross-modal feature matrix data of the monitoring equipment under normal operating conditions are collected, and the confidence score distribution under normal operating conditions is calculated by the mutual information neural network estimation model. Simultaneously, the confidence score distribution under known communication link interference conditions is also collected. With the goal of maximizing the F1-score, candidate thresholds are traversed across the confidence score interval [0,1], and the value that achieves the optimal balance between the detection rate under interference conditions and the false interception rate under normal operating conditions is selected as the preset safety threshold. In this embodiment, the safety threshold determined by the above method is 0.85.

[0071] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A fault diagnosis method for monitoring equipment based on multimodal data feature fusion, characterized in that, include: A knowledge tree is constructed based on the monitoring equipment. Based on the knowledge tree, a bidirectional long short-term memory network and a conditional random field architecture are used to perform named entity recognition on the operation and maintenance logs. The industrial entities corresponding to the knowledge tree are extracted and input into the bidirectional encoder representation model for vectorization encoding, and dense text feature vectors are output. Using dense text feature vectors as semantic anchors, the temporal waveform and infrared thermal image features collected within the same macroscopic time window are projected into the shared embedding space by the temporal encoder and visual encoder respectively. The mapping matrix is ​​optimized by using particle swarm optimization combined with information noise contrast estimation loss function, so that the temporal waveform and infrared thermal image features gather towards the semantic anchor in the shared embedding space, and the cross-modal feature matrix is ​​output. The cross-modal feature matrix is ​​input into the mutual information neural network estimation model, and the lower bound of the nonlinear semantic correlation measure between the time series waveform and the infrared thermal image features is estimated using variational representation, which serves as the confidence score for cross-modal semantic consistency. When the confidence score is lower than a preset safety threshold, it is determined that there is external interference in the monitoring link and the fault alarm is blocked; When the confidence score is not lower than the preset safety threshold, the cross-modal feature matrix is ​​matched with the abnormal event database to determine the physical fault type of the monitoring equipment.

2. The fault diagnosis method for monitoring equipment based on multimodal data feature fusion according to claim 1, characterized in that, Constructing the knowledge tree based on the monitoring equipment includes: extracting equipment components, fault phenomena, and maintenance measures of the monitoring equipment as core categories; establishing hierarchical mapping relationships between the core categories; establishing a tree-like logical topology structure between the equipment components, the fault phenomena, and the maintenance measures; and generating the knowledge tree.

3. The fault diagnosis method for monitoring equipment based on multimodal data feature fusion according to claim 1, characterized in that, Extracting industrial entities includes: capturing the contextual semantic features of the operation and maintenance logs in both positive and negative directions using the bidirectional long short-term memory network; By using the global state transition probability matrix of the conditional random field architecture to logically constrain the contextual semantic features, the optimal annotation sequence that conforms to the knowledge tree topology is output, thus obtaining the industrial entity.

4. The fault diagnosis method for monitoring equipment based on multimodal data feature fusion according to claim 1, characterized in that, Outputting a dense text feature vector includes: performing word embedding encoding, position embedding encoding, and sentence fragment embedding encoding on the industrial entities respectively to obtain corresponding word embedding vectors, position embedding vectors, and sentence fragment embedding vectors; superimposing the word embedding vectors, position embedding vectors, and sentence fragment embedding vectors to generate a composite input matrix; inputting the composite input matrix into the transformer encoder of the bidirectional encoder representation model for nonlinear feature mapping and semantic weight aggregation, and outputting the dense text feature vector.

5. The fault diagnosis method for monitoring equipment based on multimodal data feature fusion according to claim 1, characterized in that, Outputting a cross-modal feature matrix includes: extracting features from the time-series waveform using the time-series encoder to generate a time-series feature vector; and extracting features from the infrared thermal image using the visual encoder to generate a visual feature vector. Based on a dual-scale sliding time window mechanism, the temporal feature vectors and visual feature vectors that are within the same macroscopic time window and have aligned timestamps are paired as positive samples and projected onto the shared embedding space respectively. The information noise contrast estimation loss function is constructed, with feature vector pairs under the same fault state as positive samples and feature vector pairs under different fault states as negative samples. By minimizing the information noise contrast estimation loss function, positive samples move closer to the semantic anchor in the shared embedding space, while negative samples repel each other in the shared embedding space. The mapping matrix is ​​updated and the cross-modal feature matrix is ​​concatenated and output.

6. The fault diagnosis method for monitoring equipment based on multimodal data feature fusion according to claim 5, characterized in that, The mapping matrix is ​​optimized by using the particle swarm optimization algorithm combined with information-noise contrast estimation to estimate the loss function, including: A particle swarm is constructed by initializing multiple particles. The position vector of each particle is mapped to the temperature hyperparameter and learning rate of the information-noise contrast estimation loss function. The convergence error value of the information noise contrast estimation loss function is set as the fitness evaluation function of the particle; Calculate the current fitness value of each particle one by one, record the individual optimal position in the historical trajectory of a single particle, and record the global optimal position of the entire particle swarm. Based on the inertia weight, acceleration factor, random number, and the individual optimal position and the global optimal position, the velocity vector and position vector of each particle are iteratively calculated and updated; The iteration terminates when the fitness value at the global optimal position meets the preset convergence accuracy condition. The absolute optimal parameters at the global optimal position at this time are extracted and parsed. The absolute optimal parameters are then substituted into the information noise contrast estimation loss function to complete the optimization iteration of the mapping matrix.

7. The fault diagnosis method for monitoring equipment based on multimodal data feature fusion according to claim 1, characterized in that, The mutual information neural network estimation model includes a statistical evaluation network; the confidence score is calculated by: extracting aligned time-series waveform features and infrared thermal image features from the cross-modal feature matrix; and inputting the time-series waveform features and infrared thermal image features into the statistical evaluation network to obtain the network output value. Calculate the first expected value of the network output value under the joint probability distribution of the temporal waveform features and the infrared thermal image features; calculate the second expected value of the network output value under the product of the independent marginal probability distributions of the temporal waveform features and the infrared thermal image features; calculate the difference between the natural logarithm of the first expected value and the second expected value according to the Donsk-Varazan variational representation; solve for the supremum of the difference through parameter iteration optimization, and use the supremum as the lower bound of the nonlinear semantic association measure; use the value of the lower bound of the nonlinear semantic association measure as the confidence score of the cross-modal semantic consistency.

8. The fault diagnosis method for monitoring equipment based on multimodal data feature fusion according to claim 1, characterized in that, The physical fault types include: high-energy discharge, low-energy discharge, partial discharge, high-temperature overheating, medium-temperature overheating, and low-temperature overheating.