Equipment fault label determination method and device based on multi-modal feature fusion
By extracting the multimodal features of the equipment and performing cross-modal fusion, the problem of incomplete information caused by single-modal data is solved, and high-precision and real-time determination of equipment fault labels is achieved.
Patent Information
- Application Number
- CN202510699475.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-19
AI Technical Summary
In the existing technology, the determination of equipment fault labels mainly relies on single-modal data, which leads to incomplete information, affects accuracy and real-time performance, and cannot effectively capture the multi-dimensional abnormal information of the equipment.
By extracting the multimodal features of the equipment, including textual modal features, structural modal features, and temporal modal features, and fusing them using the cross-modal attention mechanism, a three-dimensional fault portrait is constructed, which is finally input into the pre-trained fault classification model to determine the fault label.
It improves the accuracy and real-time performance of equipment fault label determination, breaks through the limitations of traditional single-modal data analysis through multi-dimensional information fusion, and achieves high-precision and robust fault classification.
Smart Images

Figure CN120671033A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of equipment fault diagnosis, and in particular to a method and apparatus for determining equipment fault labels based on multimodal feature fusion. Background Art
[0002] The stable operation of industrial equipment is the core foundation for ensuring production continuity and safety, especially in key areas such as electricity, manufacturing, and energy. Equipment failure may cause serious economic losses or even safety accidents.
[0003] Currently, existing technologies for determining device fault labels primarily rely on single-modal data (such as vibration signals, current waveforms, or operation and maintenance logs). However, with the development of the Industrial Internet of Things (IIoT), the amount of multimodal data (text, time series, and structured data) generated during device operation is growing exponentially. Relying solely on single-modal data can lead to incomplete and incomplete information. For example, vibration signal analysis can miss structural issues like device communication link anomalies, while semantic analysis of text logs struggles to capture transient abnormal fluctuations in sensor data. Consequently, relying on single-modal data can compromise the accuracy and real-time nature of device fault label determination. Summary of the Invention
[0004] In view of the above problems, the present application provides a method and apparatus for determining equipment fault labels based on multimodal feature fusion, the main purpose of which is to improve the accuracy and real-time performance of equipment fault label determination through multimodal feature fusion.
[0005] To solve the above technical problems, this application proposes the following solutions:
[0006] In a first aspect, the present application provides a method for determining a device fault label based on multimodal feature fusion, the method comprising:
[0007] Extracting multimodal features corresponding to abnormal devices, wherein the multimodal features include text modal features, structural modal features, and temporal modal features;
[0008] The text modal features, the structural modal features, and the temporal modal features are fused using a cross-modal attention mechanism to obtain a final fused feature, wherein the cross-modal attention mechanism is used to characterize the cross-dimensional correlation between the text modal features, the structural modal features, and the temporal modal features;
[0009] The final fusion feature is input into the pre-trained fault classification model to obtain the target fault label corresponding to the abnormal device.
[0010] In a second aspect, the present application provides a device for determining a device fault label based on multimodal feature fusion, the device comprising:
[0011] An extraction unit, configured to extract multimodal features corresponding to abnormal devices, wherein the multimodal features include text modal features, structural modal features, and temporal modal features;
[0012] a fusion unit, configured to fuse the text modal features, the structural modal features, and the temporal modal features obtained by the extraction unit using a cross-modal attention mechanism to obtain a final fused feature, wherein the cross-modal attention mechanism is configured to characterize a cross-dimensional correlation relationship among the text modal features, the structural modal features, and the temporal modal features;
[0013] A determination unit is configured to input the final fusion feature obtained by the fusion unit into a pre-trained fault classification model to obtain a target fault label corresponding to the abnormal device.
[0014] In order to achieve the above-mentioned purpose, according to the third aspect of the present application, a storage medium is provided, which includes a stored program, wherein when the program is running, the device where the storage medium is located is controlled to execute the device fault label determination method based on multimodal feature fusion according to the above-mentioned first aspect.
[0015] In order to achieve the above-mentioned purpose, according to a fourth aspect of the present application, a processor is provided, which is used to run a program, wherein when the program is running, the method for determining a device fault label based on multimodal feature fusion according to the above-mentioned first aspect is executed.
[0016] By means of the above technical solution, the present application provides a method and device for determining equipment fault labels based on multimodal feature fusion. First, the multimodal features corresponding to the abnormal equipment are extracted. The multimodal features include text modal features, structural modal features and time series modal features. Then, the text modal features, structural modal features and time series modal features are fused using a cross-modal attention mechanism to obtain the final fused features. The cross-modal attention mechanism is used to characterize the cross-dimensional correlation between text modal features, structural modal features and time series modal features. Finally, the final fused features are input into a pre-trained fault classification model to obtain the target fault label corresponding to the abnormal equipment. The technical solution provided in this application can cover the multi-dimensional information of equipment failures by extracting multi-modal features such as text modal features, structural modal features and time series modal features of abnormal equipment, and fuse them through the cross-dimensional correlation between text modal features, structural modal features and time series modal features to obtain the final fusion features, so that abnormal equipment can break through the limitations of traditional single-modal data analysis. Based on the cross-dimensional interaction of text-structure-time series features, a three-dimensional fault portrait of the abnormal equipment is constructed. The target fault label corresponding to the final fusion feature is determined by a pre-trained fault classification model, which can ensure the high precision and robustness of fault classification, thereby improving the accuracy and real-time determination of equipment fault labels.
[0017] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present application. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:
[0019] Figure 1 A flow chart of a method for determining a device fault label based on multimodal feature fusion provided by an embodiment of the present application is shown;
[0020] Figure 2 A flowchart of another method for determining a device fault label based on multimodal feature fusion provided by an embodiment of the present application is shown;
[0021] Figure 3 The following is a block diagram showing a device for determining a device fault label based on multimodal feature fusion according to an embodiment of the present application;
[0022] Figure 4A block diagram of another device fault label determination apparatus based on multimodal feature fusion provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0023] The following describes exemplary embodiments of the present application in more detail with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.
[0024] Currently, existing technologies for determining device fault labels primarily rely on single-modal data (such as vibration signals, current waveforms, or operation and maintenance logs). However, with the development of the Industrial Internet of Things (IIoT), the amount of multimodal data (text, time series, and structured data) generated during device operation is growing exponentially. Relying solely on single-modal data can lead to incomplete and incomplete information. For example, vibration signal analysis can miss structural issues like device communication link anomalies, while semantic analysis of text logs struggles to capture transient abnormal fluctuations in sensor data. Consequently, relying on single-modal data can compromise the accuracy and real-time nature of device fault label determination.
[0025] The inventors have discovered through research that, due to the exponential growth of multimodal data (text, time series, structure, etc.) generated during the operation of the equipment, it is possible to abandon the reliance on single modal data and instead use multimodal data for feature extraction, obtaining multimodal features such as text modal features, structural modal features, and time series modal features, and capturing the invisible correlations formed by cross-dimensional interactions between multimodal features to construct a three-dimensional fault profile of the abnormal equipment, and then use the fault classification model to determine the fault label of the final fusion feature. In this way, abnormal equipment breaks through the limitations of traditional single-modal data analysis, and constructs a three-dimensional fault profile of the abnormal equipment based on the cross-dimensional interaction of text-structure-time series features, thereby improving the accuracy and real-time performance of equipment fault label determination.
[0026] Based on the above considerations, the embodiment of the present application provides a method for determining device fault labels based on multimodal feature fusion. This method can improve the accuracy and real-time performance of device fault label determination through multimodal feature fusion. The specific execution steps are as follows: Figure 1 Shown, including:
[0027] 101. Extract multimodal features corresponding to abnormal devices.
[0028] Among them, multimodal features include text modal features, structural modal features and temporal modal features.
[0029] In this step, abnormal equipment refers to equipment that shows significant deviation from the normal state during operation. This deviation may be caused by hardware failure, software error, environmental interference or human error. Multimodal features are extracted from the multimodal data of abnormal equipment. Multimodal data includes equipment log data, equipment topology communication data and equipment status data, etc. Among them, equipment log data is text modal data, equipment topology communication data is structural modal data, and equipment status data is time series modal data. Among them, text modal data is text data obtained from the equipment log system (such as operation and maintenance logs, sensor logs), such as error codes, operation records, sensor status descriptions, etc. Structural modal data is a device communication graph constructed through equipment topology communication data (such as communication links between devices, node connection relationships), while time series modal data is collected equipment status data (such as vibration signals, temperature sensor data, current waveforms, etc.) as time series input.
[0030] For text modal features, the BERT-Transformer hybrid network can be used to segment and encode device log data. BERT captures semantic information (such as the meaning of error codes), and Transformer models long-distance dependencies (such as the time series relationship of log events), thereby obtaining text modal features. For structural modal features, the dynamic GraphSAGE network can be used to construct a device communication graph based on the device topology communication data, and the degree centrality of the node (measures the connection importance of the node in the communication graph) and the mean edge transmission delay (measures the stability of the communication link) are calculated to obtain structural modal features. For temporal modal features, FFT transform can be performed on device status data (such as vibration signals) to extract Mel spectrum features (capture frequency components). The bidirectional LSTM-CRF model is then combined to model temporal dependencies (LSTM captures temporal dynamics, and CRF optimizes sequence labeling) to obtain temporal modal features.
[0031] 102. The cross-modal attention mechanism is used to fuse text modal features, structural modal features and temporal modal features to obtain the final fused features.
[0032] Among them, the cross-modal attention mechanism is used to characterize the cross-dimensional correlation between text modal features, structural modal features and temporal modal features.
[0033] In this step, multiple focus targets are predefined (e.g., 8 heads, the specific number of heads can be customized according to actual conditions), and each focus target corresponds to a cross-dimensional association task (e.g., text-structure association, structure-time series association, text-time series association). For example, heads 1-2: focus on structure-time series association (e.g., the historical association probability between device ID and vibration frequency, such as the probability of "transformer A+200Hz->winding fault" is 0.85), heads 3-5: analyze text-time series co-occurrence (e.g., the co-occurrence score of the log "high temperature" and the energy above 200Hz in the spectrum is 0.91). Heads 6-8: mine text-structure semantic associations (e.g., the cosine similarity of the semantic vector of "new transformer" and the GNN structure encoding is 0.82). Through cross-dimensional association tasks, the invisible associations of different modal features when interacting across dimensions are captured, and the interactive association features corresponding to each aggregation target are obtained.
[0034] By splicing the interactive correlation features corresponding to the above-mentioned multi-head focus targets and combining them with the preset modal weights of different modal features, we can obtain the final fusion features, namely the cross-modal correlation features, which contain multi-dimensional interactive information of "text-structure-time".
[0035] It should be noted that, since there may be multiple business scenarios in actual applications, the importance of different modal features to different business scenarios varies significantly. For example, in the fault diagnosis mode, since the failure of power equipment is usually directly related to the physical state (such as abnormal vibration, sudden temperature change) and the stability of the communication link (such as topology interruption), the text log may only serve as auxiliary information. In other words, the key modes are: structural mode (equipment communication topology) and timing mode (vibration / temperature signal), and the non-key mode is: text mode. Therefore, in order to enable the final fusion feature to accurately capture the key business features and thus improve the accuracy of subsequent fault label determination, the modal weights preset for the above-mentioned different modal features can be set in combination with the business scenario, that is, different modal weight distributions are set for different business scenarios, and arbitration of the modal weights is achieved by matching business scenarios. For example, in the steady-state operation mode, the weights are evenly distributed, that is, the text modal weight is 0.4, the structural modal weight is 0.3, and the timing modal weight is 0.3. In the fault diagnosis mode, the structural modal weight is automatically increased to 0.4, the timing modal weight is automatically increased to 0.4, and the text modal weight is reduced to 0.2.
[0036] By setting the weights of different modal features based on business scenarios, we can focus on key modalities and reduce interference from redundant features, so that the final fusion features can accurately capture key business features, thereby improving the accuracy of subsequent fault label determination.
[0037] 103. Input the final fusion features into the pre-trained fault classification model to obtain the target fault label corresponding to the abnormal device.
[0038] In this step, the fault modes corresponding to abnormalities in all devices in the system are pre-determined. Historical multimodal fault data corresponding to each fault mode is collected. Feature extraction and fusion are performed as described in steps 101-102 above. The corresponding fault mode descriptions are labeled as pre-set fault category labels, such as "transformer overheating" and "cooling system failure." These data are then divided into training, validation, and test sets in a 7:2:1 ratio to serve as training samples for the model. A fault classification model is trained based on these training samples. This model typically uses a multilayer perceptron (MLP) or deep neural network (DNN) as the base classifier. To capture complex nonlinear relationships (such as interactions between multimodal features), a Transformer or graph neural network (GNN) can be introduced to enhance the model's expressiveness. The cross-entropy loss function is used, and the optimizer uses Adam or SGD with Momentum. Hyperparameter tuning is performed using grid search or Bayesian optimization to adjust the learning rate, batch size, and regularization coefficient. The fault classification model specifically includes two hidden layers (e.g., 256 neurons in each layer, ReLU activation) and a Softmax output layer (e.g., 163 types of fault labels). The trained fault classification model is embedded in the equipment monitoring system, which receives the final fusion features in real time and outputs the fault labels.
[0039] After obtaining the final fused features, they are input into a pre-trained fault classification model. This fault classification model maps the final fused features to a latent space through two hidden layers, calculates the logarithmic probability between each preset fault label, and normalizes the logarithmic probability into a probability distribution through a Softmax output layer. This probability distribution includes the matching probability between the final fused features and each preset fault label. By setting a probability threshold, the matching probability is compared with the probability threshold. If the matching probability exceeds the probability threshold, the preset fault label corresponding to the matching probability can be used as the target fault label. If the matching probability does not exceed the probability threshold, an "unknown fault label" can be predefined and output as the target probability label, simultaneously triggering operations such as manual review.
[0040] It should be noted that if there are multiple matching probabilities exceeding the probability threshold, then all of the corresponding preset fault labels can be used as the target fault label, or the preset fault label with the highest matching probability can be used as the target fault label. This embodiment does not limit this.
[0041] Based on the above Figure 1It can be seen from the implementation method that by extracting multimodal features such as text modal features, structural modal features and time series modal features of abnormal equipment, it is possible to cover the multi-dimensional information of equipment failures, and to fuse them through the cross-dimensional correlation between text modal features, structural modal features and time series modal features to obtain the final fusion features, so that abnormal equipment can break through the limitations of traditional single-modal data analysis. Based on the cross-dimensional interaction of text-structure-time series features, a three-dimensional fault portrait of the abnormal equipment is constructed. The target fault label corresponding to the final fusion feature is determined by the pre-trained fault classification model, which can ensure the high precision and robustness of fault classification, thereby improving the accuracy and real-time performance of equipment fault label determination.
[0042] Furthermore, the preferred embodiment of the present application is in the above Figure 1 Based on this, the process of determining equipment fault labels based on multimodal feature fusion is described in detail. The specific steps are as follows: Figure 2 Shown, including:
[0043] 201. Extract multimodal features corresponding to abnormal devices.
[0044] This step is combined with the description in step 101 of the above method, and the same content will not be repeated here. It should be noted that the specific execution process of extracting the multimodal features corresponding to the abnormal device is: obtaining the multimodal data corresponding to the abnormal device, the multimodal data includes device log data, device topology communication data and device status data; using the BERT-Transformer hybrid network to segment and encode the device log data to obtain text modal features; using the dynamic GraphSAGE network to construct the device communication graph corresponding to the device topology communication data, and calculate the degree centrality of the node and the mean of the edge transmission delay to obtain the structural modal features; performing FFT transformation on the device status data to extract the Mel spectrum features, and combining the bidirectional LSTM-CRF model to obtain the time series modal features.
[0045] In this step, device log data (such as error codes, operation records, and alarm information) can be captured in real time through built-in sensors, remote monitoring systems, or operation and maintenance log platforms. This data is then de-noised and standardized, removing irrelevant characters (such as special symbols and duplicate logs), correcting spelling errors, and standardizing timestamp formats (such as ISO8601). This converts unstructured logs into semi-structured data (such as JSON format).
[0046] For device topology communication data, network traffic monitoring tools (such as Wireshark) or device communication protocol analysis (such as Modbus and MQTT) can be used to obtain the communication links, transmission delays, and data packet sizes between devices. Dynamic modeling and synchronous alignment are performed on the data, that is, a dynamic graph structure is constructed based on real-time communication data (nodes represent devices, and edges represent communication links). Communication data and device status data are aligned by timestamp to ensure multimodal data consistency. For device status data, device operating status parameters (such as voltage, current, temperature, and vibration signals) can be collected through IoT sensors (such as temperature sensors and vibration sensors) or SCADA systems. The data is also aligned for denoising and normalization, that is, a sliding window mean filter or wavelet transform is used to remove sensor noise, and the data is scaled to the [0,1] interval (such as Z-score normalization).
[0047] To extract text modal features, a BERT-Transformer hybrid network can be used, consisting of an input layer, a BERT encoding layer, a Transformer encoding layer, and an output layer. The input layer is used to input the original log text (the token sequence after word segmentation). The BERT encoding layer is used to extract context-related semantic features (dimensionality: 768) using a pre-trained BERT model (such as bert-base-uncased). The Transformer encoding layer is used to enhance the long-distance dependency modeling capabilities based on the BERT output through a multi-head attention mechanism (Multi-Head Attention). The output layer is used to pool the final hidden state of the Transformer (such as average pooling) to generate a 128-dimensional text modal feature vector, thus obtaining the text modal features.
[0048] For example, use the BERT tokenizer to tokenize the log text and add special tags. Load a pretrained BERT model, freeze the underlying parameters, and fine-tune only the top layer. Output the hidden state of each token. Build a 1-2 layer Transformer encoder and input the token embeddings output by BERT. Use the self-attention mechanism to capture key failure modes in the log (such as the association between "insufficient coolant" and "abnormal oil temperature"). Perform average pooling on the token embeddings output by the Transformer to generate a 128-dimensional text feature vector.
[0049] To extract structural modal features, a device communication graph is constructed based on the device topology communication data. Nodes represent device entities (e.g., transformers, cooling pumps), edges represent communication links (e.g., Modbus connections), and weights represent the mean transmission delay (in milliseconds). A graph neural network is constructed using the GraphSAGE model. The adjacency matrix and node feature matrix of the device communication graph are input into the GraphSAGE model. Mean Aggregator or LSTM aggregation is used to calculate the latent representation of the nodes and output an embedding vector for each node. Degree centrality is calculated to measure the importance of nodes in the communication network (e.g., core devices have higher degree centrality). The mean edge transmission delay is calculated to reflect the stability of the communication link (the higher the mean delay, the more likely a communication failure will cause a device anomaly). By concatenating the node degree centrality and the mean edge transmission delay, a 64-dimensional structural modal feature vector is generated, thus yielding the structural modal features.
[0050] To extract time series modal features, the device status time series data (such as vibration signals) is pre-windowed (such as a Hanning window) and framed (frame length: 256 points). The spectrum of each frame is extracted and calculated using FFT transform, and the spectrum is mapped to the Mel scale (usually 40 filters are used) using a triangular Mel filter bank. That is, a Mel filter bank is applied to the spectrum of each frame, and the logarithmic energy is taken to obtain the Mel spectrum. The Mel spectrum is input into the bidirectional LSTM in the bidirectional LSTM-CRF model to output a bidirectional hidden state. Based on the hidden state output by the LSTM, the transition probability of the time series label is modeled through the conditional random field (CRF) to capture the temporal dependency of the fault mode. The final hidden state of the CRF layer is globally averaged and pooled to generate a time series modal feature vector, thus obtaining the time series modal feature.
[0051] 202. Dynamic attribute reduction processing is performed on text modal features, structural modal features, and temporal modal features respectively.
[0052] Among them, dynamic attribute reduction processing includes dynamic adjustment of feature importance weights and hierarchical feature dimensionality reduction. The dynamic adjustment of feature importance weights is used to characterize the dynamic drift between frequency weights, mutual information weights and chi-square test weights.
[0053] In this step, after extracting textual, structural, and temporal modal features, each contains numerous features. However, determining which features are critical and which are non-critical depends on the device type of the abnormal device. Frequency weights reflect feature occurrence density, mutual information weights measure the correlation between features and labels, and chi-squared test weights assess feature independence. This device type includes both newly connected and existing devices. For existing devices, which have abundant historical data and stable failure modes, a balanced consideration of frequency weights, mutual information weights, and chi-squared test weights is necessary to avoid bias from a single metric (e.g., frequency weights may overlook low-frequency but critical abnormal features). For newly connected devices, which lack historical data, a reliance on mutual information weights (which directly reflect the correlation between features and fault labels) is necessary to quickly identify potential failure modes and avoid misclassifications due to data sparsity. Therefore, different feature importance weight adjustment strategies can be set to accommodate differences in device types, ensuring that various modal features and data distributions vary. This ensures robustness in various scenarios when the features are subsequently fused and fed into the fault classification model.
[0054] After the feature importance is determined, in order to reduce computing resource consumption and avoid the dimensionality disaster that may be caused by high-dimensional features, the text modal features, structural modal features, and temporal modal features can be reduced in dimension based on the above-mentioned feature importance. Specifically, multiple dimension intervals can be preset, and each dimension interval corresponds to a dimensionality reduction strategy. By counting the dimensions of the text modal features, structural modal features, and temporal modal features, and matching them with the preset multiple dimension intervals, the text modal features, structural modal features, and temporal modal features are processed according to the dimensionality reduction strategy corresponding to the hit dimension interval, thereby realizing dynamic attribute reduction of multimodal features.
[0055] Based on the above description, the specific execution process of dynamic attribute reduction processing of text modal features, structural modal features and time series modal features is as follows: if the abnormal device is an already connected device, the target weights are set for the text modal features, structural modal features and time series modal features according to the first feature importance weight adjustment strategy, and the first feature importance weight adjustment strategy is used to characterize the strategy of balanced distribution among frequency weights, chi-square test weights and mutual information weights; if the abnormal device is a newly connected device, the target weights are set for the text modal features, structural modal features and time series modal features according to the second feature importance weight adjustment strategy, and the second feature importance weight adjustment strategy is used to characterize the strategy of mutual information weight as the dominant one among frequency weights, chi-square test weights and mutual information weights; based on the target weights, the feature importance of the text modal features, structural modal features and time series modal features are ranked respectively, and the ranked text modal features, structural modal features and time series modal features are subjected to dimensionality reduction processing according to the hierarchical dimensionality reduction strategy.
[0056] In this step, whether the abnormal device is a newly connected device or an already connected device can be determined through the device ID / historical data records and dynamic detection mechanism. If it is an already connected device, the first feature importance weight adjustment strategy is adopted (balanced distribution of frequency weight, mutual information weight, and chi-square test weight). By evenly distributing the three weights (frequency, mutual information, and chi-square test), it adapts to the stable failure mode of known devices. Under this strategy, for the calculation of frequency weight, the frequency of occurrence of the feature in historical fault data (such as log keyword frequency and communication link activity) can be counted. For the calculation of mutual information, the mutual information (MI) between the feature and the fault label can be calculated to measure the contribution of the feature to fault prediction. For the calculation of the chi-square test, the independence of the feature and the fault label can be evaluated through the chi-square test. The target weight can be obtained by distributing the three weights in a preset relatively balanced ratio. For example, the frequency weight is 0.4, the mutual information weight is 0.4, and the chi-square test weight is 0.2. For newly connected devices, the second feature importance weight adjustment strategy (mutual information weight dominance) is adopted. By using mutual information weight dominance, we can adapt to the exploration needs of unknown failure modes of new devices. With this strategy, we can use the above balanced distribution as a basis, assign a higher proportion to the mutual information weight, and supplement it with frequency weight and chi-square test weight to obtain the target weight. For example, the frequency weight is 0.2, the mutual information weight is 0.6, and the chi-square test weight is 0.2.
[0057] According to the above two feature importance weight adjustment strategies, the target weight corresponding to the feature can be determined, and the text modal features, structural modal features and time series modal features can be sorted according to the target weight. Pre-set multiple dimension intervals, for example, above 500 dimensions, 500-200 dimensions, and below 200 dimensions. Set a dimensionality reduction strategy for each dimension interval. For example, when the feature dimension is higher than 500, start the t-SNE algorithm to map the data to a 64-dimensional visualization space, retaining the clustering characteristics of the device status. When the feature dimension is in the medium complexity scenario of 200-500 dimensions, PCA principal component analysis is used to extract 128 key features; when the feature dimension is lower than 200 dimensions, L1 regularization screening is directly enabled for simple scenarios, and efficient calculation is achieved by setting feature importance thresholds (such as Top50 features). That is, retain high-weight features and reduce redundant dimensions.
[0058] It should be noted that, for the subject of the dimensionality reduction process, the features of the text modal features, structural modal features, and temporal modal features after being sorted together can be processed uniformly, or the features of the text modal features, structural modal features, and temporal modal features after being sorted individually can be processed separately. This embodiment does not limit this.
[0059] 203. Determine the multi-head focus targets corresponding to the cross-modal attention mechanism.
[0060] Among them, a focus target corresponds to a cross-dimensional association task.
[0061] In this step, focus objectives can be set around multimodal complementarity and task relevance. Multimodal complementarity means that each focus objective must cover interactive tasks across different modalities (e.g., text-structure, structure-time sequence, text-time sequence). Task relevance, on the other hand, means that the focus objective must be strongly related to the fault classification task (e.g., "the relationship between abnormal oil temperature and communication topology").
[0062] For example, Head 1 shows the association between the text mode (log keywords) and the structural mode (communication graph node degree centrality). Head 2 shows the association between the structural mode (transmission delay mean) and the temporal mode (Mel spectrum high-frequency energy). Head 3 shows the association between the text mode (device status description) and the temporal mode (temperature trend slope).
[0063] 204. According to the multi-head focusing targets, cross-dimensional correlation analysis is performed on the text modal features, structural modal features and temporal modal features respectively to obtain the interactive correlation features corresponding to each focusing target.
[0064] In this step, based on the multi-head focus target set in step 203, cross-dimensional correlation analysis can be performed on text modal features, structural modal features, and time series modal features. For text modal feature processing, BERT or CLIP can be used to extract text embeddings (such as the semantic vector of "insufficient coolant"). For example, interacting with structural modalities: the text embeddings are combined with the degree centrality features of the communication graph to perform attention scoring to capture the association between "insufficient coolant" and communication stability. For structural modal feature processing, a graph neural network (GNN) can be used to extract topological features of the communication graph (such as node degree centrality and mean transmission delay). For example, interacting with time series modalities: the mean transmission delay is combined with the time series Mel-spectrogram features to perform attention fusion to identify the collaborative pattern between communication delay and oil temperature anomalies. For time series modal feature processing, LSTM or MTF (Markov transition field) can be used to convert one-dimensional time series signals into two-dimensional images (such as a time-frequency graph of oil temperature trends). For example, in interaction with text modalities, attention is matched between temporal trend slopes and keywords in text descriptions (e.g., "temperature suddenly rises") to enhance the semantic interpretation of temporal patterns. Each head outputs a fused feature vector (e.g., head 1 outputs text-structure interaction features, head 2 outputs structure-temporal interaction features), which is also known as the interactive correlation feature.
[0065] 205. The interactive correlation features corresponding to each aggregation target are spliced by dimension, and the modal weight arbitration is performed on the spliced interactive correlation features according to the business scenario to obtain the final fusion feature.
[0066] In this step, the multi-head interactive correlation features are concatenated dimensionally to form a high-dimensional fusion feature. For example, the first feature (text-structure interaction) is 128-dimensional, the first feature (structure-time series interaction) is 256-dimensional, and the first feature (text-time series interaction) is 128-dimensional. The concatenated feature is 512-dimensional. Business scenarios include fault diagnosis mode and steady-state operation mode. Fault diagnosis mode monitors, analyzes, and locates problems during system operation, and then quickly implements remedial measures to restore normal system operation. Its core goal is to respond quickly, accurately locate faults, and reduce downtime. Steady-state operation mode, on the other hand, maintains stable and efficient system operation under normal operating conditions. Its core goal is to maintain long-term system stability and optimize resource utilization, avoiding unnecessary fluctuations or resource waste. Modal weights are arbitrated by monitoring the current business scenario, that is, modal weights are adjusted (for example, in the case of equipment failure, the structural modal weight is 50%, the text modal weight is 30%, and the time series modal weight is 20%). Use the channel attention mechanism (such as SE Block) to calculate the weight coefficient of each modality, and add the weighted sum of the spliced features according to the weight coefficient of each modality calculated above to generate the final fusion feature.
[0067] 206. Calculate the log odds between the final fusion feature and all preset fault labels in the fault classification model.
[0068] The preset fault label is determined based on the equipment failure mode included in the training sample of the fault classification model, and one preset fault label corresponds to one failure mode.
[0069] This step is combined with the description of the model in step 103 of the above method, and the same content is not repeated here. The preset fault labels are defined based on the equipment failure mode in the training sample (such as "insufficient coolant", "abnormal oil temperature", "communication interruption"). Specifically, one-hot encoding can be used to represent each fault label (such as [1,0,0] represents "insufficient coolant"). The specific formula is:
[0070] logits = W clf ·F fusion +b;
[0071] Among them, W clf is the classifier weight, F fusion is the final fusion feature, b is the bias vector (biasterm) of the classifier, specifically a constant vector.
[0072] Because the fault classification model contains two hidden layers, the two hidden layers can be used to map the fused features into the fault label space and output the logarithmic probability. This logarithmic probability represents the logarithm of the ratio of the probability of the fault mode corresponding to a preset fault label to the probability of the fault mode not occurring.
[0073] 207. Determine the probability distribution between the final fusion feature and all preset fault labels based on the logarithmic probability and the total number of labels corresponding to the preset fault labels.
[0074] The probability distribution includes the matching probability between the final fusion feature and each preset fault label.
[0075] In this step, since the fault classification model also includes a Softmax output layer, the logarithmic probability is converted into a probability distribution through the Softmax output layer. For example, the matching probability of each preset fault label (such as [0.7, 0.2, 0.1] indicates that the probability of "insufficient coolant" is the highest) is calculated using the following formula:
[0076]
[0077] Among them, p(y i |F fusion ) is the matching probability of each preset fault label.
[0078] 208. If there is a matching probability higher than a preset probability threshold, the corresponding preset fault label is used as a target probability label.
[0079] In this step, the preset probability threshold is usually set based on business needs and model performance. For example, if the model predicts that the probability of a certain fault category must reach 0.9 or above, it can be trusted, and if it is lower than this threshold, it is considered "uncertain". If there is a matching probability higher than the preset probability threshold (such as 0.9), the corresponding fault label is output as the target probability threshold corresponding to the abnormal device. For example, after inputting the fusion feature, the probability distribution is [0.95, 0.1, 0.05], and the "insufficient coolant" label is output.
[0080] Based on steps 203-208 above, complementary features are extracted by designing multimodal interaction tasks (such as text-structure and structure-time series). High-precision fused features are generated by combining modal weight arbitration and feature concatenation. Robust fault label output is achieved through softmax and threshold determination. This approach offers significant advantages in scenarios such as equipment fault diagnosis and industrial monitoring, effectively improving the accuracy and interpretability of fault label determination.
[0081] 209. If there is no matching probability higher than the preset probability threshold, the predefined unknown fault label is used as the target fault label, and a manual review operation is triggered.
[0082] In this step, the "unknown fault label" is pre-defined. For example, "Unknown_Fault" or "Manual_Review_Needed" is used to mark faults that cannot be automatically classified. If the probability of no match is lower than the preset probability threshold (such as 0.9), the "unknown fault label" is output as the target fault label, which is actually outputting the "unknown fault label". Since the abnormal equipment has a fault that cannot be automatically classified, relevant personnel are required to perform manual review operations at this time. That is, the relevant fault data and model output results are pushed to the manual review queue, and the reviewer is notified. The reviewer determines the final fault label based on contextual information (such as historical data and fault characteristics), realizing efficient collaboration between the automated system and manual review, and avoiding the model outputting wrong decisions when the confidence level is low.
[0083] Furthermore, as a response to the above Figure 1-2 The implementation of the method embodiment shown in the figure, the embodiment of the present application provides a device for determining equipment fault labels based on multimodal feature fusion, which is used to improve the accuracy and real-time performance of equipment fault label determination through multimodal feature fusion. The embodiment of the device corresponds to the aforementioned method embodiment. For ease of reading, this embodiment will no longer repeat the details of the aforementioned method embodiment one by one, but it should be clear that the device in this embodiment can correspond to all the contents of the aforementioned method embodiment. Specifically, Figure 3 As shown, the device includes:
[0084] An extraction unit 31 is configured to extract multimodal features corresponding to abnormal devices, wherein the multimodal features include text modal features, structural modal features, and temporal modal features;
[0085] A fusion unit 32 is configured to fuse the text modal features, the structural modal features, and the temporal modal features obtained by the extraction unit 31 using a cross-modal attention mechanism to obtain a final fused feature, wherein the cross-modal attention mechanism is configured to characterize the cross-dimensional correlation between the text modal features, the structural modal features, and the temporal modal features;
[0086] The determination unit 33 is configured to input the final fusion feature obtained by the fusion unit 32 into a pre-trained fault classification model to obtain a target fault label corresponding to the abnormal device.
[0087] Further, such as Figure 4 As shown, the extraction unit 31 includes:
[0088] An acquisition module 311 is configured to acquire multimodal data corresponding to the abnormal device, wherein the multimodal data includes device log data, device topology communication data, and device status data;
[0089] A first extraction module 312 is configured to segment and encode the device log data obtained by the acquisition module 311 using a BERT-Transformer hybrid network to obtain the text modality feature;
[0090] A second extraction module 313 is configured to construct a device communication graph corresponding to the device topology structure communication data obtained by the acquisition module 311 using a dynamic GraphSAGE network, and calculate the degree centrality of the nodes and the mean value of the edge transmission delay to obtain the structural modal features;
[0091] The third extraction module 314 is used to perform FFT transformation on the device status data obtained by the acquisition module 311 to extract Mel spectrum features, and combine the bidirectional LSTM-CRF model to obtain the time series modal features.
[0092] Further, such as Figure 4 As shown, the device also includes:
[0093] The processing unit 34 is used to perform dynamic attribute reduction processing on the text modal features, the structural modal features and the temporal modal features respectively. The dynamic attribute reduction processing includes dynamic adjustment of feature importance weights and hierarchical feature dimensionality reduction. The dynamic adjustment of feature importance weights is used to characterize the dynamic drift between frequency weights, mutual information weights and chi-square test weights.
[0094] Further, such as Figure 4 As shown, the processing unit 34 includes:
[0095] A first processing module 341 is configured to, if the abnormal device is a connected device, set target weights for the text modal feature, the structural modal feature, and the temporal modal feature according to a first feature importance weight adjustment strategy, wherein the first feature importance weight adjustment strategy is used to represent a strategy for balancing the frequency weight, the chi-square test weight, and the mutual information weight;
[0096] A second processing module 342 is configured to, if the abnormal device is a newly connected device, set target weights for the text modal feature, the structural modal feature, and the temporal modal feature according to a second feature importance weight adjustment strategy, wherein the second feature importance weight adjustment strategy is used to represent a strategy in which the mutual information weight is dominant among the frequency weight, the chi-square test weight, and the mutual information weight;
[0097] The third processing module 343 is used to sort the feature importance of the text modal features, the structural modal features and the temporal modal features based on the target weights obtained by the first processing module 341 or the second processing module 342, and perform dimensionality reduction processing on the sorted text modal features, the structural modal features and the temporal modal features according to a hierarchical dimensionality reduction strategy.
[0098] Further, such as Figure 4 As shown, the fusion unit 32 includes:
[0099] A first determination module 321 is configured to determine a multi-head focus target corresponding to the cross-modal attention mechanism, where one focus target corresponds to one cross-dimensional association task;
[0100] An analysis module 322 is configured to perform cross-dimensional correlation analysis on the text modal features, the structural modal features, and the temporal modal features according to the multiple focus targets obtained by the first determination module 321, to obtain interactive correlation features corresponding to each focus target;
[0101] The fusion module 323 is used to splice the interactive correlation features corresponding to each aggregation target obtained by the analysis module 322 by dimension, and perform modal weight arbitration on the spliced interactive correlation features according to the business scenario to obtain the final fusion feature.
[0102] Further, such as Figure 4 As shown, the determining unit 33 includes:
[0103] a calculation module 331 configured to calculate the logarithmic probability between the final fusion feature and all preset fault labels in the fault classification model, wherein the preset fault labels are determined based on the device failure modes included in the training samples of the fault classification model, and each preset fault label corresponds to one failure mode;
[0104] A second determination module 332 is configured to determine a probability distribution between the final fused feature and all the preset fault labels based on the logarithmic probability obtained by the calculation module 331 and the total number of labels corresponding to the preset fault labels, wherein the probability distribution includes a matching probability between the final fused feature and each of the preset fault labels;
[0105] The first confirmation module 333 is configured to use the corresponding preset fault label as the target probability label if the matching probability obtained by the second determination module 332 is higher than a preset probability threshold.
[0106] Further, such as Figure 4 As shown, the device also includes:
[0107] The second confirmation module 334 is configured to use a predefined unknown fault label as the target fault label and trigger a manual review operation if the matching probability obtained by the second determination module 332 does not exist and is higher than the preset probability threshold.
[0108] Furthermore, the embodiment of the present application also provides a storage medium, which is used to store a computer program, wherein when the computer program is running, the device where the storage medium is located is controlled to execute the above Figure 1-2 The device fault label determination method based on multimodal feature fusion described in.
[0109] Furthermore, the embodiment of the present application also provides a processor, which is used to run a program, wherein the program executes the above Figure 1-2 The device fault label determination method based on multimodal feature fusion described in.
[0110] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0111] It is understood that the relevant features of the above methods and devices can be referenced to each other. In addition, the terms "first" and "second" in the above embodiments are used to distinguish between the embodiments, and do not represent the advantages and disadvantages of the embodiments.
[0112] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0113] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems may also be used together with the teachings herein. Based on the above description, it is apparent that the structure required for constructing such systems is suitable. In addition, the present application is not directed to any specific programming language. It should be understood that various programming languages may be utilized to implement the present application described herein, and the description of the specific languages above is provided for the purpose of disclosing the preferred embodiment of the present application.
[0114] In addition, the memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0115] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0116] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0117] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0118] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0119] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0120] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0121] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0122] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0123] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0124] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A method for determining equipment fault labels based on multimodal feature fusion, characterized in that: The method comprises: Extracting multimodal features corresponding to abnormal devices, wherein the multimodal features include text modal features, structural modal features, and temporal modal features; The text modal features, the structural modal features, and the temporal modal features are fused using a cross-modal attention mechanism to obtain a final fused feature, wherein the cross-modal attention mechanism is used to characterize the cross-dimensional correlation between the text modal features, the structural modal features, and the temporal modal features; The final fusion feature is input into a pre-trained fault classification model to obtain the target fault label corresponding to the abnormal device.
2. The method according to claim 1, characterized in that Extract multimodal features corresponding to abnormal devices, including: Acquire multimodal data corresponding to the abnormal device, the multimodal data including device log data, device topology communication data, and device status data; Using a BERT-Transformer hybrid network to segment and encode the device log data to obtain the text modality features; Based on the structural data, a device communication graph corresponding to the device topology structure communication data is constructed using a dynamic GraphSAGE network, and the degree centrality of the nodes and the mean value of the edge transmission delay are calculated to obtain the structural modal characteristics; An FFT transform is performed on the device status data to extract the Mel spectrum features, and the time series modal features are obtained by combining the bidirectional LSTM-CRF model.
3. The method according to claim 1, characterized in that Before fusing the text modality feature, the structural modality feature, and the temporal modality feature using a cross-modal attention mechanism to obtain a final fused feature, the method further includes: Dynamic attribute reduction processing is performed on the text modal features, the structural modal features and the temporal modal features respectively. The dynamic attribute reduction processing includes dynamic adjustment of feature importance weights and hierarchical feature dimensionality reduction. The dynamic adjustment of feature importance weights is used to characterize the dynamic drift between frequency weights, mutual information weights and chi-square test weights.
4. The method according to claim 3, characterized in that Performing dynamic attribute reduction processing on the text modal feature, the structural modal feature, and the temporal modal feature includes: If the abnormal device is a connected device, target weights are set for the text modal feature, the structural modal feature, and the temporal modal feature according to a first feature importance weight adjustment strategy, where the first feature importance weight adjustment strategy is used to represent a strategy for balanced distribution among the frequency weight, the chi-square test weight, and the mutual information weight; If the abnormal device is a newly connected device, target weights are set for the text modal feature, the structural modal feature, and the temporal modal feature according to a second feature importance weight adjustment strategy, wherein the second feature importance weight adjustment strategy is used to characterize a strategy in which the mutual information weight is dominant among the frequency weight, the chi-square test weight, and the mutual information weight; The text modal features, the structural modal features and the temporal modal features are sorted by feature importance based on the target weights, and the sorted text modal features, the structural modal features and the temporal modal features are subjected to dimensionality reduction processing according to a hierarchical dimensionality reduction strategy.
5. The method according to any one of claims 1 to 4, characterized in that The text modality features, the structural modality features, and the temporal modality features are fused using a cross-modal attention mechanism to obtain a final fused feature, including: Determine a multi-head focus target corresponding to the cross-modal attention mechanism, where one focus target corresponds to a cross-dimensional association task; Performing cross-dimensional correlation analysis on the text modal features, the structural modal features, and the temporal modal features according to the multiple focus targets, to obtain interactive correlation features corresponding to each focus target; The interactive correlation features corresponding to each aggregation target are spliced by dimension, and the modal weight arbitration is performed on the spliced interactive correlation features according to the business scenario to obtain the final fusion feature.
6. The method according to claim 1, characterized in that The final fusion feature is input into the pre-trained fault classification model to obtain the target fault label corresponding to the abnormal device, including: Calculating the logarithmic probability between the final fusion feature and all preset fault labels in the fault classification model, where the preset fault labels are determined based on device failure modes included in training samples of the fault classification model, and one preset fault label corresponds to one failure mode; Determining a probability distribution between a final fused feature and all the preset fault labels based on the logarithmic probability and the total number of labels corresponding to the preset fault labels, wherein the probability distribution includes a matching probability between the final fused feature and each of the preset fault labels; If there is a matching probability higher than a preset probability threshold, the corresponding preset fault label is used as the target probability label.
7. The method according to claim 6, characterized in that The method further comprises: If the matching probability does not exist and is higher than the preset probability threshold, the predefined unknown fault label is used as the target fault label, and a manual review operation is triggered.
8. A device for determining equipment fault labels based on multimodal feature fusion, characterized in that: The device comprises: An extraction unit, configured to extract multimodal features corresponding to abnormal devices, wherein the multimodal features include text modal features, structural modal features, and temporal modal features; a fusion unit, configured to fuse the text modal features, the structural modal features, and the temporal modal features obtained by the extraction unit using a cross-modal attention mechanism to obtain a final fused feature, wherein the cross-modal attention mechanism is configured to characterize a cross-dimensional correlation relationship among the text modal features, the structural modal features, and the temporal modal features; A determination unit is configured to input the final fusion feature obtained by the fusion unit into a pre-trained fault classification model to obtain a target fault label corresponding to the abnormal device.
9. A storage medium, characterized in that: The storage medium includes a stored program, wherein when the program is executed, the device where the storage medium is located is controlled to execute the device fault label determination method based on multimodal feature fusion according to any one of claims 1 to 7.
10. A processor, characterized in that: The processor is configured to run a program, wherein the program, when running, executes the method for determining a device fault label based on multimodal feature fusion according to any one of claims 1 to 7.
Citation Information
Cited By
Abnormality diagnosis method for label printing equipment based on multi-modal fusion
CN120921815A
Industrial equipment fault detection method, device and equipment based on large vertical domain model
CN120974433A
Industrial equipment fault detection method, device and equipment based on vertical domain large model
CN120974433B
Operation and maintenance analysis method and device of new energy equipment, and electronic equipment
CN121388878A
Multi-modal operation and maintenance data fault determination method and system based on large model
CN121456793A