Industrial equipment fault detection method, device and equipment based on large vertical domain model

By adopting a fault detection method for industrial equipment based on a vertical domain large model, multimodal data is acquired and preprocessed. Combined with structural perception enhancement and multimodal fusion modules, the problem of high false alarm rate and low accuracy of general models in complex equipment scenarios is solved, and high-precision fault identification and diagnosis are achieved.

CN120974433AActive Publication Date: 2025-11-18广东知业科技有限公司

Patent Information

Application Number
CN202511480219.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2025-11-18
Estimated Expiration
2045-10-16

AI Technical Summary

Technical Problem

Existing general-purpose large models are difficult to adapt to complex equipment structures, variable industrial operating conditions, and complex data modalities in vertical domain scenarios. In particular, when the number of equipment fault data samples is small, it leads to a high false alarm rate and poor fault identification accuracy.

Method used

An industrial equipment fault detection method based on a vertical domain large model is adopted. Multimodal data is acquired and preprocessed, and domain prior data is introduced as constraint parameters during the model training stage. Target pseudo-samples are generated by combining fault sample data with partial fault labels. Feature enhancement and semantic alignment are performed using a structure perception enhancement module and a multimodal fusion module to generate structural feature vectors and joint representation vectors.

Benefits of technology

It significantly improves fault identification accuracy, reduces the risk of false alarms, provides more reliable fault diagnosis support, adapts to the structural characteristics of complex equipment, and enhances the ability to correlate and complement information from multi-dimensional data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974433A_ABST
    Figure CN120974433A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial equipment fault detection method, device and equipment based on a large vertical domain model, and relates to the field of artificial intelligence, and the method comprises the steps: obtaining to-be-detected multi-modal data, and carrying out the preprocessing of the to-be-detected multi-modal data into processed data; processing the processed data into a fault category result through a fault detection model; the fault detection model comprises a structure perception enhancement module and a multi-modal fusion module which are connected with each other; the fault detection model is obtained based on training of fault sample data and target pseudo samples, part of the fault sample data is marked with fault label results, and the target pseudo samples are obtained according to the fault label results and the fault sample data by taking obtained field prior data as constraint parameters; the structure perception enhancement module is used for performing feature enhancement processing on the structure data to obtain a structure feature vector; and the multi-modal fusion module is used for performing semantic fusion processing on the structural feature vector, the time sequence data, the image data and the text data. According to the invention, the fault identification precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, in particular to an industrial equipment fault detection method, device and equipment based on a vertical domain large model. BACKGROUND

[0003] With the rapid development of industrial technology, industrial equipment, as the core carrier of the manufacturing production system, its running state directly determines the production efficiency, product quality and job safety. In the process of deepening industrialization and intelligent manufacturing, key equipment such as numerical control machine tools, heavy-duty motors, wind power equipment and rail transit traction systems are developing towards high integration, high automation and high complexity. Once a fault occurs, it may not only cause production line downtime and huge economic losses, but also cause equipment damage, personnel casualties and other safety accidents. Therefore, in order to ensure the continuous and stable operation of industrial production, how to accurately, efficiently and in advance diagnose the fault of industrial equipment is particularly important.

[0004] At present, in the industrial production scene, the related technology adopts a general large model to diagnose the fault of industrial data. However, this model is difficult to adapt to complex equipment structure, variable industrial working conditions, complex data modal and other vertical domain scenes. Especially in the case of a small number of equipment fault data samples, using a general large model to diagnose faults results in a high false positive rate and poor fault recognition accuracy. SUMMARY

[0006] The purpose of the present application is to provide an industrial equipment fault detection method, device and equipment based on a vertical domain large model.

[0007] To achieve the above purpose, the present application provides the following solutions: In a first aspect, the present application provides an industrial equipment fault detection method based on a vertical domain large model, comprising: obtaining to-be-detected multi-modal data in the running process of a target industrial equipment; preprocessing the to-be-detected multi-modal data to obtain processed data; the processed data includes structure data, time series data, image data and text data; The processed data is processed through a fault detection model to obtain a fault category result; the fault detection model comprises a structure perception enhancement module and a multi-modal fusion module connected with each other; the fault detection model is trained based on fault sample data and target pseudo samples, part of the fault sample data is labeled with a fault label result, the target pseudo samples are obtained by taking acquired domain prior data as a constraint parameter and according to the fault label result and the fault sample data; the structure perception enhancement module is used for performing feature enhancement processing on the structure data to obtain a structure feature vector; and the multi-modal fusion module is used for performing semantic alignment and fusion processing on the structure feature vector, the time series data, the image data and the text data.

[0008] In a second aspect, the present application provides an industrial equipment fault detection device based on a vertical domain large model, comprising: An acquisition module is configured to acquire multi-modal data to be detected in a running process of a target industrial equipment. A preprocessing module is configured to perform preprocessing on the multi-modal data to be detected to obtain processed data; the processed data comprises structure data, time series data, image data and text data. A fault detection module is configured to process the processed data through a fault detection model to obtain a fault category result; the fault detection model comprises a structure perception enhancement module and a multi-modal fusion module connected with each other; the fault detection model is trained based on fault sample data and target pseudo samples, part of the fault sample data is labeled with a fault label result, the target pseudo samples are obtained by taking acquired domain prior data as a constraint parameter and according to the fault label result and the fault sample data; the structure perception enhancement module is used for performing feature enhancement processing on the structure data to obtain a structure feature vector; and the multi-modal fusion module is used for performing semantic alignment and fusion processing on the structure feature vector, the time series data, the image data and the text data.

[0009] In a third aspect, the present application provides a computer device, comprising a memory, a processor, a computer program stored on the memory and executable on the processor, and the processor executes the computer program to implement the steps of the industrial equipment fault detection method based on the vertical domain large model according to any one of the above embodiments.

[0010] According to the embodiments of the present application, the following technical effects are disclosed: The application provides an industrial equipment fault detection method, device and equipment based on a vertical domain large model, which comprises the following steps: obtaining to-be-detected multi-modal data in the running process of a target industrial equipment; pre-processing the to-be-detected multi-modal data to obtain processed data; the processed data comprises structural data, time series data, image data and text data; processing the processed data in a fault detection model to obtain a fault category result; the fault detection model comprises a structure perception enhancement module and a multi-modal fusion module connected with each other; the fault detection model is obtained by training based on fault sample data and target pseudo samples; part of the fault sample data is labeled with a fault label result, and the target pseudo samples are obtained by taking acquired domain prior data as a constraint parameter and according to the fault label result and the fault sample data; the structure perception enhancement module is used for performing feature enhancement processing on the structural data to obtain a structural feature vector; and the multi-modal fusion module is used for performing semantic alignment and fusion processing on the structural feature vector, the time series data, the image data and the text data.

[0011] Compared with the prior art, in the scheme, the to-be-detected multi-modal data of the target industrial equipment including structure, time series, image, text and other dimensions are acquired and pre-processed, so that the characteristics of complex data modal in the industrial scene can be accurately covered, and comprehensive and actual running state of the equipment consistent data source can be provided for subsequent diagnosis; in the model training stage, the domain prior data is introduced as a constraint parameter, the target pseudo samples are generated in combination with the fault sample data labeled with part of the fault labels, the shortage of few fault data samples in the industrial scene is effectively made up, and the false alarm risk caused by insufficient samples is greatly reduced; and the structure perception enhancement module and the multi-modal fusion module connected with each other are arranged in the fault detection model, the structure perception enhancement module can perform feature enhancement on the structural data and generate a structural feature vector, so that the structural characteristic information of the image target industrial equipment is deeply mined, the structural characteristics of the complex equipment are effectively adapted, and the limitation that the traditional general model is difficult to match the complex equipment structure is broken; the multi-modal fusion module performs semantic alignment and fusion processing on the structural feature vector and the other three types of data, the deep correlation and information complement of multi-dimensional data are realized, the information deviation of single modal data is avoided, and then the fault recognition precision is significantly improved, and more reliable fault diagnosis support is provided for stable operation of the industrial equipment. BRIEF DESCRIPTION OF DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort.

[0013] Figure 1A structural schematic diagram of an application environment of an industrial equipment fault detection method based on a vertical domain large model according to an embodiment of the present application; Figure 2 A flowchart of an industrial equipment fault detection method based on a vertical domain large model according to an embodiment of the present application; Figure 3 A flowchart of a method for obtaining a fault category result by processing processed data in a fault detection model according to an embodiment of the present application; Figure 4 A flowchart of a method for obtaining a target pseudo sample according to an embodiment of the present application; Figure 5 A functional module schematic diagram of an industrial equipment fault detection device based on a vertical domain large model according to an embodiment of the present application; Figure 6 A structural schematic diagram of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION

[0015] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0016] The above purposes, features and advantages of the present application will be more apparent and easy to understand. The present application will be described in further detail below with reference to the drawings and specific embodiments.

[0017] In related art, equipment fault diagnosis generally relies on a large number of historical fault samples and high-quality labeled data. However, in actual industrial production scenarios, there are obvious small sample, class imbalance and modal heterogeneity problems in equipment fault data, which leads to weak generalization ability of traditional supervised learning or specific scene model and high deployment difficulty. Generally, a general large model is used for fault diagnosis of industrial data. However, this model is difficult to adapt to complex equipment structure, variable industrial working conditions, complex data modal and other vertical scene, especially in the case of small number of equipment fault data samples. Using a general large model to diagnose faults results in high false positive rate and poor fault recognition accuracy.

[0018] Based on the above defects, the application provides an industrial equipment fault detection method based on a vertical domain large model. Compared with the prior art, in the present scheme, the multi-dimensional data to be detected including structure, time sequence, image, text and the like of the target industrial equipment are acquired and preprocessed, which can accurately cover the characteristics of complex data modalities in industrial scenes, and provide comprehensive and actual running state of the equipment for subsequent diagnosis. In the model training stage, the domain prior data is introduced as a constraint parameter, and the target pseudo sample is generated by combining the fault sample data with part of the labeled fault labels, which effectively makes up for the shortage of few fault data samples in the industrial scene, and greatly reduces the false alarm risk caused by insufficient samples. In the fault detection model, a structure perception enhancement module and a multi-modal fusion module are connected with each other, the structure perception enhancement module can enhance the features of the structure data and generate a structure feature vector, thereby deeply mining the structure characteristic information of the image target industrial equipment, effectively adapting to the structure characteristics of complex equipment, and breaking the limitation of traditional general models that are difficult to match complex equipment structures. The multi-modal fusion module performs semantic alignment and fusion processing on the structure feature vector and other three types of data, realizes deep correlation and information complementation of multi-dimensional data, avoids information deviation of single modal data, and further significantly improves the fault recognition accuracy, and provides more reliable fault diagnosis support for stable operation of industrial equipment.

[0019] The industrial equipment fault detection method based on a vertical domain large model provided by the embodiments of the application can be applied to the application environment of the industrial equipment fault detection method based on a vertical domain large model as shown in Figure 1 The application environment includes a terminal 102, a server 104 and a data storage system. The terminal 102 communicates with the server 104 through a network. The data storage system can store the multi-modal data to be detected in the running process of the target industrial equipment acquired by the server 104. The data storage system can be separately arranged, or integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the acquired multi-modal data to be detected to the server 104. After the server 104 acquires the multi-modal data to be detected, it is preprocessed and processed by a fault detection model to obtain a fault category result. In addition, in some embodiments, the industrial equipment fault detection method can also be realized by the server 104 or the terminal 102 alone, such as direct preprocessing by the terminal 102 and processing by the fault detection model to obtain a fault category result. The fault detection model can be understood as a vertical domain large model, which is suitable for vertical domain scenes such as complex equipment structure, variable industrial conditions and complex data modalities.

[0020] The terminal 102 can be, but is not limited to, various desktop computers, notebook computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things device can be a smart speaker, a smart television, a smart air conditioner, a smart vehicle-mounted device, etc. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The server 104 can be implemented by a stand-alone server or a server cluster composed of multiple servers, and can also be a cloud server.

[0021] In an exemplary embodiment, as shown in Figure 2 , an industrial equipment fault detection method based on a vertical domain large model is provided, which is executed by a computer device, specifically by a terminal or a server, or by a terminal and a server together. In the embodiment of the present application, the method is applied to the server 104 in Figure 1 , including the following steps S201 to S203. Wherein: Step S201, obtaining multi-modal data to be detected in the running process of the target industrial equipment; It should be noted that the multi-modal data covering the equipment state collected from the target industrial equipment includes original structure data, original time series data, original image data, and original text data, which can be, for example, time series signals, equipment images, text descriptions, operation and maintenance logs, etc. It can also include audio data, video data, radar atlas and other industrial modal information. The present embodiment does not make any limitation on each industrial modal of the multi-modal data.

[0022] The structure data can include equipment BOM list and topology graph, which can record the hierarchical relationship of "equipment → component → sub-component" and the functional connection relationship between components, such as "numerical control machine tool → spindle system → bearing". The functional connection relationship between components can be, for example, the transmission connection between the bearing and the gear. The data format of the structure data can include graphical format and structured table form, which is used for subsequent modeling of the physical structure of the equipment. The time series data is collected by sensors in real time to collect dynamic signals in the running process of the target industrial equipment, including temperature, current, vibration, pressure, etc. The data format can be represented by "timestamp + value sequence", for example, 1 vibration value is collected every 10 ms, forming a time series sequence of 1 hour.

[0023] Image data refers to visual information about the appearance and internal components of target industrial equipment acquired through industrial cameras, infrared thermal imagers, etc. This can include surface defect images and infrared thermal images. Surface defect images include optical images of gear tooth wear, and infrared thermal images include the distribution of heating areas in motors. The data format can be represented by a pixel matrix, such as a 256×256 pixel grayscale / color image, used to capture visually visible fault features of the target industrial equipment, such as cracks, deformation, and areas of abnormal temperature. Text data refers to unstructured text collected during equipment operation and maintenance, which can include maintenance records, abnormal alarm logs, etc. Its data format can be string text, containing semantic information such as fault phenomena, handling methods, and fault codes, used to supplement empirical features that mechanistic data cannot cover. By collecting multi-dimensional data in this step, we can obtain more comprehensive fault characteristics of the target industrial equipment, providing good data guidance for subsequent multimodal fusion.

[0024] Step S202: Preprocess the multimodal data to be detected to obtain processed data; the processed data includes: structural data, time series data, image data and text data.

[0025] Understandably, the original multimodal data to be detected has problems such as heterogeneous format, asynchronous timing, and missing data. For example, vibration data timestamps are at 10ms intervals, while temperature data is at 100ms intervals; text data has typos, and time series data has null values. Therefore, it is necessary to preprocess the data to transform it into processed data. The processed data can be standardized data with a unified format, high quality, and can be directly input into the model. In one embodiment, the multimodal data to be detected is preprocessed to obtain processed data, including: performing timestamp alignment processing on the multimodal data to be detected to obtain aligned data; performing null value imputation processing on the aligned data to obtain imputed data; the imputed data includes temporal imputed data and other imputed data; performing segmentation and slicing processing on the temporal imputed data to obtain multiple continuous segments of fixed length; and performing feature standardization processing based on the continuous segments and other imputed data to obtain processed data.

[0026] Specifically, after acquiring the multimodal data to be detected, data synchronization and format unification can be performed first. This involves aligning the timestamps of the multimodal data to be detected, i.e., time-series synchronization, to obtain aligned data. For example, using high-frequency time-series data as a reference, low-frequency data can be padded with the same timestamps using interpolation, thus ensuring that multimodal data at the same moment correspond to the same device state. High-frequency time-series data could be, for example, vibration data at 10ms intervals, and low-frequency data could be, for example, temperature data at 100ms intervals. The interpolation method can be linear interpolation. There can be missing values in the aligned data. After the aligned data is determined, the aligned data is processed for null value filling, so as to obtain filled data. The filling methods for different modal data can be different. For missing values in time series data and image data, mean filling, interpolation filling repair can be used. The missing values can be, for example, 10 null values caused by temporary sensor failure. The mean filling method is suitable for stationary data, and the interpolation filling method is suitable for time series trend data. For ambiguous records in text data, the records can be supplemented in association with fault codes. For example, the ambiguous record in the text data is main shaft abnormality, without specifying the specific phenomenon, and the associated alarm code obtained is F03. The associated alarm code can be used to supplement it as "main shaft abnormality: bearing wear". The filled data can include time series filling data and other filling data. The time series filling data refers to the filled time series data, and the other filling data refers to the filled data of other modalities.

[0027] The time series filling data is segmented and sliced by using the interpolation method or the sliding window method, to obtain a plurality of fixed-length segmented slices. For example, 1 hour of time series data is divided into 100 "6-minute / segment" subsequences by using the sliding window technology. After the continuous segments are obtained, the continuous segments and the other filling data are subjected to feature standardization processing, to obtain processed data. The processed data can be that the filled structural data is converted into graph structure data, the filled image data is scaled to a unified resolution, for example, the pixel value of the 256x256 image data is normalized to the range of 0-1; and the filled text data is converted into a vector by using word segmentation, word embedding and the like, so as to ensure that the different modal data formats are adapted to subsequent encoders. In this step, by performing timestamp alignment, null value filling, segmentation and slicing, and feature standardization processing on the to-be-detected multi-modal data, the "format barrier" and "quality defect" of the multi-modal data can be eliminated, so that the processed data meets the input requirements of the model, and the data can be effectively read and utilized by the model, while accurately covering the characteristics of complex data modalities in industrial scenes, to provide a comprehensive and actual operation state-adapted data source for subsequent diagnosis.

[0028] In step S203, the processed data is processed by a fault detection model to obtain a fault category result. The fault detection model includes a structure perception enhancement module and a multi-modal fusion module connected to each other. The fault detection model is trained based on fault sample data and target pseudo sample data. Part of the fault sample data is labeled with a fault label result. The target pseudo sample data is obtained by using the acquired domain prior data as a constraint parameter and based on the fault label result and the fault sample data. The structure perception enhancement module is used for feature enhancement processing of the structure data to obtain a structure feature vector. The multi-modal fusion module is used for semantic alignment and fusion processing of the structure feature vector, time series data, image data and text data.

[0029] It should be noted that the pre-processed structure data needs to be further converted into a vector feature that can be understood by the model, and associated with other modal data, so as to enhance the physical correlation between fault features. The multi-modal fusion module is used to eliminate the semantic barriers of multi-modal data through sub-modal encoding and cross-attention alignment processing, so that multi-modal features can reflect the fault state cooperatively, and the limitations of single-modal features can be avoided.

[0030] In one of the embodiments, the above-mentioned fault detection model comprises a fault classification module connected with the multi-modal fusion module. The application also provides a specific implementation manner of processing the processed data through the fault detection model to obtain a fault category result, as shown in Figure 3 The method comprises the following steps: In step S301, the structure data is subjected to feature extraction and component relationship strengthening processing by the structure perception enhancement module to obtain a structure feature vector.

[0031] The above-processed data comprises structure data, text data, image data and time series data. After obtaining the processed data, the structure data is subjected to feature extraction processing by the structure perception enhancement module, and the hierarchical structure data is subjected to component relationship strengthening processing to obtain a structure feature vector.

[0032] Optionally, in the process of extracting and strengthening the component relationship of the structure data by the structure perception enhancement module to obtain the structure feature vector, the structure data is first subjected to topological modeling to construct a topological graph; the topological graph comprises nodes and edges, the nodes are used to represent component information in the equipment, and the edges are used to represent structure information between components; the structure information is encoded and processed by a graph neural network to obtain a structure embedding vector; for the hierarchical structure data in the structure data, the hierarchical information and the position information of the components in the topological graph are obtained; the hierarchical position encoding is generated according to the hierarchical information and the position information; the hierarchical position encoding is added to the component information and subjected to fusion processing to obtain a structure fusion vector; the structure fusion vector and the structure embedding vector are subjected to attention allocation processing by an attention layer to obtain a structure feature vector.

[0033] Specifically, the processed data described above includes structure data, text data, image data and time series data. After obtaining the processed data, the structure data needs to be feature encoded. The device topology structure in the structure data is first converted into a graph structure G=(V, E), where V is a node representing component information in the device, and E is a connection edge representing component information in the device. The component information can be the attribute features of each node, for example, for a bearing node, its component information can include model, rated speed, material, and a connection feature can also be added to each edge, such as transmission connection, fixed connection, etc. In this embodiment, by modeling the device topology as a graph structure, the physical dependency of the core components can be preserved, so that the subsequent model can more comprehensively capture the graph structure relationship of fault conduction, and avoid misjudgment caused by relying only on single component features; and the graph neural network (such as GCN / GAT) is used for node feature aggregation, which can capture the dependency relationship and conduction path characteristics between each component of the industrial device, strengthen the correlation between local and global features, and can be used to realize neighborhood node feature weighted aggregation through the graph convolution network GCN model, and assign higher attention weights to the key path through the graph attention network GAT, thereby effectively identifying component correlation features and improving fault recognition accuracy. Combined with the hierarchical position encoding and attention fusion mechanism, the feature differences of components at different functional levels can be distinguished through hierarchical position encoding, and the global and local feature weights can be dynamically adjusted through the attention fusion mechanism to focus on the fault sensitive area, so that the model considers the spatial hierarchy and functional coupling of the device while extracting features, thereby obtaining consistent structure representation from global to local. Compared with the model based on only single modal signal, the structure perception enhancement module provided in the present application can improve the fault recognition accuracy by about 8%~15%, and significantly reduce the misjudgment rate in complex devices (such as multi-stage transmission, thermoelectric coupling system), thereby providing accurate structure feature basis for multi-modal fusion.

[0034] It should be noted that the structure feature vector described above is a multi-dimensional vector formed after graph network coding, which is used to comprehensively reflect the topological connection relationship, functional level and signal interaction strength of the device components, and is the core structural representation for the model to perform fault identification and causal inference.

[0035] After the graph structure is constructed, the graph neural network can be used to encode and process the structure information in the graph structure. By calculating the association weight of the node and the neighbor node, the node attribute and global connection relationship are fused to generate a structure embedding vector S, S∈R d, d denotes the dimension of the structure embedding vector, and R represents the real number set. Among them, the vector of each node not only contains its own attribute, but also contains its position and function in the global structure of the device, for example, the vector of the main shaft bearing assembly reflects the structure information of “connected with motor transmission, linked with gear”. The structure embedding vector is encoded by the graph neural network, so that each node feature includes the global structure context, thereby enhancing the small sample fault recognition capability. In the subsequent feature extraction process of time sequence, image and text modalities, the relevance weight of the structure embedding vector and each modality feature is calculated through the attention fusion mechanism, the modality information most useful for current fault recognition is dynamically amplified, and selective interaction and enhancement of cross-modality features are realized. For example, when a bearing fault occurs, the relevance weight of the structure vector of the bearing node and the vibration feature increases. The graph neural network can include GAT or GCN, or can be implemented by device path embedding, structure Transformer or topological constraint based sparse representation method.

[0036] After constructing the structure embedding vector, for the hierarchical structure data in the hierarchical structure, the hierarchical position coding unique to each component is constructed by first obtaining the hierarchical information of the component in the subsystem and its position information in the topological graph. This hierarchical position coding can effectively represent the relative relationship and spatial layout of the component in the entire device structure system. Then, the coding is fused with the original component information to form a structure fusion vector that not only retains the inherent attributes of the component, but also integrates the structural context information in which the component is located, making the feature expression more holistic and relevant. The structure fusion vector and the structure embedding vector are processed by attention allocation through the attention layer, which can adaptively focus on the structural features that are more discriminative for fault diagnosis and weaken irrelevant information interference. The finally generated structure feature vector not only accurately captures the hierarchical topological relationship of the device, but also strengthens the weight of key structural information through the attention mechanism, providing more targeted and discriminative structure feature representation for subsequent fault detection, and significantly improving the perception and understanding ability of the model for complex device structures.

[0037] For example, for the hierarchical structure data of “device → component → sub-component”, the hierarchical position coding is added to the node, for example, the “main shaft system” belongs to the second component, and the hierarchical position coding reflects its “higher than bearing, lower than numerical control machine tool” hierarchical relationship. After fusion with the original node feature, the model is input into the model to perceive the hierarchical association between components, thereby obtaining the structure feature vector. The structure data is encoded by the graph neural network in this step, which can convert the abstract device structure into a numerical vector, and link with other modal features, so that the model can combine the device physical structure and hierarchical position relationship when extracting fault features, enhance the module's perception ability of the relative position between components, avoid the physical logic disconnection caused by relying only on data statistical features, dynamically amplify the modal information most useful for current fault identification, and realize selective interaction and enhancement of cross-modal features.

[0038] In step S302, the image data, time series data and text data are processed by the multi-modal fusion module for feature extraction, modal missing repair and semantic alignment to obtain a joint representation vector with consistent semantics.

[0039] It can be understood that the multi-modal fusion module is used to realize cross-modal feature alignment and fusion, and the preprocessed time series data, image data and text data need to extract features respectively and combine with the structure feature vector to fuse into a joint feature in a unified semantic space. The multi-modal fusion module includes an encoder, a missing repair unit and a semantic alignment unit connected in sequence. The encoders corresponding to each modal data are different.

[0040] In this embodiment, the image data, time series data and text data are processed by the multi-modal fusion module for feature extraction, modal missing repair and semantic alignment to obtain a joint representation vector with consistent semantics, including: the image data, time series data and text data are processed by the corresponding encoders respectively to extract image feature vectors, time series feature vectors and text feature vectors; the missing vector is determined from the structure feature vector, image feature vector, time series feature vector and text feature vector by the missing repair unit, and the missing vector is processed by the modal repair unit to obtain a target vector; the target vector includes a target image vector, a target time series vector, a target text vector and a target structure vector; the target image vector, the target time series vector, the target text vector and the target structure vector are aligned in a common semantic space by the semantic alignment unit to extract a joint representation vector with consistent semantics.

[0041] Specifically, the data of each modality is provided with a corresponding encoder, and feature extraction of each modality needs to be realized through the respective encoder. The encoder corresponding to the image data can include a CNN / ResNet encoder, the encoder corresponding to the time series data can include a Transformer / LSTM encoder, and the encoder corresponding to the text data can include a BERT / GRU encoder. The image data is input into the CNN / ResNet encoder to extract visual features and obtain an image feature vector, such as a high-temperature area feature vector extracted from an infrared thermal image. The time series data is input into the Transformer / LSTM encoder to capture time trend features and obtain a time series feature vector, such as a "2x frequency peak feature vector" in vibration data. The text data is input into the BERT / GRU encoder to extract semantic features and obtain a text feature vector, such as a keyword vector of "bearing wear" in a maintenance record.

[0042] Among them, the multi-modal fusion module extracts the feature vectors of each modality through the corresponding encoder, providing accurate basic data for subsequent feature fusion, that is, complementary enhancement is realized through the three types of encoders: ① the encoder corresponding to the time series data is used to capture the dynamic change trend of fault occurrence and the periodicity of fault correlation, improving the sensitivity of the model to abnormal fluctuations; ② the encoder corresponding to the image data can extract the device topology and the correlation between components, and strengthen the positioning accuracy of conduction type faults; ③ the encoder corresponding to the text / semantic data can incorporate expert experience and alarm semantics to make up for the knowledge gap that data is difficult to quantify. By extracting the feature vectors of each modality, the overall detection accuracy can be improved.

[0043] And the missing repair unit determines and repairs the missing vector through the feature vectors of each modality, effectively solving the pain point that multi-modal data is prone to missing in industrial scenarios, so that the missing vector can be obtained in time and repaired, so that the obtained target vector is more complete and comprehensive, which is convenient for subsequent fault detection based on more complete data, thereby greatly improving the fault detection accuracy. Subsequently, the target vectors of each modality are unified in different modal feature spaces through attention fusion and semantic alignment mechanisms, reducing modal bias and noise interference, and facilitating the acquisition of more effective and reliable modal feature information, further improving the fault detection accuracy. The overall scheme can improve the fault recognition accuracy by about 10%~18%, and significantly enhance the robustness of the multi-source heterogeneous data scene.

[0044] The target vector can include target feature vectors of each modality. In this embodiment, a modality loss tolerance mechanism is also provided, which is embodied by a loss repair unit. After obtaining the image feature vector, the time sequence feature vector and the text feature vector, there can be loss of modality data. Such loss often occurs in industrial scenarios due to sensor failure, data transmission interruption and other problems. The loss repair unit determines whether there is a loss in each feature vector. When there is no loss, the feature vector of this modality at this time is determined as the target feature vector. For example, when there is no loss in the time sequence feature vector, the time sequence feature vector is determined as the target structure vector. When there is a loss, a loss vector is determined and a repair strategy suitable for the characteristics of industrial data is used for modality repair processing to obtain the target feature vector. For example, when there is no image data due to failure of an infrared camera, the image data is repaired to obtain the target image vector.

[0045] After obtaining the target vector, the semantic alignment unit uses a cross-modal attention mechanism to map the target feature vectors of each modality to a common semantic space for alignment and extract a joint representation vector with consistent semantics. For example, the similarity between the vibration time sequence feature vector and the text semantic feature vector is calculated, and the vector representations of the two are adjusted so that the vibration peak feature and the bearing wear text description are closer in the common space, achieving semantic consistency, thereby outputting a joint representation vector with consistent semantics. The joint representation vector integrates the core features of the time sequence, image, text and structure four modalities, and fully reflects the multi-dimensional features of the fault. In this step, the image, time sequence, text, structure and other multi-modal information are used for fault judgment, so that the recognized information is more comprehensive. The exclusive modality encoder is used to extract features, and the cross-modal attention mechanism is used for semantic fusion, effectively solving the problem of inconsistent modal semantics. The modality loss tolerance mechanism is provided to ensure that the judgment can still be made in the case of loss of modalities or partial loss of signals, and the robustness of the system is improved.

[0046] Optionally, in addition to the cross-modal attention mechanism, the multi-modal fusion module can also use gate fusion, modality complementary encoder, general modality mapping space and other methods to achieve equivalent multi-modal data fusion.

[0047] In this embodiment, the integrity and semantic consistency of multi-modal data are effectively improved through the cooperative operation of the missing data repair unit and the semantic alignment unit. Among them, the missing data repair unit first conducts comprehensive detection on each feature vector, which can accurately identify the possible missing vectors. For the identified missing vectors, the unit uses a repair strategy suitable for the characteristics of industrial data to repair the modal, and finally generates a complete set of target vectors, ensuring the integrity of multi-modal data and avoiding information gaps caused by partial modal missing. On this basis, the semantic alignment unit maps the four types of target vectors to a unified common semantic space for alignment processing, extracts semantic consistent joint representation vectors by mining the potential association between different modalities, such as the correspondence between equipment vibration time series data and component wear image, the semantic matching between fault text description and structural abnormal features, etc. This process not only eliminates the semantic bias of multi-modal data due to differences in sources, but also realizes the deep association and complementarity of cross-modal information, so that the fused features not only retain the unique information of each modality, but also form a comprehensive representation with unified semantic dimension, providing more comprehensive, consistent and robust input features for subsequent fault detection models, significantly enhancing the model's ability to understand and utilize multi-source heterogeneous data in complex industrial scenarios.

[0048] In step S303, the joint representation vector is processed by the fault classification module to obtain the fault category result.

[0049] After obtaining the joint representation vector, the joint representation vector has fully integrated core information such as device structure, runtime sequence, appearance image and text description, and has the same semantic dimension after semantic alignment, which can comprehensively and accurately reflect the overall operation state and potential abnormal features of the device.

[0050] Optionally, the above fault classification module can include a fully connected layer, an attention layer, a regularization layer and a classification layer. The joint representation vector is transformed by the fully connected layer to mine deep fault association features, and the attention layer focuses on key fault features with high discrimination, and the regularization layer (such as Dropout or Batch Normalization) can inhibit overfitting and enhance the model's adaptability to data noise and distribution differences in industrial scenarios, to obtain enhanced features, and the classification decision layer maps the enhanced features to specific fault categories based on the industrial fault type system, outputs the final fault category result through multi-class discrimination logic, and realizes accurate conversion from fused features to diagnostic conclusions.

[0051] It should be noted that the above fault category result can include a fault category and a corresponding confidence, and the fault category can include mechanical wear, circuit short circuit, sealing failure, etc. The category with a confidence greater than a preset confidence threshold can be determined as the final target fault category. The fault category result can not only clearly locate the device fault type, but also provide a direct and reliable basis for subsequent troubleshooting and maintenance scheme formulation. In addition, the fault category result can ensure the accuracy and robustness of the classification result with the support of the pre-processing of the multi-modal data, effectively solve the problems of poor adaptability and high false positive rate of general models in industrial fault classification, and further improve the practical value of the overall fault diagnosis system.

[0052] Exemplarily, in a small sample scene with only 100 fault samples, the existing general large model and the fault detection model in the present application are respectively used for fault detection processing, the detection data of different models are counted, and the corresponding detection results are obtained. The general large model can include a general model and a general visual model. The general model can be, for example, BERT-base, and the general visual model can be a ResNet50 multi-modal fusion model. The detection results can include accuracy, precision, recall, F1 score, false positive rate and the like. The detection results can be seen from Table 1 as shown below: Table 1

[0053] As can be seen from Table 1, compared with the existing general large model, the accuracy, precision, recall and F1 score of the fault detection model in the present application are all higher than those of the general large model, and the false positive rate is significantly smaller than that of the general large model. Therefore, compared with the existing general large model, the fault detection effect of the fault detection model provided in the present application is better.

[0054] The application provides an industrial equipment fault detection method based on a vertical domain large model. Compared with the prior art, in the scheme, multi-dimensional to-be-detected multi-modal data including structure, time sequence, image, text and the like of a target industrial equipment are acquired and preprocessed, the characteristics of complex data modal in an industrial scene are accurately covered, comprehensive and actual operation state of the equipment are provided for subsequent diagnosis, domain prior data is introduced as a constraint parameter in a model training stage, target pseudo samples are generated in combination with part of fault sample data labeled with fault labels and structure feature vectors, the shortage of a small number of fault data samples in the industrial scene is effectively made up, and the false alarm risk caused by insufficient samples is greatly reduced. A structure perception enhancement module and a multi-modal fusion module are provided in the fault detection model and are connected with each other, the structure perception enhancement module can perform feature enhancement on structure data and generate a structure feature vector, thereby deeply mining the structure characteristic information of the image target industrial equipment, effectively adapting to the structure characteristics of complex equipment, and breaking the limitation that a traditional general model is difficult to match the structure of complex equipment. The multi-modal fusion module performs semantic alignment and fusion processing on the structure feature vector and the other three types of data, realizes deep correlation and information complementation of multi-dimensional data, avoids information deviation of single modal data, and further significantly improves fault recognition precision, thereby providing more reliable fault diagnosis support for stable operation of the industrial equipment.

[0055] In one of the embodiments, a specific implementation of constructing a fault detection model is also provided, and the method comprises the following steps: The target pseudo sample and the original equipment data are acquired, and the original equipment data is preprocessed into fault sample data; the fault sample data comprises first data labeled with fault label results, second data without labels and equipment operation data; equipment attribute data is acquired from a preset industry knowledge base, and the equipment attribute data is vectorized into knowledge vectors; the knowledge vectors, the fault sample data and the target pseudo sample are divided into a training set and a verification set according to a preset division rule; the training set is input into an initial model to obtain fault output results; a loss function is constructed according to the fault output results and the fault label results, and the parameters of each module in the initial model are iteratively optimized by using a small sample learning optimization strategy according to minimization of the loss function to obtain a to-be-verified model; the loss function comprises a combined classification loss, a feature consistency loss and a prototype constraint loss; the verification set is input into the to-be-verified model for verification to obtain a fault detection model.

[0056] Optionally, the original equipment data comprises data of each modal of the equipment, which can be acquired from an external device, acquired by real-time parameter acquisition of the equipment, or acquired from a blockchain or a database, and the acquisition mode of the fault sample data is not limited in the embodiment.

[0057] Specifically, in the process of obtaining the fault sample data, the original equipment data can be obtained first, and the original equipment data is preprocessed, including timestamp alignment, null value filling, feature standardization processing, and segmenting and slicing the time series data in the original equipment data to obtain the processed sample data. Then, the processed sample data is subjected to sample labeling and cleaning processing to obtain the fault sample data. The processed sample data includes equipment fault data and normal operation data. The equipment fault data is divided into first data and second data. The first data is labeled to form a fault label result, and the second data is not labeled. Then, the normal operation data other than the first data and the second data is retained, thereby forming the fault sample data.

[0058] The fault label result can include a fault category, such as bearing wear fault, and abnormal samples with contradictory labels and features are removed, thereby forming a first data set with a small amount of fault label results, a large amount of unlabeled second data, and a part of normal equipment operation data, providing a data basis for subsequent semi-supervised pre-training and small sample fine-tuning, so that the model still maintains high recognition accuracy and robustness in the case of few fault samples.

[0059] In one of the embodiments, as shown in Figure 4 The specific implementation of obtaining the target pseudo sample is also provided, and the method includes: Step S401, obtaining domain prior data and converting the domain prior data into constraint parameters; the constraint parameters include conditional constraint parameters, latent variable distribution constraint parameters, and multi-modal consistency rule parameters.

[0060] Step S402, inputting the fault label result, the fault sample data, and the structure sample vector into the autoencoder for processing with the constraint parameters as the generation condition to obtain pseudo sample data; the pseudo sample data is configured with a weight value and a sample parameter.

[0061] Step S403, calculating a confidence according to the weight value and the sample parameter, filtering the pseudo sample data with a confidence less than a preset threshold from all the pseudo sample data, and obtaining the target pseudo sample through contrastive learning constraint processing; the contrastive learning constraint is used to narrow the semantic distance between the pseudo sample data and the fault label result.

[0062] It can be understood that in the small sample scene of industrial equipment fault detection, there are problems of insufficient training and poor generalization of directly training the model due to the extremely small number of real fault samples. Pseudo samples are needed to supplement the data. In order to avoid pseudo samples from deviating from the industrial actual scene, the domain prior data of equipment failure needs to be obtained first. The domain prior data can include the physical mechanism of equipment failure, the failure mode and the statistical law. The physical mechanism is that bearing wear will be accompanied by a specific frequency vibration. The failure mode is that gear tooth breakage has an impact pulse. The statistical law is that the fault current fluctuation range.

[0063] The above-mentioned autoencoder can include a variational autoencoder (CVAE) or a generative adversarial network (GAN), and can also be used to enhance sample distribution by means of graph enhancement, contrastive learning, random interpolation (such as Mixup), and small sample synthesis method (such as SMOTE). When constructing small sample pseudo samples based on domain prior data, the domain prior data is converted into specific constraint parameters, including conditional constraints, latent variable distribution constraints, and multimodal consistency rules. The conditional constraints correspond to the fault type, and the latent variable distribution constraints conform to the statistical characteristics. The constraint parameters are used as generation conditions, and the fault label results and fault sample data are input into the autoencoder for processing to guide the autoencoder to generate pseudo sample data that conforms to the physical and statistical characteristics of the industrial actual scene. For example, the conditional constraint is that the generated pseudo sample corresponds to a specific fault type, such as a motor overload sample that needs to meet the synchronous rise of current and temperature. The latent variable distribution constraint is that the pseudo sample value conforms to the statistical range of the real fault, such as the vibration peak value not exceeding the equipment limit. The multimodal consistency rule includes, for example, that the linkage of vibration, temperature and other signals conforms to the physical logic, and does not appear contradictory such as high current but low temperature. These pseudo samples can supplement the gap of real samples, balance the sample distribution, and provide the AI model with enough reliable learning materials, so as to finally improve the precision and generalization ability of fault identification in the small sample scene.

[0064] Taking the small sample scenario of equipment fault diagnosis as an example, the input parameters of CVAE or GAN can include a small number of real fault samples to provide the real sample-based feature template for the generation tool, ensuring that the pseudo sample is consistent with the real sample in terms of data format and basic features; It can also include constraint parameters transformed from domain prior, such as condition constraint parameters, latent variable distribution constraint parameters, and multimodal consistency rule parameters corresponding to “bearing wear fault”, which are embedded in the model structure of the generation tool (such as the condition layer of CVAE and the generator loss function of GAN), thereby limiting the generation range of pseudo samples; It can also include random noise or latent variables, such as the latent variable vector of CVAE and the initial noise vector of the GAN generator, to provide randomness for generating diverse pseudo samples and avoid complete repetition of generated pseudo samples. At the same time, the random noise will be transformed into feature signals that meet the industrial rules under the guidance of the constraint parameters, outputting pseudo-synthetic fault sample feature vectors, i.e. pseudo sample data. Among them, the output pseudo sample has the same format and dimension as the real fault sample data, and meets the constraint parameters corresponding to the domain prior data, which not only has the feature attributes of the real fault sample, but also supplements the quantity gap of the real sample, providing sufficient and reliable materials for subsequent AI model training.

[0065] After generating pseudo sample data, the pseudo sample quality control mechanism can be used to control the quality of the pseudo sample and obtain the target pseudo sample. Each pseudo sample corresponding to each device can be configured with a corresponding weight value and sample parameter. The weight value is used to represent the correlation strength between the sample and the real fault mode, and the sample parameter is used to represent the key attributes such as fault duration and impact range. In order to ensure the quality of the pseudo sample, the confidence of each pseudo sample needs to be calculated according to the weight value and the sample parameter, so as to measure its matching degree with the real fault feature, and filter out low-quality samples with a confidence lower than a preset threshold. Subsequently, the remaining samples are further optimized through contrast learning constraints: by constructing a loss function to narrow the distance between the pseudo sample data and the corresponding fault label result in the semantic space, the correlation between the pseudo sample and the real fault class is strengthened, thereby obtaining the target pseudo sample. For example, the “bearing wear” class pseudo sample is closer to the real bearing wear label feature in the feature space, and the final target pseudo sample not only retains the typical features of industrial faults, but also has consistent distribution characteristics with real samples, which can effectively expand the training data set and improve the generalization ability and diagnosis accuracy of the fault detection model in the small sample scenario.

[0066] In order to improve the model precision and generalization ability, industry knowledge can be injected during data training, such as obtaining an industry knowledge base, and a rule engine, an expert system, a domain ontology graph, and a fault reasoning tree can be used as priori representations to improve the model explanation ability and practicality. Taking the industry knowledge base as an example, fault codes, operation and maintenance rules, and fault and component causal relationship graphs are extracted from the industry knowledge base, aligned and encoded into learnable knowledge vectors K∈R d , d represents the dimension of the knowledge vector, which is used to define the length of the feature in the mathematical space, and R is a real set. Among them, the fault code is, for example, F01 = current over limit, the operation and maintenance rule is, for example, “temperature > 80℃ needs to be shut down”, and the fault and component causal relationship graph is, for example, “bearing wear → main shaft vibration anomaly”, and the knowledge vector corresponding to the fault code F01 is [1, 0, 0,...]. The knowledge vector, the fault sample data and the target pseudo sample are mixed as the training data of the initial model to solve the problems of insufficient sample quantity and unbalanced categories, and are divided into a training set and a validation set according to a preset division rule, for example, 8:2.

[0067] The initial model includes an initial structure perception enhancement module, an initial multi-modal fusion module and an initial fault classification module connected in sequence. The training set is input into the initial model to obtain a fault output result, including: the structure sample data of the training set is subjected to structure enhancement processing by the initial structure perception enhancement module to obtain a structure sample vector; other sample data in the training set and the structure sample vector are subjected to feature extraction, modal missing repair and semantic alignment processing by the initial multi-modal fusion module to obtain a joint sample vector; and the joint sample vector is subjected to classification processing by the initial fault classification module to obtain the fault output result.

[0068] The training set can include structure sample data, time sequence sample data, text sample data and image sample data. The structure sample data is subjected to structure enhancement processing by the initial structure perception enhancement module. A structure vector is first generated, and is interactively enhanced with modal features through an attention fusion mechanism. A hierarchical position code is added to the nodes of the device topology graph, and is added to or spliced with the original features of the nodes. The model can perceive the functional hierarchy and relative position relationship of the components during feature calculation, thereby outputting a structure sample vector. The time sequence sample data, the text sample data and the image sample data are subjected to feature extraction by the initial multi-modal fusion module. After being encoded by a corresponding encoder and subjected to modal missing repair processing, the results of different modalities are aligned in a common semantic space by a semantic alignment unit to obtain a joint sample vector. Then, the joint sample vector is subjected to classification processing by the initial fault classification module to obtain a fault output result.

[0069] Optionally, for each type of failure, the class center vector Ck is calculated based on the real failure sample data and the target pseudo sample, for example, the class center vector of "bearing wear" can be the mean of all sample features of this class, and the distance (Euclidean distance / cosine distance) between the to-be-detected sample and Ck is calculated to determine the failure category, thereby improving the recognition accuracy of boundary samples. A loss function is constructed according to the fault output result and the fault label result, and the loss function can include a classification loss, a feature consistency loss, and a prototype constraint loss. The classification loss is used to optimize the fault category prediction accuracy, the feature consistency loss is used to ensure the semantic consistency after the multi-modal feature fusion, and the prototype constraint loss is used to reduce the distance between the sample and the class center, so that the model quickly converges under a small sample and does not bias. According to the loss function minimization, a small sample learning optimization strategy is adopted to iteratively optimize the parameters of the initial structure perception enhancement module, the initial multi-modal fusion module, and the initial fault classification module in the initial model to obtain a to-be-verified model, and the verification set is input into the to-be-verified model for verification to obtain a fault detection model. The small sample learning optimization strategy can also use a matching network, a relationship network, a same-class contrast loss, TripletLoss, or an attention mechanism-based metric network to complete similarity-driven learning.

[0070] It can be understood that for a new device with similar structure features, the meta-learning module or the Adapter module can be used to migrate and generalize the model without retraining, so that the model can quickly adapt to new data with very few new data and maintain high performance, thereby greatly reducing the training cost and deployment time. The meta-learning module, for example, MAML algorithm, quickly fine-tunes the model parameters with a small amount of new samples; the Adapter module, for example, can be a lightweight fully connected layer inserted, which only fine-tunes the layer to adapt to the new device.

[0071] In the embodiment, industry fault codes, maintenance experience, knowledge graph, etc. are used to construct knowledge vectors as guidance for feature extraction, so that the model output is closer to expert experience, reduces false correlation and "biased learning", and improves the result reliability; and through domain prior constraint, it is ensured that the generated pseudo sample conforms to the industrial physical law, avoids misleading the model by false samples, and supplements the sample size, so that the main model has enough data to learn the fault features. By taking the knowledge vector as part of the training data, the training data can be closer to the actual properties of the device, and through the small sample optimization strategy, the model can efficiently learn and flexibly migrate under a small amount of data, solving the problem of weak generalization ability in a small sample scenario. Further, after the fault detection model is constructed through training, it can be processed for lightweight: through knowledge distillation, quantization, compression of model size, and the like, the lightweight model can be deployed on the edge device in the industrial field and run on the edge side device, so that real-time fault diagnosis is realized, and delay and security risks of data uploading to the cloud are avoided. The quantization can be converting 32-bit floating point numbers to 16-bit, the compression of model size can refer to compressing the size from the original 1GB to 200MB, and the edge side device can include PLC, industrial gateway, edge server, and the like.

[0072] Optionally, the edge device can collect new data in real time, the new data including new fault samples and new working condition data, and the collected new data is flowed back to the cloud; the cloud model updates the training regularly with the flowed back data, optimizes the model parameters, adapts to new working conditions and new fault types, and realizes continuous iteration of the model. In the embodiment, the fault detection model is deployed to the edge device in the industrial scene through lightweight, the real-time requirement and hardware adaptation requirement of the industrial field are met, the model is continuously optimized with the device running through continuous learning, and the decline of diagnosis accuracy caused by working condition change is avoided. Through the construction of target pseudo samples, the real shortage problem is solved, the diversity of training data is expanded, the class center prototype network is introduced, the discrimination of the model under the boundary fuzzy category is improved, a multi-task loss function is constructed, the adaptability and convergence speed of the model to small sample classes are obviously enhanced under the guidance of the multi-task loss function, and thus the trained fault detection model has high accuracy and strong generalization ability.

[0073] Based on the same inventive concept, the embodiment of the present application also provides an industrial equipment fault detection device for realizing the above-mentioned industrial equipment fault detection method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, and therefore the specific limitations in one or more industrial equipment fault detection device embodiments provided below can refer to the limitations of the industrial equipment fault detection method described above, which will not be described here again.

[0074] In one exemplary embodiment, as shown in Figure 5 a vertical domain large model-based industrial equipment fault detection device is provided, which comprises: An acquisition module 510 is configured to acquire multi-modal data to be detected in the running process of a target industrial equipment; A preprocessing module 520 is configured to preprocess the multi-modal data to be detected to obtain processed data; the processed data includes structural data, time series data, image data, and text data. The fault detection module 530 is configured to process the processed data through a fault detection model to obtain a fault category result. The fault detection model comprises a structure perception enhancement module and a multi-modal fusion module connected with each other. The fault detection model is trained based on fault sample data and target pseudo sample data. The fault sample data are partially labeled with fault label results. The target pseudo sample data are obtained by taking the acquired domain prior data as constraint parameters and based on the fault label results and the fault sample data. The structure perception enhancement module is configured to perform feature enhancement processing on the structure data to obtain a structure feature vector. The multi-modal fusion module is configured to perform semantic alignment and fusion processing on the structure feature vector, time series data, image data, and text data.

[0075] As an optional implementation, the fault detection module 530 is specifically configured to: perform feature extraction and component relationship strengthening processing on the structure data through the structure perception enhancement module to obtain a structure feature vector; perform feature extraction, modal missing repair, and semantic alignment processing on the image data, the time series data, and the text data through the multi-modal fusion module to obtain a semantic consistent joint representation vector; perform fault classification processing on the joint representation vector through the fault classification module to obtain a fault category result.

[0076] As an optional implementation, the fault detection module 530 is further configured to: perform topology modeling on the structure data to construct a topology graph. The topology graph comprises nodes and edges. The nodes are configured to represent component information in the device, and the edges are configured to represent structure information between components. perform encoding processing on the structure information through a graph neural network to obtain a structure embedding vector; for hierarchical structure data in the structure data, obtain hierarchical information of a subsystem where a component is located and position information of the component in the topology graph; generate hierarchical position encoding according to the hierarchical information and the position information; add the hierarchical position encoding to the component information and perform fusion processing on the hierarchical position encoding and the component information to obtain a structure fusion vector; perform attention allocation processing on the structure fusion vector and the structure embedding vector through an attention layer to obtain a structure feature vector.

[0077] As an optional implementation, the fault detection module 530 is further configured to: perform processing on the image data, the time series data, and the text data through corresponding encoders respectively to extract an image feature vector, a time series feature vector, and a text feature vector; The missing vector is determined from the structure feature vector, the image feature vector, the time sequence feature vector and the text feature vector by a missing repair unit, and the missing vector is modality repaired to obtain a target vector; the target vector includes a target image vector, a target time sequence vector, a target text vector and a target structure vector; The target image vector, the target time sequence vector, the target text vector and the target structure vector are aligned in a common semantic space by a semantic alignment unit to extract a joint representation vector consistent in semantics.

[0078] As an optional implementation, the apparatus is further configured to: Obtain the target pseudo sample and the original equipment data, and pre-process the original equipment data into fault sample data; the fault sample data includes first data labeled with a fault label result, second data without labeling, and equipment operation data; Obtain equipment attribute data from a preset industry knowledge base, and vectorize the equipment attribute data into a knowledge vector; Divide the knowledge vector, the fault sample data and the target pseudo sample into a training set and a validation set according to a preset division rule; Input the training set into an initial model to obtain a fault output result; According to the fault output result and the fault label result, a loss function is constructed, and parameters of each module in the initial model are iteratively optimized according to a small sample learning optimization strategy to obtain a to-be-verified model; the loss function includes a combined classification loss, a feature consistency loss and a prototype constraint loss; Input the validation set into the to-be-verified model for verification to obtain a fault detection model.

[0079] As an optional implementation, the apparatus is further configured to: Obtain domain prior data, and convert the domain prior data into constraint parameters; the constraint parameters include conditional constraint parameters, latent variable distribution constraint parameters and multi-modal consistency rule parameters; Input the fault label result and the fault sample data into a self-encoder for processing by taking the constraint parameters as generation conditions to obtain pseudo sample data; the pseudo sample data is configured with a weight value and a sample parameter; According to the weight value and the sample parameter, a confidence is calculated, pseudo sample data with a confidence less than a preset threshold is filtered from all the pseudo sample data, and target pseudo sample data is obtained through contrast learning constraint processing; the contrast learning constraint is used to narrow the semantic distance between the pseudo sample data and the fault label result.

[0080] As an optional implementation, the apparatus is further configured to: The structural sample data of the training set is subjected to structural enhancement processing through an initial structure perception enhancement module to obtain a structural sample vector; The other sample data in the training set and the structural sample vector are subjected to feature extraction, mode missing repair and semantic alignment processing through an initial multi-modal fusion module to obtain a joint sample vector; The joint sample vector is subjected to classification processing through an initial fault classification module to obtain a fault output result.

[0081] As an optional implementation, the preprocessing module 520 is specifically configured to: The multi-modal data to be detected is subjected to timestamp alignment processing to obtain aligned data; The aligned data is subjected to null value filling processing to obtain filled data; the filled data includes time series filling data and other filling data; The time series filling data is subjected to segmentation and slicing processing to obtain a plurality of continuous segments of fixed lengths; The continuous segments and the other filling data are subjected to feature standardization processing to obtain processed data.

[0082] The industrial equipment fault detection device provided by the embodiments of the present application can accurately cover the characteristics of complex data modalities in an industrial scene by obtaining multi-dimensional multi-modal data to be detected including structure, time sequence, image, text and the like of a target industrial equipment and performing preprocessing, thereby providing a comprehensive and device-actual-operation-state-adapted data source for subsequent diagnosis; and introducing field prior data as a constraint parameter in the model training stage, generating target pseudo samples in combination with fault sample data with partially labeled fault labels, effectively making up for the short board of a small number of fault data samples in the industrial scene, and greatly reducing the false alarm risk caused by insufficient samples; and the structure perception enhancement module and the multi-modal fusion module are connected with each other in the fault detection model, the structure perception enhancement module can perform feature enhancement on structural data and generate a structural feature vector, thereby deeply mining the structural characteristic information of the image target industrial equipment, effectively adapting to the structural characteristics of complex equipment, and breaking the limitation that a traditional general model is difficult to match the structure of complex equipment; the multi-modal fusion module performs semantic alignment and fusion processing on the structural feature vector and the other three types of data, realizes deep correlation and information complementation of multi-dimensional data, avoids information deviation of single modal data, and further significantly improves the fault recognition precision, thereby providing more reliable fault diagnosis support for stable operation of industrial equipment.

[0083] In an exemplary embodiment, a computer device, which can be a server or a terminal, is provided, and an internal structure diagram of the computer device can be as shown in Figure 6As shown in the figure. The computer device includes a processor, a memory, an input / output interface (I / O for short) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through the system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capability. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store video tag processing data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through network connection. The computer program is executed by the processor to implement an industrial equipment fault detection method based on a vertical domain large model.

[0084] Those skilled in the art can understand that, Figure 6 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0085] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps in each of the method embodiments described above.

[0086] In one exemplary embodiment, a computer readable storage medium is provided, storing a computer program, which is executed by a processor to implement the steps in each of the method embodiments described above.

[0087] In one exemplary embodiment, a computer program product is provided, including a computer program, which is executed by a processor to implement the steps in each of the method embodiments described above.

[0088] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.

[0089] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to a memory, a database or other medium used in the embodiments provided in the present application can include at least one of a non-volatile and a volatile memory. The non-volatile memory can include a read-only memory (ROM), a magnetic tape, a floppy disk, a flash memory, an optical storage, a high-density embedded non-volatile memory, a resistive random access memory (ReRAM), a magnetoresistive random access memory (MRAM), a ferroelectric random access memory (FRAM), a phase change memory (PCM), a graphene memory, etc. The volatile memory can include a random access memory (RAM) or an external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as a static random access memory (SRAM) or a dynamic random access memory (DRAM), etc.

[0090] The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a blockchain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.

[0091] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combinations of the technical features do not exist contradictory, they should be considered as the scope of the present application.

[0092] The principles and implementation modes of the present application are described by applying specific examples herein, and the above-mentioned embodiments are only used to help understand the method and its core idea of the present application; meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range can be changed. In conclusion, the content of the present application should not be understood as a limitation.

Claims

1. A method for fault detection of industrial equipment based on a large vertical domain model, characterized in that, The industrial equipment fault detection method based on a vertical domain large model includes: Acquire multimodal data to be detected during the operation of the target industrial equipment; The multimodal data to be detected is preprocessed to obtain processed data; the processed data includes: structural data, time series data, image data, and text data. The processed data is processed through a fault detection model to obtain fault category results. The fault detection model includes an interconnected structure-aware enhancement module and a multimodal fusion module. The fault detection model is trained based on fault sample data and target pseudo-samples. Some of the fault sample data is labeled with fault tags. The target pseudo-samples are obtained by using acquired domain prior data as constraint parameters and based on the fault tag results and fault sample data. The structure-aware enhancement module is used to perform feature enhancement processing on the structure data to obtain a structure feature vector. The multimodal fusion module is used to perform semantic alignment and fusion processing on the structure feature vector, the time-series data, the image data, and the text data.

2. The industrial equipment fault detection method based on a large vertical domain model according to claim 1, characterized in that, The fault detection model further includes: a fault classification module; the fault classification module is connected to the multimodal fusion module; The processed data is then processed through a fault detection model to obtain fault category results, including: The structural awareness enhancement module performs feature extraction and component relationship enhancement processing on the structural data to obtain a structural feature vector. The multimodal fusion module performs feature extraction, modality missing repair, and semantic alignment on the image data, time-series data, and text data to obtain a semantically consistent joint representation vector. The joint representation vector is processed by the fault classification module to obtain the fault category result.

3. The industrial equipment fault detection method based on a large vertical domain model according to claim 2, characterized in that, The structure-aware enhancement module performs feature extraction and component relationship enhancement on the structural data to obtain a structural feature vector, including: The structural data is used to perform topological modeling to construct a topological graph; the topological graph includes nodes and edges, the nodes are used to represent component information in the device, and the edges are used to represent structural information between the components; The structural information is encoded using a graph neural network to obtain a structural embedding vector; For the hierarchical structure data in the structural data, obtain the hierarchical information of the subsystem where the component is located and its position information in the topology diagram; Generate a hierarchical location code based on the hierarchical information and the location information; The hierarchical position code is added to the component information and fused with the component information to obtain a structural fusion vector; The structural fusion vector and the structural embedding vector are processed by attention layer for attention allocation to obtain the structural feature vector.

4. The industrial equipment fault detection method based on a large vertical domain model according to claim 2, characterized in that, The multimodal fusion module includes: an encoder, a missing component repair unit, and a semantic alignment unit connected in sequence; The multimodal fusion module performs feature extraction, modality missing repair, and semantic alignment on the image data, time-series data, and text data to obtain a semantically consistent joint representation vector, including: The image data, time-series data, and text data are processed by their respective encoders to extract image feature vectors, time-series feature vectors, and text feature vectors. The missing vector is determined from the structural feature vector, image feature vector, temporal feature vector, and text feature vector by the missing vector repair unit, and modal repair processing is performed on the missing vector to obtain the target vector; the target vector includes: target image vector, target temporal vector, target text vector, and target structural vector; The semantic alignment unit aligns the target image vector, target temporal vector, target text vector, and target structure vector in a common semantic space, and extracts a semantically consistent joint representation vector.

5. The industrial equipment fault detection method based on a large vertical domain model according to claim 2, characterized in that, The fault detection model is constructed through the following steps: Obtain target pseudo-samples and original equipment data, and preprocess the original equipment data into fault sample data; the fault sample data includes first data labeled with fault tags, second data without labels, and equipment operation data; Obtain equipment attribute data from a preset industry knowledge base, and vectorize the equipment attribute data into knowledge vectors; The knowledge vector, the fault sample data, and the target pseudo-sample are divided into a training set and a validation set according to a preset partitioning rule; The training set is input into the initial model to obtain the fault output result; Based on the fault output and fault label results, a loss function is constructed. By minimizing the loss function, a few-shot learning optimization strategy is used to iteratively optimize the parameters of each module in the initial model to obtain the model to be validated. The loss function includes: combined classification loss, feature consistency loss, and prototype constraint loss. The validation set is input into the model to be validated for validation to obtain the fault detection model.

6. The industrial equipment fault detection method based on a large vertical domain model according to claim 5, characterized in that, Obtaining target pseudo-samples includes: Acquire domain prior data and transform the domain prior data into constraint parameters; the constraint parameters include: conditional constraint parameters, latent variable distribution constraint parameters, and multimodal consistency rule parameters. Using the constraint parameters as generation conditions, the fault label results and fault sample data are input into the autoencoder for processing to obtain pseudo sample data; the pseudo sample data is configured with weight values ​​and sample parameters. Based on the weight values ​​and sample parameters, the confidence level is calculated, and pseudo-sample data with confidence levels less than a preset threshold are filtered out from all pseudo-sample data. The target pseudo-sample is obtained by processing it through contrastive learning constraints. The contrastive learning constraints are used to narrow the semantic distance between the pseudo-sample data and the fault label results.

7. The industrial equipment fault detection method based on a large vertical domain model according to claim 5, characterized in that, The initial model includes: an initial structure-aware enhancement module, an initial multimodal fusion module, and an initial fault classification module connected in sequence; the training set is input into the initial model to obtain fault output results, including: The structural sample data of the training set is processed by the initial structure-aware enhancement module to obtain a structural sample vector. The other sample data in the training set and the structured sample vector are processed by the initial multimodal fusion module for feature extraction, modality missing repair and semantic alignment to obtain a joint sample vector; The joint sample vector is classified by the initial fault classification module to obtain the fault output result.

8. The industrial equipment fault detection method based on a large vertical domain model according to claim 1, characterized in that, The multimodal data to be detected is preprocessed to obtain processed data, including: The multimodal data to be detected is timestamped to obtain aligned data. The aligned data is then padded with null values ​​to obtain padded data; the padded data includes time-series padded data and other padded data. The time-series filling data is segmented and sliced ​​to obtain multiple continuous segments of fixed length; The processed data is obtained by performing feature standardization processing based on the continuous segments and the other imputed data.

9. An industrial equipment fault detection device based on a large vertical domain model, characterized in that, The industrial equipment fault detection device based on a vertical domain large model includes: The acquisition module is used to acquire the multimodal data to be detected during the operation of the target industrial equipment; The preprocessing module is used to preprocess the multimodal data to be detected to obtain processed data; the processed data includes: structural data, time series data, image data and text data; A fault detection module is used to process the processed data through a fault detection model to obtain fault category results. The fault detection model includes an interconnected structure-aware enhancement module and a multimodal fusion module. The fault detection model is trained based on fault sample data and target pseudo-samples. Some of the fault sample data is labeled with fault tags. The target pseudo-samples are obtained by using acquired domain prior data as constraint parameters and based on the fault tag results and fault sample data. The structure-aware enhancement module is used to perform feature enhancement processing on the structure data to obtain a structure feature vector. The multimodal fusion module is used to perform semantic alignment and fusion processing on the structure feature vector, the time-series data, the image data, and the text data.

10. A computer device, comprising: The memory and processor contain a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the steps of the industrial equipment fault detection method based on a large vertical domain model as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Sensor fault diagnosis method based on single-domain generalization under uncertainty guidance adversarial enhancement domain

    CN119377669A

  • Finite data-oriented semi-supervised rolling bearing cross-domain fault diagnosis method

    CN119901491A

  • Industrial data identification method, device, equipment, medium and product

    CN120148043A

  • Communication fault identification method and device based on cross-modal fusion, and electronic equipment

    CN120455241A

  • Power plant intelligent maintenance method and system based on multi-modal dynamic graph learning

    CN120494806A

Cited By

  • Railway safety monitoring method and device based on multi-modal fusion, equipment and medium

    CN121959260A

  • Power transmission line fault unmanned aerial vehicle inspection and intelligent diagnosis method and system based on visual Transform

    CN122023380A

  • Unmanned aerial vehicle inspection and intelligent diagnosis method and system for power transmission line fault based on visual transformer

    CN122023380B