Fault diagnosis method of industrial equipment and related device
By employing predictive models and multimodal descriptive models, the problem of low accuracy and efficiency of manual analysis in industrial equipment fault diagnosis is solved, achieving a transparent and causal fault diagnosis process that is understandable to non-professional users.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-03
- Publication Date
- 2026-03-10
AI Technical Summary
Current industrial equipment fault diagnosis relies on manual analysis, which suffers from insufficient accuracy and low efficiency, making it difficult to balance cost and efficiency.
By employing an operational prediction model and a multimodal descriptive model, and through feature extraction and fusion of the monitoring signal interface of industrial equipment, the expert click path is reproduced, and a natural language feature description is generated, thus achieving a transparent and causal fault diagnosis process.
It enables non-professional users to intuitively understand the source of fault diagnosis results and its supporting evidence, and has the ability to go through the entire process from multi-round interactive operation prediction to final diagnostic reasoning, simulating the complete analysis process of diagnostic engineers.
Smart Images

Figure CN121638459A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of fault monitoring, in particular, to a fault diagnosis method of industrial equipment and a related device. BACKGROUND
[0002] At present, with the continuous development of intelligent manufacturing, the intelligent degree of industrial systems is getting higher and higher, and at the same time, it is becoming more and more complex, and the loss caused by the damage of industrial equipment is also getting bigger and bigger. Therefore, it is necessary to intelligently monitor the industrial equipment to ensure the safe and stable operation of the equipment. Fault diagnosis of industrial equipment is an important link in intelligent monitoring.
[0003] The current mainstream monitoring method still needs manual analysis of data and relies on manual monitoring. The monitoring method based on manual diagnosis has two obvious limitations. First, manual diagnosis is limited by the professional level, experience difference and work intensity of the diagnosis personnel, and is subjective, which may result in missed detection of equipment failure, thereby posing a hidden danger to the safe operation of the equipment. Second, the manual diagnosis method has problems in the training cycle of diagnosis personnel and the efficiency of manual processing, which greatly limits the number of monitored equipment and the speed of business expansion.
[0004] Under this circumstance, how to ensure the accuracy of the diagnosis result in the monitoring process while taking into account the cost and efficiency has become a difficult problem for those skilled in the art. SUMMARY
[0005] The purpose of the present application is to provide a fault diagnosis method of industrial equipment and a related device to improve the above problems.
[0006] In order to achieve the above purpose, the technical solutions adopted by the embodiments of the present application are as follows: In a first aspect, the embodiments of the present application provide a fault diagnosis method of industrial equipment, which comprises: inputting a current display interface and its interface description into an operation prediction model to perform operation prediction, so as to obtain current operation prediction information, wherein the current display interface is an interface for displaying monitoring signals of industrial equipment, the interface description is a summary description of the monitoring signals displayed in the current display interface related to fault analysis, and the current operation prediction information includes a current expansion operation, its corresponding operation data type and operation execution position; obtaining a new current display interface according to the current prediction information; inputting the new current display interface into a multi-modal description derivation model to perform feature extraction and fusion on the current display interface, so as to obtain the interface description of the current display interface, and repeatedly inputting the current display interface and its interface description into the operation prediction model as input to perform operation prediction, so as to obtain the current operation prediction information. After the observation phase, the target sequence is used as input to the fault reasoning model to obtain the fault diagnosis results. The target sequence includes the operation data type and interface description corresponding to each display interface.
[0007] Secondly, embodiments of the present invention provide a fault diagnosis device for industrial equipment, comprising: The first processing unit is used to input the current display interface and its interface description into the operation prediction model to perform operation prediction in order to obtain current operation prediction information. The current display interface is an interface used to display monitoring signals of industrial equipment. The interface description is a summary description of the monitoring signals displayed in the current display interface and related to fault analysis. The current operation prediction information includes the current operation, its corresponding operation data type, and the operation execution position. The first processing unit is further configured to obtain a new current display interface based on the current prediction information; The first processing unit is further configured to input the new current display interface into the multimodal description derivation model to extract and fuse features of the current display interface to obtain the interface description of the current display interface, and repeatedly use the current display interface and its interface description as input to the operation prediction model to perform operation prediction to obtain the current operation prediction information. The second processing unit is used to take the target sequence as input to the fault reasoning model after the observation phase ends in order to obtain the fault diagnosis result. The target sequence includes the operation data type and interface description corresponding to each display interface.
[0008] Thirdly, embodiments of the present invention provide a storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method.
[0009] Fourthly, embodiments of the present invention provide an electronic device, the electronic device comprising: a processor and a memory, the memory being used to store one or more programs; when the one or more programs are executed by the processor, the above-described method is implemented.
[0010] Compared to existing technologies, the fault diagnosis method and related apparatus for industrial equipment provided in this embodiment of the invention inputs the current display interface and its description into an operation prediction model to perform operation prediction and obtain current operation prediction information. The current display interface is used to display monitoring signals of the industrial equipment, and the interface description is a summary description of the monitoring signals displayed in the current display interface and related to fault analysis. The current operation prediction information includes the currently unfolded operation, its corresponding operation data type, and the operation execution position. A new current display interface is obtained based on the current prediction information. The new current display interface is input into a multimodal description derivation model to extract and fuse features of the current display interface to obtain an interface description. The current display interface and its interface description are repeatedly used as input to the operation prediction model to perform operation prediction and obtain current operation prediction information. After the observation phase, the target sequence is used as input to a fault inference model to obtain the fault diagnosis result. The target sequence includes the operation data type and interface description corresponding to each display interface. By predicting operations using an operational prediction model, the system reproduces the expert's click path and generates natural language feature descriptions using a multimodal model. This makes the diagnostic process transparent and causal, allowing users to intuitively understand the source of the diagnostic results and its supporting evidence. The system realizes the entire process from multi-round interactive operation prediction and feature analysis description to final diagnostic reasoning, and has the ability to reason about faults and generate conclusions.
[0011] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0012] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.
[0014] Figure 2 This is a flowchart illustrating the fault diagnosis method for industrial equipment provided in an embodiment of the present invention.
[0015] Figure 3 This is a schematic diagram illustrating the training process of the operation prediction model provided in an embodiment of the present invention.
[0016] Figure 4 This is a schematic diagram illustrating the training process of the derivation model provided in an embodiment of the present invention.
[0017] Figure 5 This is a schematic diagram of a fault diagnosis device for industrial equipment provided in an embodiment of the present invention.
[0018] In the diagram: 10-Processor; 11-Memory; 12-Bus; 13-Communication interface; 501-First processing unit; 502-Second processing unit. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0020] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0021] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0022] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0023] In the description of this invention, it should be noted that the terms "upper," "lower," "inner," "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship in which the product of this invention is usually placed when in use. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting this invention.
[0024] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set" and "connection" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0025] The following detailed description of some embodiments of the present invention is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0026] This invention provides an electronic device, which may be a server device, a computer device, or a mobile phone device, etc. Please refer to... Figure 1 This is a schematic diagram of the structure of an electronic device. The electronic device includes a processor 10, a memory 11, and a bus 12. The processor 10 and the memory 11 are connected via the bus 12. The processor 10 is used to execute executable modules, such as computer programs, stored in the memory 11.
[0027] Processor 10 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the fault diagnosis method for industrial equipment can be completed through the integrated logic circuits in the hardware or software instructions within processor 10. The aforementioned processor 10 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0028] The memory 11 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage.
[0029] Bus 12 can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. Figure 1 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus 12 or one type of bus 12.
[0030] The memory 11 is used to store programs, such as programs corresponding to fault diagnosis devices for industrial equipment. The fault diagnosis device for industrial equipment includes at least one software functional module that can be stored in the memory 11 in the form of software or firmware, or embedded in the operating system (OS) of the electronic device. Upon receiving an execution instruction, the processor 10 executes the program to implement the fault diagnosis method for the industrial equipment.
[0031] The electronic device provided in this embodiment of the invention may further include a communication interface 13. The communication interface 13 is connected to the processor 10 via a bus.
[0032] It should be understood that, Figure 1 The structure shown is only a partial schematic diagram of the electronic device; the electronic device may also include components that are larger than... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown. Figure 1 The components shown can be implemented using hardware, software, or a combination thereof.
[0033] The fault diagnosis method for industrial equipment provided in this embodiment of the invention can be applied to, but is not limited to, [various applications]. Figure 1 For the specific process of the electronic devices shown, please refer to [link / reference]. Figure 2 The fault diagnosis methods for industrial equipment include: S11, S12, S13, S14 and S15, which are described in detail below.
[0034] S11. Input the current display interface and its description into the operation prediction model to perform operation prediction and obtain the current operation prediction information.
[0035] The current display interface showcases monitoring signals from industrial equipment. The interface description provides a summary of the monitoring signals and fault analysis information displayed on the current screen. The current operation prediction information includes the current expanded operation, its corresponding data type, and the operation execution location. The operation data type refers to the type of data selected for the expanded operation, which can be any of the following: monitoring signal indicator trend, monitoring waveform signal, monitoring spectrum signal, or monitoring envelope spectrum signal. The operation execution location is the coordinate position of the selected object within the display interface.
[0036] The expand operation is used to expand the selected signal data. It should be understood that the initial display interface needs to show the trend of the monitored signal indicators and the interface description that triggers fault analysis. The interface description of the initial display interface can be generated by the alarm device and the location of the measuring point, such as "An alarm has occurred at a certain measuring point of a certain device, and fault diagnosis is required."
[0037] Expanding operations can be, but are not limited to, mouse clicks, mouse zooming, mouse trajectory drawing, voice commands, and mouse hover operations. When the data type of the expanding operation is an indicator trend, and the signal data also includes any one or more of the monitored waveform signal, monitored spectrum signal, and monitored envelope spectrum signal, the monitored waveform signal, monitored spectrum signal, and monitored envelope spectrum signal corresponding to the position where the expanding operation is executed will be expanded. When the data type of the expanding operation is any one of the monitored waveform signal, monitored spectrum signal, and monitored envelope spectrum signal, the data at the operation execution position will be scaled.
[0038] When the monitored signal is a vibration signal, the corresponding monitored signal includes the vibration signal index trend, vibration waveform signal, vibration spectrum signal, and vibration envelope spectrum signal; when the monitored signal is a current signal, the corresponding monitored signal includes the current signal index trend, current waveform signal, current spectrum signal, and current envelope spectrum signal; when the monitored signal is a temperature signal, the corresponding monitored signal includes the temperature signal index trend.
[0039] In this embodiment of the invention, the operation prediction model may employ an agent or a reinforcement learning model.
[0040] S12, Based on the current prediction information, obtain the new current display interface.
[0041] S13, input the new current display interface into the multimodal description derivation model to extract and fuse features of the current display interface to obtain the interface description of the current display interface.
[0042] By describing and deriving a model, features of the current display interface are extracted and fused to obtain an interface description of the current display interface, which is used to simulate the annotations written by diagnostic engineers when observing the display interface.
[0043] S14 determines whether the observation phase has ended. If the observation phase has ended, execute S15; otherwise, execute S11.
[0044] Optionally, the end of the observation can be determined by whether the number of times the interface description is generated reaches a preset threshold. Alternatively, the observation can be determined to end when the operation prediction model stops making predictions and no new operation prediction information is obtained. If the observation phase has not ended, after obtaining a new interface description in S13, S11 needs to be executed again to repeatedly use the currently displayed interface and its interface description as input to the operation prediction model to perform operation predictions and obtain the current operation prediction information.
[0045] S15. After the observation phase ends, the target sequence is used as the input to the fault reasoning model to obtain the fault diagnosis result.
[0046] The target sequence includes the data type of operation and the interface description corresponding to each display interface. The fault inference model can be a pre-trained inference model based on open source (such as DeepSeek-R1 or Qwen3), which is then fine-tuned and trained using LoRa based on the fault diagnosis data designed above.
[0047] In one alternative implementation, the fault reasoning model can also simultaneously output maintenance suggestions corresponding to the fault diagnosis results.
[0048] In the fault diagnosis method for industrial equipment provided in this embodiment of the invention, the operation is predicted by the operation prediction model, the expert's click path is reproduced, and natural language feature description is generated by combining the multimodal model, so that the diagnosis process is transparent and causal, allowing users to intuitively understand the source of the diagnosis results and its supporting evidence, realizing the whole process from multi-round interactive operation prediction and feature analysis description to final diagnosis reasoning, and having the ability to reason about faults and generate conclusions.
[0049] It can simulate the complete process of fault analysis by diagnostic engineers and intuitively present the unfolding operation, intermediate analysis results and final diagnostic conclusions. It enables the algorithm to perform fully automatic analysis of signals generated during equipment operation and allows non-professional users to clearly understand the fault diagnosis methods and conclusions.
[0050] The following example illustrates the data processing of the fault reasoning model. For each complete interactive data, the data type of the operation corresponding to each display interface output by the operation prediction model and its corresponding interface description are extracted and used as input to the fault reasoning model. The fault reasoning model outputs the fault diagnosis result and its corresponding maintenance suggestions.
[0051] The following is an example of the data type (data_type) for the operation interface and its corresponding interface description (data_result): { "data_type": "Total value trend" and "data_result": "During the period from 12:00:00 on 2025-01-01 to 10:30:00 on 2025-09-24, the total value of low-frequency acceleration showed a slow but significant upward trend"}; {"data_type": "Waveform" and "data_result": "Periodic impacts are visible in the low-frequency acceleration waveform at 10:30:00 on 2025-9-24, with the impact frequency being the rotational frequency of the gearbox's first shaft."}; {"data_type": "Spectrum" and "data_result": "Gear meshing characteristics are visible in the low-frequency acceleration spectrum at 10:30:00 on 2025-09-24, including first-order meshing and its harmonic bands and shaft rotation frequency sidebands, with a rich number of sidebands."}; { "data_type": "Waveform" and "data_result": "No obvious abnormalities were found in the low-frequency acceleration waveform at 18:20:00 on 2025-06-10."}; {"data_type": "Spectrum" and "data_result": "No obvious anomalies were found in the low-frequency acceleration spectrum at 18:20:00 on 2025-06-10."}; The following is an example of the fault diagnosis results and corresponding maintenance suggestions output by the fault reasoning model: response": { "fault_reasoning": "Gear meshing characteristics have recently appeared in the spectrum, with a large number of sidebands. Combined with the periodic impacts of frequency intervals that appear from nothing to something in the waveform, it is speculated that this is related to severe gear wear failure."; "diagnose_result": "Severe gear wear fault exists"; "maintenance_suggestions": "1. During routine inspections, pay attention to abnormal noises from the gearbox and changes in lubricating oil temperature. 2. Check for damage to the gear teeth and any abnormalities in the meshing clearance as appropriate. Repair or replace as needed, depending on the inspection results."
[0052] Building upon the preceding discussion of how to train an operational prediction model to accurately reproduce the operational behavior of expert diagnostic engineers, this invention also provides an optional implementation method, which please refer to. Figure 3 , Figure 3This is a schematic diagram illustrating the training process of the operation prediction model provided in an embodiment of the present invention. The training process of the operation prediction model includes S21, S22, and S23, which are described in detail below.
[0053] S21. Obtain the six-tuple information corresponding to each display interface when the diagnostic engineer performs fault diagnosis. The six-tuple information includes a screenshot of the display interface, the engineer's annotation on the display interface (a summary description of the monitored signal performance after the diagnostic engineer observes the current interface), the expansion operation performed by the engineer on the display interface (click or zoom), the data type of the operation (any one of the following: monitored signal indicator trend, monitored waveform signal, monitored spectrum signal, and monitored envelope spectrum signal), the operation execution location (coordinate position in the display interface), and the purpose of the operation in the engineer's annotation. The purpose of the operation is added to help the model make correct predictions during the training process.
[0054] S22, arrange the obtained six-tuples in chronological order to obtain the first training sequence; S23, the training of the operation prediction model is completed based on the first training sequence.
[0055] The operation prediction model is trained to learn and reproduce the actual operation path of the diagnostic engineer. Based on open-source pre-trained large-scale graphical and textual modal models (such as MiniCPM-V and Qwen-VL), and combined with the behavioral modeling data designed above, LoRa fine-tuning training is performed to obtain the operation prediction model.
[0056] Based on the foregoing, regarding the content in S11, this embodiment of the invention also provides an optional implementation method, please refer to the following: S11, inputting the current display interface and its interface description into the operation prediction model to perform operation prediction, so as to obtain the current operation prediction information, including: S111, specifically as follows.
[0057] S11, The operation prediction model combines the historical display interface, the interface description of the historical display interface, the current display interface, and the interface description of the current display interface of the monitoring signals of this batch to perform operation prediction in order to obtain the current operation prediction information.
[0058] The following is an explanation of the six-tuples corresponding to the operation prediction model. The six-tuples include: the current display screen (cur_screenshot), the screen description (last_state_description), the currently expanded operation (action) and its corresponding action, the operation data type (data_type), the operation execution position (action), and the operation objective (objective).
[0059] The initial input to the operational prediction model is: "query": {"history": [], "cur_screenshot": " / data / screen_0.jpg","last_state_description": "An alarm occurred at a certain measuring point of a certain device, requiring fault diagnosis."}; The initial output of the operational prediction model is: "response": { "action": "Click", "place": "[30,100], [35,110]", "data_type": "Other", "objective": "Select analysis measurement point"}.
[0060] The subsequent inputs to the operational prediction model are: "query": {"history": [{"data_type": "Other", "data_result": "",},{"data_type": "Total Value Trend","data_result": "During the period from 12:00:00 on 2025-01-01 to 10:30:00 on 2025-09-24, the total value of low-frequency acceleration showed a slow but significant upward trend."},{"data_type": "Waveform","data_result": "Periodic impacts can be seen in the low-frequency acceleration waveform at 10:30:00 on 2025-09-24, with the impact frequency being the gearbox shaft frequency."},{"data_type": "Spectrum","data_result": "Gear meshing characteristics can be seen in the low-frequency acceleration spectrum at 10:30:00 on 2025-09-24, with primary meshing and its harmonic bands and shaft frequency sidebands, and the number of sidebands is abundant."},{"data_type": {"waveform", "data_result": "No obvious abnormalities were found in the low-frequency acceleration waveform at 18:20:00 on 2025-06-10."}, {"data_type": "Spectrum", "data_result": "No obvious abnormalities were found in the low-frequency acceleration spectrum at 18:20:00 on 2025-06-10."}]; "cur_screenshot": "The storage address of the current screenshot, such as: / data / screen_3.jpg"; "last_state_description": "No obvious abnormalities were found in the low-frequency acceleration spectrum at 18:20:00 on 2025-06-10."
[0061] The subsequent output of the operational prediction model is: "response": {"action": "scaling", "place": "[500,700], [600,750]", "data_type": "spectrum", "objective": "spectrum band scaling analysis"}.
[0062] The fields are explained as follows: query represents the model input information, and response represents the model output information. The training goal is to enable the model to generate the corresponding response based on the given query. The "action" field represents the interface expansion action, with two categories: click and zoom. The "place" field indicates the location of the operation on the top-left and bottom-right axes of the screen. "data_type" indicates the data type of the operation, with values including indicator trends (further subdivided based on indicator type, such as total value trend, impact ratio trend, etc.), waveform, spectrum, and envelope spectrum. "objective" indicates the purpose of the operation, manually labeled in the training set, such as selecting analysis points or scaling analysis of the spectrum band, used by the model to learn and understand the meaning of this action.
[0063] In one optional implementation, the description derivation model includes a signal encoder, a visual encoder, a text encoder, and a language model. Based on this, regarding the content of S13, this embodiment of the invention also provides an optional implementation, please refer to the following. S13, the new current display interface is input into the multimodal description derivation model to extract and fuse features of the current display interface to obtain an interface description of the current display interface, including: S131 to S134, specifically described below.
[0064] S131, input the raw data of the monitoring signal corresponding to the current display interface into the signal encoder to obtain the signal characteristics.
[0065] S132, input the screenshot of the current display interface into the visual encoder to obtain image features.
[0066] S132, input the relevant text information of the observation equipment (corresponding to the monitoring signal) in the current display interface into the text encoder to obtain text features.
[0067] S133 maps signal features, image features, and text features to the high-dimensional space of the language model, performs cross-modal feature alignment, and obtains multi-source signal features.
[0068] S134: Input the multi-source signal features into the language model for understanding and representation to obtain the interface description of the currently displayed interface.
[0069] By combining the generative capabilities of large language models, it achieves a comprehensive understanding and representation of the features of multi-source signals, and has the ability to automatically generate textual descriptions of signal representations based on multimodal inputs.
[0070] Optionally, the signal encoder includes a waveform signal encoder and a spectrum signal encoder (spectral signal encoder or envelope spectrum signal encoder), and the signal features include waveform features and spectrum features. S131, the raw data of the monitoring signal corresponding to the currently displayed interface is input into the signal encoder to obtain the signal features, including: S131A and S131B, which are described in detail below.
[0071] S131A inputs the raw data of the monitoring waveform signal in the monitoring signal into the waveform signal encoder to obtain the waveform characteristics; S131B inputs the raw data of the monitoring spectrum signal (monitoring spectrum signal and monitoring envelope spectrum signal) in the monitoring signal into the spectrum signal encoder to obtain the spectral features.
[0072] The waveform signal encoder is composed of a stack of convolutional layers (Conv1D), max pooling layers (MP), and fully connected layers (MLP), and its structure is as follows.
[0073] The input layer of the waveform signal encoder receives the raw data of the monitored waveform signal. The data length can be, but is not limited to, 65536.
[0074] The first convolutional layer (Conv1D_1) of the waveform signal encoder has 8 convolutional kernels, a kernel size of 64, and a stride of 8. The first convolutional layer uses a large convolutional kernel (64) and a stride (8) to capture low-frequency, macroscopic features in long sequences and perform significant dimensionality reduction. The length of a single convolution and the corresponding data length is 8192. The output of the first convolutional layer is represented as 8192×8.
[0075] The kernel size of the first maximum pooling layer (MP_1) of the waveform signal encoder is 8, which further reduces the length of the sequence output by the first convolutional layer to 1 / 8 of the original, that is, 1024×8.
[0076] The second convolutional layer (Conv1D_2) of the waveform signal encoder has 16 convolutional kernels, a kernel size of 32, and a stride of 4. The second convolutional layer reduces the kernel size and increases the number of channels (i.e., the number of convolutional kernels) to extract more complex features from the shorter sequence output by the first max pooling layer. The output of the second convolutional layer is represented as 256×16.
[0077] The second maximum pooling layer (MP_2) of the waveform signal encoder has a convolution kernel size of 8, which further reduces the length of the sequence output by the second convolution layer to 1 / 8 of the original, i.e., 32×16.
[0078] The third convolutional layer (Conv1D_2) of the waveform signal encoder has 32 convolutional kernels with a kernel size of 16 and a stride of 2. The third convolutional layer continues to extract higher-level abstract features, which are used to extract more complex features from the short sequence output by the second max pooling layer. The output of the third convolutional layer is represented as 16×32.
[0079] The third maximum pooling layer (MP_3) of the waveform signal encoder has a convolution kernel size of 4, which further reduces the sequence length output by the third convolution layer to 1 / 4 of the original length, shortening the sequence length to only 4 time steps, i.e., 4×32.
[0080] Flatten (implicit) of waveform signal encoder: Before entering the fully connected layer, the output sequence of the third max pooling layer is flattened from 4*32 to a one-dimensional vector of 4*32=128 neurons.
[0081] The fully connected layer of the waveform signal encoder maps the 128-dimensional feature map language model (which can use embedding) extracted from the convolutional block into a high-dimensional space, which can use 3584-dimensional feature vectors.
[0082] The spectral signal encoder comprises, in sequence, an input layer, a first convolutional layer (Conv1D_1), a first max pooling layer (MP_1), a second convolutional layer (Conv1D_2), a second max pooling layer (MP_2), a third convolutional layer (Conv1D_3), a third max pooling layer (MP_3), a fourth convolutional layer (Conv1D_4), a global average pooling layer (GAP1D), a sampling frequency information addition layer (Metadata Input), a concatenation layer, and a fully connected layer (MLP).
[0083] The input layer of the spectrum encoder receives the raw data of the monitored spectrum signal, with a data length of L. Considering that the spectrum and envelope spectrum data are currently calculated using Global Fourier Transform (FFT), they can be regarded as a one-dimensional frequency sequence, and thus can be processed using Conv1D.
[0084] The first convolutional layer (Conv1D_1) of the spectral signal encoder has 8 convolutional kernels, a kernel size of 64, and a stride of 8. The output of the first convolutional layer is represented as L1×8.
[0085] The kernel size of the first max pooling layer (MP_1) of the spectral signal encoder is 4, and the output of the first max pooling layer is represented as L2×8.
[0086] The second convolutional layer (Conv1D_2) of the spectral signal encoder has 16 convolutional kernels, a kernel size of 32, and a stride of 4. The output of the second convolutional layer is represented as L3×16.
[0087] The kernel size of the second max pooling layer (MP_2) of the spectral signal encoder is 4, and the output of the second max pooling layer is represented as L4×16.
[0088] The third convolutional layer (Conv1D_3) of the spectral signal encoder has 32 convolutional kernels, a kernel size of 16, and a stride of 2. The output of the third convolutional layer is represented as L5×32.
[0089] The kernel size of the third max pooling layer (MP_3) of the spectral signal encoder is 2, and the output of the third max pooling layer is represented as L6×32.
[0090] The fourth convolutional layer (Conv1D_4) of the spectral signal encoder has 64 convolutional kernels, a kernel size of 8, and a stride of 2. The output of the third convolutional layer is represented as L7×64.
[0091] The output of the Global Average Pooling (GAP1D) layer is represented in 64 dimensions.
[0092] The output of the sampling frequency information addition layer (Metadata Input) corresponds to the sampling frequency and can be, but is not limited to, 2-dimensional.
[0093] The concatenation layer is used to concatenate the output of the sampling frequency information addition layer with the output of the global average pooling layer, and can be, but is not limited to, 66 dimensions.
[0094] The fully connected layer (MLP) maps the feature output of the concatenation layer to the high-dimensional space of the language model (which can use embedding), and this high-dimensional space can use 3584-dimensional feature vectors.
[0095] Where L represents the varying length of the input spectrum. L1 to L7 are the intermediate feature map lengths that vary with L.
[0096] Because the data lengths collected by different monitoring devices vary, a Global Average Pooling (GAP1D) layer is introduced into the spectral signal encoder to process variable-length data into intermediate dimensions of the same length. Considering that sensors can have multiple sampling frequencies (which can be, but are not limited to, two), a sampling frequency information addition layer is introduced into the spectral signal encoder. This layer uses the one-hot encoding corresponding to the sampling frequency as prior information. Then, the 64-dimensional spectral features output by the concatenation layer and the GAP1D layer are merged and input into the fully connected layer (MLP). This guides the MLP to learn how to weightedly combine the 64-dimensional spectral features output by GAP1D according to the current sampling rate, thereby improving the model's discriminative and generalization abilities. Finally, the dimension of the fully connected layer (MLP) is set to 3584, the same dimension used in the subsequent text large-scale model embedding.
[0097] Visual Encoder (Ev): When performing data analysis and fault diagnosis, diagnostic engineers can intuitively see the visual representation of the raw data. In order to simulate the visual feature extraction capabilities of diagnostic engineers, in addition to using a signal encoder to extract features from the raw signal, the four types of data—indicator trend, waveform, spectrum, and envelope spectrum—are directly converted into image form. High-dimensional feature extraction is performed through the ViT model, and the output dimension is also set to 3584 dimensions, which is used in the subsequent large text model Embedding.
[0098] The text encoder (Et) is used to inject device modeling information, current alarm information, and currently processed data, allowing the model to understand the content to be analyzed from more dimensions. Device modeling information mainly consists of text-based descriptions of the device structure and the connection relationships between device components; current alarm information mainly consists of text-based descriptions of the alarm content, such as the alarm level generated by a certain indicator; currently processed data information mainly includes data types (indicator trends, waveforms, spectra, envelope spectra) and their time information. This information is necessary because although a unified structure is used to train the model, not all data needs to be analyzed completely in every analysis. Instead, targeted analysis is performed based on the data currently being analyzed, which aligns with the step-by-step analysis habits in human diagnostics. Therefore, the data type and time information currently being analyzed need to be explicitly passed to the model for training and inference. The encoder already present in a pre-trained large text model can be reused here.
[0099] Building upon the preceding discussion of how to train a descriptive inference model to accurately reproduce the descriptive annotation behavior of expert diagnostic engineers, this invention also provides an optional implementation method, please refer to... Figure 4 , Figure 4This is a schematic diagram illustrating the training process of the derivation model provided in an embodiment of the present invention. The training process of the derivation model includes steps S31, S32, S33, and S34, which are described in detail below.
[0100] S31. Obtain the original monitoring signal data, screenshot of the display interface, and related text information corresponding to the interface displayed when the diagnostic engineer performs fault diagnosis, and combine them into the second training sample.
[0101] S32, use the description of the engineer's annotation interface displayed when the diagnostic engineer performs fault diagnosis as the label of the corresponding second training sample.
[0102] S33, arrange the labeled second training samples in chronological order to obtain the second training sequence.
[0103] S34, the descriptive inference model is trained using the second training sequence.
[0104] This allows the trained model to automatically generate textual descriptions of signal representations based on multimodal inputs.
[0105] Since the derivation model supports data inputs of different modalities, and different data formats require feature extraction through different encoders, a unified format and identifier are designed to distinguish the input data:<wave_signal> 'and'< / wave_signal> 'Used to frame waveform signal data,'<spectrum_signal> 'and'< / spectrum_signal> 'Used to frame spectral or envelope spectral data,' 'and' are used to frame image data. <text>'And'< / text> This is used to frame text data. During training, the data is sent to the corresponding encoder for processing based on the identifier. During training, it is not required that all four data types exist simultaneously in each training sample; data that does not exist can be left blank. The output is a signal characteristic description written by a diagnostic engineer, presented in natural language, such as "Periodic impacts are visible in the low-frequency acceleration waveform at 10:30:00 on 2025-9-24, with the impact frequency being the rotational frequency of the gearbox's first shaft."
[0106] Optionally, the loss function describing the derived model is:
[0107] in, Represents the loss function. This indicates the time step length of the second training sequence. This represents the label of the second training sample at time step t. This represents the label of the second training sample before time step t (all samples). This represents the multi-source signal features obtained by fusing the signal features, image features, and text features at time step t. Indicating the characteristics of multi-source signals and Under the condition, the model output and The consistent probability, the model training objective is to minimize During the training of the derivation model, the text encoder (which is already quite mature) is frozen, and the signal encoder and visual encoder are adjusted during optimization.
[0108] Please see Figure 5 , Figure 5 The present invention provides a fault diagnosis device for industrial equipment, which may optionally be applied to the electronic equipment described above.
[0109] The fault diagnosis device for industrial equipment includes: a first processing unit 501 and a second processing unit 502.
[0110] The first processing unit 501 is used to input the current display interface and its interface description into the operation prediction model to perform operation prediction in order to obtain the current operation prediction information. The current display interface is an interface used to display the monitoring signals of industrial equipment. The interface description is a summary description of the monitoring signals displayed in the current display interface and related to fault analysis. The current operation prediction information includes the current operation, its corresponding operation data type, and the operation execution position.
[0111] The first processing unit 501 is also used to obtain a new current display interface based on the current prediction information.
[0112] The first processing unit 501 is also used to input the new current display interface into the multimodal description derivation model to extract and fuse features of the current display interface to obtain the interface description of the current display interface, and repeatedly use the current display interface and its interface description as input to the operation prediction model to perform operation prediction to obtain the current operation prediction information.
[0113] The second processing unit 502 is used to take the target sequence as input to the fault reasoning model after the observation phase ends in order to obtain the fault diagnosis result. The target sequence includes the operation data type and interface description corresponding to each display interface.
[0114] It should be noted that the industrial equipment fault diagnosis device provided in this embodiment can execute the method flow shown in the above-described method flow embodiment to achieve the corresponding technical effects. For the sake of brevity, any parts not mentioned in this embodiment can be referred to the corresponding content in the above-described embodiments.
[0115] This invention also provides a storage medium storing computer instructions and programs, which, when read and executed, perform the fault diagnosis method for industrial equipment described in the above embodiments. The storage medium may include memory, flash memory, registers, or a combination thereof.
[0116] The following provides an electronic device, which may be a server device, a computer device, or a mobile phone device, etc. This electronic device, for example... Figure 1 As shown, the above-described fault diagnosis method for industrial equipment can be implemented. Specifically, the electronic device includes: a processor 10, a memory 11, and a bus 12. The processor 10 may be a CPU. The memory 11 is used to store one or more programs, which, when executed by the processor 10, perform the fault diagnosis method for industrial equipment described in the above embodiment.
[0117] In summary, the fault diagnosis method and related apparatus for industrial equipment provided by this invention involves inputting the current display interface and its description into an operation prediction model to perform operation prediction and obtain current operation prediction information. The current display interface is used to display monitoring signals of the industrial equipment, and the interface description is a summary description of the monitoring signals displayed on the current display interface and related to fault analysis. The current operation prediction information includes the currently executed operation, its corresponding operation data type, and the operation execution position. A new current display interface is obtained based on the current prediction information. The new current display interface is input into a multimodal description derivation model to extract and fuse features of the current display interface to obtain an interface description. The current display interface and its description are repeatedly used as input to the operation prediction model to perform operation prediction and obtain current operation prediction information. After the observation phase, a target sequence is used as input to a fault inference model to obtain the fault diagnosis result. The target sequence includes the operation data type and interface description corresponding to each display interface. By predicting operations using an operational prediction model, the system reproduces the expert's click path and generates natural language feature descriptions using a multimodal model. This makes the diagnostic process transparent and causal, allowing users to intuitively understand the source of the diagnostic results and its supporting evidence. The system realizes the entire process from multi-round interactive operation prediction and feature analysis description to final diagnostic reasoning, and has the ability to reason about faults and generate conclusions.
[0118] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
[0119] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. A failure diagnosis method of an industrial device, characterized by, The method comprises: inputting a current display interface and an interface description thereof into an operation prediction model to perform operation prediction to obtain current operation prediction information, wherein the current display interface is an interface for displaying monitoring signals of an industrial device, the interface description is a summary description of the monitoring signals displayed in the current display interface in relation to fault analysis, and the current operation prediction information comprises a current expansion operation and operation data types and operation execution positions corresponding to the current expansion operation; obtaining a new current display interface according to the current prediction information; inputting the new current display interface into a multi-modal description derivation model to perform feature extraction and fusion on the current display interface to obtain an interface description of the current display interface, and repeatedly inputting the current display interface and the interface description thereof into the operation prediction model to perform operation prediction to obtain current operation prediction information; after the observation stage ends, inputting a target sequence into a fault reasoning model to obtain a fault diagnosis result, wherein the target sequence comprises operation data types and interface descriptions corresponding to each display interface.
2. The failure diagnosis method of an industrial device according to Claim 1, characterized by, The training process of the operation prediction model comprises: obtaining six-tuple information corresponding to each display interface when a diagnostic engineer performs fault diagnosis, wherein the six-tuple information comprises a screenshot of a display interface, an engineer-annotated interface description of the display interface, an expansion operation performed by the engineer on the display interface, operation data types, operation execution positions, and an operation purpose annotated by the engineer; arranging the obtained six-tuple information in chronological order to obtain a first training sequence; completing the training of the operation prediction model based on the first training sequence.
3. The failure diagnosis method of an industrial device according to Claim 1, characterized by, The operation prediction model combines historical display interfaces of the current batch of monitoring signals, interface descriptions of the historical display interfaces, the current display interface, and the interface description of the current display interface to perform operation prediction to obtain current operation prediction information. The description derivation model comprises a signal encoder, a visual encoder, a text encoder, and a language model, and the inputting of the new current display interface into the multi-modal description derivation model to perform feature extraction and fusion on the current display interface to obtain an interface description of the current display interface comprises:
4. The failure diagnosis method of an industrial device according to Claim 1, characterized by, inputting original data of monitoring signals corresponding to the current display interface into the signal encoder to obtain signal features; inputting a screenshot of the current display interface into the visual encoder to obtain image features; inputting relevant text information of an observed device in the current display interface into the text encoder to obtain text features; mapping the signal features, the image features, and the text features to a high-dimensional space of the language model to perform cross-modal feature alignment to obtain multi-source signal features; inputting the multi-source signal features into the language model to perform understanding and representation to obtain the interface description of the current display interface. 5. The failure diagnostic method of an industrial device according to Claim 4, characterized by, The signal encoder comprises a waveform signal encoder and a spectrum signal encoder, the signal features comprise waveform features and spectrum features, and the monitoring signal original data corresponding to the current display interface is input into the signal encoder to obtain signal features, including: monitoring waveform signal original data in the monitoring signal is input into the waveform signal encoder to obtain waveform features; monitoring spectrum signal original data in the monitoring signal is input into the spectrum signal encoder to obtain spectrum features.
6. The failure diagnostic method of an industrial device according to Claim 4, characterized by, The training process of the description derivation model is: monitoring signal original data corresponding to the display interface, display interface screenshots, and related text information when the diagnostic engineer performs fault diagnosis are obtained, which are combined as second training samples; the engineer's comment interface description of the display interface when the diagnostic engineer performs fault diagnosis is taken as a label of the corresponding second training sample; the second training samples with labels are arranged in chronological order to obtain a second training sequence; the description derivation model is trained by using the second training sequence.
7. The failure diagnosis method of an industrial device according to Claim 6, characterized by, The loss function of the description derivation model is: wherein, represents a loss function, represents a time step length of the second training sequence, represents a label of the second training sample at the t-th time step, represents a label of the second training sample before the t-th time step, represents a multi-source signal feature after fusing the signal feature, the image feature and the text feature at the t-th time step, represents a multi-source signal feature after fusing the signal feature, the image feature and the text feature at the t-th time step, and the probability that the model output is consistent with under the condition.
8. A failure diagnosing apparatus of an industrial device, characterized by comprising: The apparatus comprises: a first processing unit configured to input a current display interface and an interface description thereof into an operation prediction model to perform operation prediction, so as to obtain current operation prediction information, wherein the current display interface is an interface for displaying monitoring signals of an industrial device, the interface description is a summary description of the monitoring signals displayed in the current display interface and related to fault analysis, and the current operation prediction information comprises a current expansion operation, an operation data type corresponding to the current expansion operation, and an operation execution position; the first processing unit is further configured to obtain a new current display interface according to the current prediction information; the first processing unit is further configured to input the new current display interface into a multi-modal description derivation model to perform feature extraction and fusion on the current display interface, so as to obtain an interface description of the current display interface, and repeatedly input the current display interface and the interface description thereof into the operation prediction model as input to perform operation prediction, so as to obtain current operation prediction information; a second processing unit configured to input a target sequence into a fault reasoning model as input after an observation stage ends, so as to obtain a fault diagnosis result, wherein the target sequence comprises an operation data type and an interface description corresponding to each display interface.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by a processor to implement the method of any one of claims 1-7.
10. An electronic device, comprising: comprise: a processor and a memory, the memory being configured to store one or more programs; when the one or more programs are executed by the processor, the method of any one of claims 1-7 is implemented.