Equipment remote fault diagnosis method based on large model

By fusing multimodal data into a large model and using an attention mechanism for fault reasoning, combined with human-machine collaborative verification, the problems of geographical and knowledge dependence in traditional equipment diagnosis are solved, and efficient and accurate remote fault diagnosis and support operations are achieved.

CN121765236APending Publication Date: 2026-03-31CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719
View PDF 0 Cites 2 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-06
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Traditional equipment fault diagnosis methods rely on on-site expert diagnosis, which suffers from geographical limitations, knowledge dependence, and insufficient data utilization, resulting in low efficiency and high costs, and failing to achieve systematic and sustainable equipment support.

Method used

A remote fault diagnosis method for equipment based on a large model is adopted. By acquiring real-time sensor data, historical text data and on-site multimedia data, and aligning and fusing them based on a unified time reference, a multimodal feature vector is generated. Then, a multimodal reasoning model with attention mechanism is used for fault reasoning. Combined with a human-machine collaborative verification mechanism, the final diagnosis result is generated, and resource scheduling instructions are automatically generated.

Benefits of technology

It enables automatic reasoning of fault causes and handling suggestions from multi-source heterogeneous data, improves the objectivity and automation level of diagnosis, ensures the accuracy and reliability of diagnostic results, realizes rapid linkage from problem identification to resource scheduling, and improves the efficiency and reliability of remote support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121765236A_ABST
    Figure CN121765236A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of equipment fault diagnosis, and particularly discloses an equipment remote fault diagnosis method based on a large model. Through multi-modal data fusion, an attention mechanism reasoning model, man-machine cooperation verification and intelligent resource scheduling, the problems that maintenance excessively depends on expert experience and a mature remote scheme is lacked are solved. Specifically, according to the scheme, firstly, multi-source data are aligned and fused to form a unified feature vector, and the defect of information isolation is overcome. Afterwards, a model based on an attention mechanism can automatically focus key features, preliminary diagnosis is generated, and dependence on expert experience is reduced. And then, through a man-machine interaction verification mechanism, an expert can remotely check and correct a result, so that the diagnosis reliability is ensured. And finally, the system automatically schedules resources according to a verified result, so that rapid linkage from diagnosis to disposal is realized, remote guarantee is accurately executed, and equipment deterioration is effectively restrained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of equipment fault diagnosis, and more specifically, relates to a method for remote equipment fault diagnosis based on a large model. Background Technology

[0002] Currently, with the rapid expansion of modern equipment systems in terms of complexity and geographical distribution, traditional fault diagnosis and maintenance methods are increasingly revealing their systemic limitations. This model heavily relies on professionals going to the site for diagnosis and operation, and its inherent defects are amplified dramatically in the context of the new era.

[0003] First, severe geographical limitations make it difficult to balance the support effectiveness of equipment. When critical equipment is deployed in remote, high-risk, or overseas areas, organizing technical experts for on-site support is not only slow to respond and costly in terms of logistics, but may also delay the best maintenance window due to harsh environments, directly threatening the continuity of the mission and the availability of the equipment.

[0004] Secondly, this model relies heavily on the experience and knowledge of individual experts, creating a bottleneck in knowledge transfer. The accuracy and efficiency of fault diagnosis often depend on the long-term experience accumulation of individual experts. This tacit knowledge is difficult to standardize, replicate on a large scale, and effectively pass on, resulting in the support capability being constrained by the mobility and unevenness of human resources, and failing to form a systematic and sustainable collective wisdom.

[0005] Finally, traditional methods of utilizing data value are extremely superficial and isolated. Diagnostic processes often rely on limited real-time sensor alarms or fragmented experience-based judgments, failing to deeply mine information throughout the equipment's entire lifecycle. This results in a significant amount of potential value that could be used to predict faults and optimize management being overlooked, and decision-making lacking sufficient data support. These background deficiencies collectively constitute the core driving force for the modernization and intelligent upgrading of traditional support systems. Summary of the Invention

[0006] In view of the shortcomings of the existing technology, the purpose of this application is to provide a remote equipment fault diagnosis method based on a large model, which aims to solve the problem of equipment damage caused by reliance on individual expert experience and lack of mature remote maintenance solutions in modern maintenance processes.

[0007] To achieve the above objectives, in a first aspect, this application provides a method for remote equipment fault diagnosis based on a large model, comprising: acquiring real-time sensor data, historical text data, and on-site multimedia data of the equipment; aligning the real-time sensor data and on-site multimedia data with historical text data based on a unified time reference to generate a multimodal feature vector; inputting the multimodal feature vector into a multimodal inference model based on an attention mechanism for processing to obtain an initial diagnostic result containing fault causes and preliminary handling suggestions; responding to a verification request for the initial diagnostic result, providing an interactive interface to receive verification instructions, and generating a final diagnostic result based on the verification instructions; and generating resource scheduling instructions based on the final diagnostic result to execute remote support operations.

[0008] In one embodiment, real-time sensor data, historical text data, and on-site multimedia data are aligned and fused based on a unified time reference to generate a multimodal feature vector. This includes: assigning a unified timestamp to the real-time sensor data and on-site multimedia data and performing spatiotemporal alignment; extracting the temporal features of the aligned sensor data, the visual features of the multimedia data, and the semantic features of the text data, respectively; and fusing the temporal features, semantic features, and visual features to generate a multimodal feature vector.

[0009] In one embodiment, the multimodal feature vector is input into a multimodal inference model based on an attention mechanism for processing, including: using the attention mechanism to calculate the association weights between temporal features, semantic features and visual features; performing weighted fusion of temporal features, semantic features and visual features according to the association weights to obtain context-aware fusion features; and performing fault inference through the multimodal inference model based on the context-aware fusion features to output an initial diagnosis result.

[0010] In one embodiment, the step of providing an interactive interface to receive verification instructions in response to a verification request for an initial diagnostic result includes: establishing a data channel with a remote user terminal; displaying the initial diagnostic result and the reasoning basis of the multimodal reasoning model on the interactive interface; receiving feedback data from the remote user terminal through the data channel; and integrating the feedback data with the initial diagnostic result to generate a final diagnostic result.

[0011] In one embodiment, the feedback data includes confirmation, correction, or supplementary information on the initial diagnostic result; the step of integrating the feedback data with the initial diagnostic result includes: when the feedback data is correction or supplementary information, optimizing the relevant parameters of the multimodal inference model using the correction or supplementary information, and regenerating the diagnostic result as the final diagnostic result.

[0012] In one embodiment, after the step of automatically generating resource scheduling instructions based on the final diagnosis results, the method further includes: recording the final diagnosis results, resource scheduling instructions, and execution results of this fault as new fault cases; and adding the new fault cases to the fault case knowledge base after characterization processing to achieve self-updating of the fault case knowledge base.

[0013] In one embodiment, before acquiring real-time sensor data, historical text data, and on-site multimedia data of the equipment, the method further includes: retrieving historical failure cases similar to the current equipment state from a failure case knowledge base; and inputting the feature information of the retrieved historical failure cases as prior knowledge into a multimodal reasoning model to assist in failure reasoning.

[0014] Secondly, this application provides a remote fault diagnosis device for equipment, comprising: a multimodal data fusion module for acquiring real-time sensor data, historical text data, and on-site multimedia data of the equipment, and aligning and fusing the real-time sensor data, historical text data, and on-site multimedia data based on a unified time reference to generate a multimodal feature vector; an intelligent reasoning module for inputting the multimodal feature vector into a multimodal reasoning model based on an attention mechanism for processing to obtain an initial diagnosis result containing the cause of the fault and preliminary handling suggestions; an interactive verification module for responding to a verification request for the initial diagnosis result, providing an interactive interface to receive verification instructions, and generating a final diagnosis result based on the verification instructions; and a support decision module for generating resource scheduling instructions based on the final diagnosis result to execute remote support operations.

[0015] Thirdly, this application provides an electronic device including a memory and one or more processors; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code including computer instructions; the one or more processors invoke the computer instructions to cause the electronic device to perform the method of the first aspect.

[0016] Fourthly, this application provides a computer-readable storage medium including instructions that, when executed on an electronic device, cause the electronic device to perform the method of the first aspect.

[0017] Overall, the technical solutions conceived in this application have the following beneficial effects compared with the prior art: This application, by employing a technology that aligns and fuses real-time sensor data, historical text, and on-site multimedia data based on a unified time benchmark, first addresses the shortcomings of isolated information and single-dimensionality in traditional diagnostics. It integrates fragmented information into a unified multimodal feature vector, laying a comprehensive and accurate data foundation for subsequent in-depth analysis. Furthermore, by inputting these multimodal feature vectors into an attention-based inference model, which simulates expert thinking and autonomously focuses on the key features most relevant to the fault, it solves the problem of strong dependence on expert on-site experience. This enables the automatic inference of initial diagnostic results containing fault causes and handling suggestions from multi-source heterogeneous data, improving the objectivity and automation level of the diagnosis. Subsequently, by introducing a mechanism for responding to verification requests and providing an interactive interface to receive verification instructions, this solution constructs a human-machine collaborative decision-making closed loop. By granting experts remote review and correction permissions, it addresses the reliability risks that may exist in purely automated diagnosis, ensuring the accuracy and trustworthiness of the final diagnostic results. Finally, resource scheduling instructions are automatically generated based on the verified diagnostic results.

[0018] Compared with existing technologies, this solution seamlessly integrates intelligent diagnosis with support execution, solving the problem of disconnect between diagnosis and action in the traditional model. It enables rapid linkage from problem identification to resource scheduling, thereby accurately and efficiently executing remote support operations and effectively curbing the deterioration of equipment damage. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating the remote equipment fault diagnosis method based on a large model provided in an embodiment of this application. Figure 2 This is a schematic diagram of the process for generating multimodal feature vectors provided in an embodiment of this application; Figure 3 This is a schematic diagram illustrating the process of the multimodal reasoning model based on the attention mechanism provided in the embodiments of this application. Figure 4 This is a schematic diagram of the process for generating the final diagnostic result provided in an embodiment of this application; Figure 5 This is a schematic diagram of the self-updating process of the fault case knowledge base provided in the embodiments of this application; Figure 6 This is a schematic diagram of the prior knowledge processing provided in the embodiments of this application; Figure 7 This is a schematic diagram of the structure of the remote fault diagnosis device for equipment provided in the embodiments of this application; Figure 8 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0021] In this article, the term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The symbol " / " in this article indicates that the related objects are in an "or" relationship; for example, A / B means A or B.

[0022] The terms "first" and "second," etc., used in the specification and claims herein are used to distinguish different objects, not to describe a specific order of objects. For example, "first response message" and "second response message," etc., are used to distinguish different response messages, not to describe a specific order of response messages.

[0023] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0024] In the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more, for example, multiple processing units means two or more processing units, multiple elements means two or more elements, etc.

[0025] Existing remote fault diagnosis methods for equipment have significant drawbacks and limitations due to geographical restrictions and reliance on specialized knowledge. Furthermore, as equipment systems become increasingly complex and geographically dispersed, current methods primarily rely on local technical personnel or expert teams for on-site diagnosis. Therefore, equipment diagnosis is not only inefficient and costly in terms of manpower, but also impacts equipment operation and various missions.

[0026] Based on this, this application provides an embodiment of a remote equipment fault diagnosis method based on a large model. Please refer to... Figure 1 , Figure 1 This is a flowchart illustrating the remote equipment fault diagnosis method based on a large model provided in this application embodiment. In this embodiment, the remote equipment fault diagnosis method based on a large model includes steps S10 to S40.

[0027] Step S10: Acquire real-time sensor data, historical text data, and on-site multimedia data of the equipment, and align the real-time sensor data and on-site multimedia data with historical text data based on a unified time reference to generate a multimodal feature vector.

[0028] Understandably, the ultimate goal of step S10 is to generate a multimodal feature vector. This vector is a standardized data representation that can be understood and processed by computers, especially large models. It is a highly condensed dataset containing all the key information about the past (historical text), the present (real-time sensors), and the situation on-site (multimedia).

[0029] It should be noted that real-time sensor data is a numerical data stream collected in real time from various sensors installed on the equipment, such as temperature, pressure, vibration, voltage, and current sensors. It reflects the most accurate physical operating state of the equipment at present. For example, real-time engine speed, oil temperature, and vibration amplitude. This is the core objective basis for determining whether a fault has occurred and its nature.

[0030] It should be noted that historical text data is recorded in text form and contains historical information related to the equipment. This includes: past fault records, maintenance measures, replaced parts, standard operating procedures, maintenance specifications, equipment technical parameters, system schematics, and successfully resolved similar fault diagnosis cases. It provides background knowledge and historical experience for diagnosis. The large model can learn from this data "under what circumstances is this type of equipment prone to what problems?" and "how were similar symptoms resolved last time?" This simulates the experience of human experts.

[0031] It should be noted that on-site multimedia data consists of images, videos, and audio data captured by on-site cameras, microphones, or mobile phones used by inspection personnel. This provides intuitive on-site situational information, compensating for the limitations of sensor and text data. For example, images / videos can reveal visible anomalies such as oil leaks, smoke, loose or damaged parts; audio can capture unusual noises, such as abnormal friction or impact sounds. This is crucial for diagnosing faults that sensors cannot directly detect, such as cosmetic damage.

[0032] Understandably, aligning real-time sensor data and on-site multimedia data based on a unified time reference means assigning a unified, high-precision time stamp to all data—real-time data streams and multimedia data—to ensure that information can be correctly correlated. This involves matching data from different sources that belong to the same point in time or the same time period.

[0033] Understandably, this fusion is not a simple patchwork, but rather the integration of aligned multimodal information into a coherent data representation with inherent logic. For example, it combines information such as the frequency characteristics of vibration signals, the location of components identified in images, and common failure modes of those components from historical logs.

[0034] Specifically, the fusion of historical text data does not refer to precise point-in-time matching with traditional time-series data, which is neither possible nor necessary. Its core implementation method is a semantic context-based association and activation. This process can be understood as follows: the system first performs deep analysis on real-time collected sensor data and on-site multimedia data, extracting key feature patterns, such as "abnormal engine vibration frequency" and "loose components identified in on-site images." These features together constitute a real-time context vector describing the current fault situation. Simultaneously, in the system backend, all historical textual materials, including maintenance records, technical manuals, and case studies, are stored in a dedicated vector database. The system then performs rapid semantic similarity matching on the real-time context vector representing the current situation, searching for historical experiences that are closest in deeper meaning to the current fault phenomenon.

[0035] Understandably, data from each modality is processed to extract the numerical features that best represent its key information. Sensor data can have its mean, variance, and spectral features extracted. Text data can be converted into word vectors or semantic vectors using natural language processing techniques. Multimedia data can have its image and audio features extracted using computer vision and audio processing techniques. These features extracted from different modalities are then combined to generate a multimodal feature vector. This vector can describe the overall state of the equipment at a specific moment.

[0036] In one implementation, please refer to Figure 2 , Figure 2 This is a schematic diagram of the process for generating multimodal feature vectors provided in an embodiment of this application. Step S10 includes steps S11 to S13.

[0037] Step S11: Add a unified timestamp to the real-time sensor data and the on-site multimedia data, and perform spatiotemporal alignment.

[0038] Understandably, this step primarily addresses the issue of data correspondence. The system first assigns a precise timestamp from the same clock source to all real-time sensor readings and on-site photos and videos. This ensures that "abnormal vibration data recorded at 2:05:30 PM" and "equipment photo taken at 2:05:31 PM" originate from the same scene, laying the foundation for subsequent analysis by linking them together. It's like compiling information from different sources into a single record sheet marked with a unified time point.

[0039] Specifically, data can be transmitted to a cloud platform in real time via wireless sensor networks or wired connections. Video and audio streams are transmitted via RTSP or HTTP protocols and undergo image processing such as object detection and anomaly recognition to aid in fault diagnosis. All of this data is aligned with a unified timestamp to ensure synchronization of multi-source data over time.

[0040] Step S12: Extract the temporal features of the aligned sensor data, the visual features of the multimedia data, and the semantic features of the text data, respectively.

[0041] Understandably, after data time alignment, the system begins extracting key information from the three types of data. For sensor data, it analyzes the patterns of change over time, such as whether the values ​​spike suddenly or decline slowly, and whether there are specific fluctuation cycles; these are called time-series features. For on-site photos and videos, the system uses image recognition technology to identify key visual information, such as whether there is an oil leak, whether screws are loose, and the position of dashboard pointers; these are summarized as visual features. For historical text, the system uses natural language processing technology to understand the meaning of the text; for example, it understands the problem and cause described in the sentence "the filter is clogged, causing excessive oil pressure"; this is called semantic features.

[0042] Specifically, sensor data can be analyzed using a long short-term memory network model to extract potential fault trends during equipment operation; text data can be semantically understood using pre-trained natural language processing models such as BERT and BioClinicalBERT to identify fault modes; and video data can be extracted for image features using convolutional neural networks to further help identify the physical fault manifestations of the equipment.

[0043] Step S13: The temporal features, semantic features and visual features are fused to generate a multimodal feature vector.

[0044] Understandably, the system integrates the three key pieces of information extracted in the previous step—the sensor's temporal patterns, the visual conditions at the scene, and the descriptive meaning of historical text—into a comprehensive data package, namely a multimodal feature vector. This data package is no longer an isolated data point, but a structured collection of information containing the device's current state, its appearance at the scene, and relevant historical knowledge. It will be fed into the subsequent large model as a complete basis for fault analysis and judgment.

[0045] Step S20: Input the multimodal feature vector into the multimodal inference model based on the attention mechanism for processing to obtain the initial diagnostic results containing the cause of the fault and preliminary handling suggestions.

[0046] It should be noted that step S20 is the core reasoning step in the entire large-model-based remote equipment fault diagnosis process. Its core function is to transform the previously fused multimodal feature vectors into structured diagnostic conclusions. Specifically, this step receives input from step S10, which consists of multimodal feature vectors such as sensor time-series data, historical text records, and on-site images / videos aligned to a unified time reference. These data have been transformed into high-dimensional semantic representations through a feature extraction network.

[0047] Understandably, attention-based multimodal inference models automatically learn the contribution of different modal features to fault diagnosis through dynamic weight allocation. For example, vibration sensor data is given higher weight to mechanical wear faults, while image features are emphasized for appearance damage faults. Simultaneously, cross-modal attention is used to capture implicit correlations between modalities, such as associating the description of "bearing overheating" in the text with high-temperature signals detected by sensors and discoloration of the bearing area in images. The model can employ possible architectures including dual-tower attention networks, Transformer fusion models, or graph neural networks, using deep learning techniques to model complex fault modes and ultimately outputting an initial diagnostic result comprising two parts.

[0048] It should be noted that: firstly, the classification and fine-grained description of the fault causes, such as "gearbox seal failure leading to lubricating oil contamination and causing bearing wear," are accompanied by a 92% confidence level; secondly, preliminary handling suggestions based on historical cases and knowledge reasoning, such as "immediately stop the machine, replace the seals, and clean the gearbox." The "preliminary" nature of this result is reflected in its need for further confirmation through subsequent human-computer interaction verification. On the one hand, the model may be affected by data noise or rare fault modes, leading to misjudgments; on the other hand, complex faults require correction based on expert experience.

[0049] Understandably, rapidly generating candidate conclusions significantly reduces manual analysis time and provides a focused direction for the verification process, requiring only the verification of high-confidence results. Compared to traditional methods relying on single-modality or rule matching, this approach reduces diagnostic blind spots through multimodal complementary information, utilizes deep learning to model complex relationships, and enhances interpretability through attention weight visualization. Ultimately, it forms a standardized output from data perception to decision execution, laying a reliable foundation for subsequent resource scheduling and remote support operations.

[0050] In one implementation, please refer to Figure 3 , Figure 3 This is a schematic flowchart illustrating the processing of a multimodal inference model based on an attention mechanism provided in this application embodiment. The process includes steps S21 to S23, where multimodal feature vectors are input into the attention-based multimodal inference model for processing.

[0051] Step S21: Calculate the association weights between temporal features, semantic features, and visual features using an attention mechanism.

[0052] Understandably, the purpose of this step is to dynamically quantify the correlation strength between temporal features, semantic features, and visual features through an attention mechanism, so as to provide a precise data association foundation for subsequent cross-modal reasoning.

[0053] Specifically, this step first applies a cross-modal attention mechanism to the aligned and serialized trimodal features of the input. The feature vector of one modality is used as the query, and the feature vectors of other modalities are used as the key and value. The association weight matrix is ​​generated by calculating dot product similarity and softmax normalization. For example, the weight of the crack region in the temporal segment and the visual image may be as high as 0.8, while the weight of the irrelevant background is close to 0, thus highlighting the fault-related features. In order to capture multi-level semantic associations, a multi-head attention mechanism is further introduced. The features are divided into multiple subspaces, the weights are calculated independently, and then fused to enhance the model's ability to perceive local details and global patterns, thereby generating the association weight matrix.

[0054] Understandably, the final generated correlation weight matrix can suppress noise modal interference through dynamic weight allocation; it can strengthen key evidence chains; and it can support interpretability. Compared with traditional fixed-weight fusion methods, this mechanism enables the model to adapt to modal interaction patterns under different fault scenarios, significantly improving diagnostic accuracy and robustness, and laying a solid data fusion foundation for generating high-confidence initial diagnostic conclusions in subsequent steps.

[0055] Step S22: The temporal features, semantic features and visual features are weighted and fused according to the association weights to obtain context-aware fused features.

[0056] It should be noted that this step first constructs a cross-modal interaction graph based on the attention weight matrix. Taking temporal features as an example, each segment not only needs to integrate contextual information from its own sequence, such as modeling temporal dependencies using LSTM or Transformer, but also needs to extract fault description keywords from semantic features based on association weights, such as semantic vectors for "overheating" and "abnormal noise," and extract key region features from the corresponding time-time device image from visual features, forming a "temporal-semantic-visual" triple association. Subsequently, a weighted summation strategy is used to achieve feature aggregation. For a certain temporal segment Ti, the calculation method for its fused feature Fi is Fi=α. i,s *Sj+α i,v *Vk+λ*Ti, where α i,s and α i,vThe association weights between the segment and semantic feature Sj and visual feature Vk, respectively, are calculated in step S21. λ is the retention coefficient for preserving the original temporal information, which is usually set to 0.3-0.5 to avoid excessive information dilution.

[0057] Understandably, through this dynamic weighting mechanism, the fused features can adaptively highlight fault-related modal information while suppressing irrelevant noise. To further enhance the feature representation capability, residual connections are used to superimpose the original features and the fused features, and a nonlinear transformation is performed through a multilayer perceptron to generate a final context-aware fused feature vector with unified dimensions.

[0058] Step S23: Based on context-aware fusion features, fault reasoning is performed through a multimodal reasoning model to output initial diagnostic results.

[0059] It should be noted that the core task of step S23 is to perform deep analysis and logical reasoning on the context-aware fusion features through a multimodal reasoning model, ultimately outputting a structured initial diagnostic result. The multimodal reasoning model can adopt a hybrid architecture, combining the deterministic logic of a rule engine with the probabilistic reasoning capabilities of deep learning.

[0060] Specifically, in the feature decoding layer, the fused features are nonlinearly transformed using a multilayer perceptron or Transformer decoder to extract high-order semantics related to faults, such as "high-frequency harmonics + tooth surface cracks + insufficient lubrication → tooth surface pitting". In the knowledge graph reasoning layer, a pre-built equipment fault knowledge graph is introduced, which typically contains triples such as fault phenomenon, cause, and solution. Symbolic reasoning is achieved through a graph neural network, for example, matching the causal path of "gear fault → wear / crack" based on "sideband frequency". In the attention-driven decision layer, the model dynamically adjusts the reasoning path according to the weight distribution of each modality in the fused features.

[0061] Specifically, taking CNC machine tool spindle fault diagnosis as an example: First, activate the neurons corresponding to key evidence such as "2nd harmonic vibration + bearing spalling + abnormal noise" in the fusion features. Then, search for similar patterns in the training data and find that this combination corresponds to "spindle bearing cage fracture" in 85% of the cases. Next, use Monte Carlo dropout to quantify the uncertainty (if there is conflicting evidence such as high temperature but sufficient coolant, the confidence level drops from 95% to 78%). Finally, generate a structured report containing fault type (spindle bearing cage fracture), location (spindle front end), severity (moderate, raceway spalling area 15%), confidence level (78%), and recommended measures (replace the bearing and check the lubrication system).

[0062] Understandably, compared to traditional single-modal diagnostic methods, this model can simultaneously capture temporal dynamic changes, visual-spatial impairments, and semantic historical context through cross-modal collaborative reasoning. For example, it can jointly diagnose gearbox lubrication system faults through "temperature fluctuations + oil emulsification + seal failure." Furthermore, its interpretability is enhanced through attention weight visualization and knowledge graph path tracing, enabling experts to quickly verify the rationality of the conclusions. In addition, by combining symbolic reasoning from the knowledge graph with feature learning from deep learning, the model can still generate reliable diagnoses through knowledge transfer even in small-sample scenarios. For instance, it can use a general gear fault knowledge graph to perform preliminary diagnoses of novel gearboxes that have never been seen before.

[0063] Step S30: In response to the verification request for the initial diagnostic result, an interactive interface is provided to receive the verification instruction, and the final diagnostic result is generated based on the verification instruction.

[0064] It should be noted that step S30, as the core link in realizing the human-machine collaborative closed loop in the remote equipment fault diagnosis process, combines the automated advantages of model reasoning with the deep insights of expert experience by constructing an interactive interface and a dynamic verification mechanism, ultimately generating reliable diagnostic conclusions that conform to engineering practice. When the system receives a verification request for the initial diagnostic results, the interactive interface will present key information in a layered visual format.

[0065] In one implementation, please refer to Figure 4 , Figure 4 This is a schematic diagram of the process for generating a final diagnostic result provided in an embodiment of this application. The step of providing an interactive interface to receive verification instructions in response to a verification request for the initial diagnostic result includes steps S31 to S33.

[0066] Step S31: Establish a data channel with the remote user terminal and display the initial diagnostic results and the reasoning basis of the multimodal reasoning model on the interactive interface. Step S32: Receive feedback data from the remote user terminal through the data channel.

[0067] It's important to note that by establishing a secure, real-time data channel, a collaborative cognitive workspace is built for deep interaction with remote domain experts. For example, a high-performance protocol data channel enables a stable, low-latency, two-way real-time connection with the remote user. Once the connection is successful, the system doesn't simply present an isolated diagnostic conclusion, but rather displays it on the interactive interface as a highly integrated diagnostic decision support panel.

[0068] It should be noted that this display can be presented at the top of the panel as structured cards showing the initial conclusions of the model output, such as fault type, location, severity, and confidence level, with key information enhanced through color coding and icon labeling; the middle area integrates a multimodal evidence chain, such as locating abnormal frequency bands in time-series signals through a time axis, annotating crack areas in visual images through heat maps, and highlighting repair records in semantic logs through text boxes, allowing experts to quickly trace the basis of the model's reasoning; the bottom provides a verification instruction input area, supporting single / multiple selection buttons for quick confirmation or correction of conclusions, free text boxes to supplement on-site observation details, such as "there is oil on the gearbox housing, it is recommended to check the lubrication system", and multimodal annotation tools to select minor faults not recognized by the model, such as early cracks in visual images, while dynamically loading historical data, knowledge base entries, and real-time operating conditions for reference.

[0069] Step S33: Integrate the feedback data with the initial diagnostic results to generate the final diagnostic results.

[0070] Understandably, after receiving an instruction, the system identifies the intent type through semantic analysis, and the feedback data includes confirmation, correction, or supplementary information regarding the initial diagnostic result. If the instruction conforms to preset rules, such as prioritizing expert conclusions when the confidence level is <70%, the final diagnosis is generated directly. When the feedback data is correction or supplementary information, complex corrections are involved. For example, if an expert supplements the visual crack but the model does not detect it, the corrected multimodal data is re-input into the inference model. The relevant parameters of the multimodal inference model are optimized using the correction or supplementary information, and the feature weights are updated through incremental learning, such as increasing the priority of the visual modality in crack detection. The diagnostic result is then regenerated as the final diagnostic result.

[0071] Understandably, when expert and model conclusions conflict, the system calculates the confidence difference. If expert experience carries higher weight, such as based on on-site inspection evidence, the expert conclusion is adopted; otherwise, multi-expert consultation or further data collection is triggered. The final structured report output contains four layers of information: initial model conclusions, expert validation opinions, collaborative correction process, and final conclusions.

[0072] Step S40: Generate resource scheduling instructions based on the final diagnosis results to execute remote support operations.

[0073] It should be noted that this step transforms diagnostic conclusions into executable tasks through a three-layer deconstruction model. First, fault characteristics are extracted and associated with standard maintenance plans; second, resource requirements are mapped, and the optimal combination is selected based on the real-time status of the resource database; finally, operational parameters and collaborative instructions are generated. Simultaneously, in terms of resource scheduling strategies, priority-driven, dynamic allocation, timing optimization, and fault-tolerant design are adopted to ensure efficient task execution. Compared to the traditional manual dispatch mode, this step significantly shortens response time, improves resource utilization and collaborative efficiency, and fully records the instruction generation, scheduling, and execution process, supporting the accumulation and optimization of operational data.

[0074] Specifically, the system combines equipment failure modes with feedback from on-site personnel to automatically generate maintenance tasks and allocate spare parts and tools, ensuring timely provision of necessary resources. The system also integrates with inventory management and logistics systems to automatically schedule spare parts delivery, ensuring all resources are available promptly during maintenance. In certain special circumstances, the system can deploy robots or drones to perform simple maintenance operations such as replacing parts and inspecting internal equipment, reducing the workload of on-site personnel and improving maintenance accuracy and safety.

[0075] In this embodiment, by employing a technique that aligns and fuses real-time sensor data, historical text, and on-site multimedia data based on a unified time benchmark, the shortcomings of isolated information and single-dimensionality in traditional diagnosis are first addressed. Fragmented information is integrated into a unified multimodal feature vector, laying a comprehensive and accurate data foundation for subsequent in-depth analysis. Furthermore, by inputting these multimodal feature vectors into an attention-based inference model, which can simulate expert thinking and autonomously focus on the key features most relevant to the fault, the problem of strong reliance on expert on-site experience is solved. This enables the automatic inference of initial diagnostic results containing fault causes and handling suggestions from multi-source heterogeneous data, improving the objectivity and automation level of the diagnosis. Subsequently, by introducing a mechanism for responding to verification requests and providing an interactive interface to receive verification instructions, this solution constructs a human-machine collaborative decision-making closed loop. By granting experts remote review and correction permissions, the reliability risks that may exist in purely automated diagnosis are resolved, ensuring the accuracy and trustworthiness of the final diagnostic results. Finally, resource scheduling instructions are automatically generated based on the verified diagnostic results.

[0076] Compared with existing technologies, this solution seamlessly integrates intelligent diagnosis with support execution, solving the problem of disconnect between diagnosis and action in the traditional model. It enables rapid linkage from problem identification to resource scheduling, thereby accurately and efficiently executing remote support operations and effectively curbing the deterioration of equipment damage.

[0077] Furthermore, based on the previous embodiment, please refer to... Figure 5 and Figure 6 , Figure 5 This is a schematic diagram of the self-updating process of the fault case knowledge base provided in the embodiments of this application; Figure 6 This is a schematic flowchart illustrating the prior knowledge processing provided in an embodiment of this application. This application incorporates historical data and proposes an embodiment to improve processing speed. Figure 5 In this process, steps S40 is followed by steps S41 and S42.

[0078] Step S41: Record the final diagnosis result, resource scheduling instructions and execution results of this fault as a new fault case.

[0079] Specifically, step S41 systematically records the complete closed-loop data of the fault, including the final diagnostic results (such as fault type, location, and severity), the generated resource scheduling instructions (such as spare parts type, tool list, and operating parameters), and the execution results (such as task completion time and post-repair equipment status), as a new fault case. This process is achieved through standardized data templates to ensure the completeness and structure of the case information, providing a high-quality data foundation for subsequent analysis.

[0080] Step S42: After the new fault cases are characterized, they are added to the fault case knowledge base to achieve self-updating of the fault case knowledge base.

[0081] Understandably, newly added fault cases undergo feature processing to transform them into quantifiable feature vectors, and finally, the processed cases are added to the fault case knowledge base. The knowledge base adopts a hierarchical storage architecture, with high-frequency cases placed in a high-speed cache layer to accelerate retrieval, and low-frequency cases archived in a long-term storage layer. Simultaneously, implicit relationships between cases are automatically established through association rule mining, such as the mapping relationship between specific fault modes and optimal resource combinations. This self-updating mechanism allows the knowledge base to dynamically accumulate operational experience, forming a closed loop: when the feature matching degree between a new case and historical cases exceeds a threshold, the system directly calls the solution of similar cases, avoiding redundant analysis. If the matching degree is low, a prior knowledge processing flow can be triggered, generating new knowledge rules through collaborative verification by an expert system and a machine learning model, and updating the knowledge base in reverse.

[0082] exist Figure 6 In this process, steps S01 and S02 are included before step S10.

[0083] Step S01: Retrieve historical failure cases similar to the current equipment status from the failure case knowledge base. Step S02: Input the feature information of the retrieved historical failure cases as prior knowledge into the multimodal reasoning model to assist in failure reasoning.

[0084] It should be noted that the system can retrieve highly relevant historical failure cases from the failure case knowledge base based on the real-time status data of the current equipment using a similarity matching algorithm. The retrieval process employs a hierarchical screening strategy: first, a coarse screening is performed based on explicit tags such as equipment model and component type to exclude irrelevant cases; then, a fine screening is performed using distance metrics in the feature vector space to select the most similar cases.

[0085] Understandably, the retrieved historical case feature information is transformed into structured prior knowledge and input into the multimodal reasoning model. This process involves three layers of knowledge injection: first, aligning the fault characteristics of historical cases with current real-time data to mark potential correlations; second, extracting causal reasoning chains from historical cases as logical constraints for model reasoning; and third, integrating resource scheduling efficiency data from historical cases to optimize current instructions.

[0086] In addition, all data is encrypted during remote communication and data storage to prevent leakage or tampering. Furthermore, the system employs a role-based access control mechanism to manage user permissions, ensuring that only authorized personnel can access specific data and perform specific operations. Each operation is logged, allowing administrators to review all records and ensure compliance, thus preventing malicious activity.

[0087] In this embodiment, by preloading prior knowledge, the multimodal reasoning model can obtain a high-confidence starting point for reasoning in the initial stage, avoiding data exploration from scratch and significantly improving reasoning speed and convergence. This provides more efficient and reliable technical support for the intelligent operation and maintenance of complex equipment.

[0088] The equipment remote fault diagnosis device provided in this application is described below. The equipment remote fault diagnosis device described below corresponds to and can be referred to in conjunction with the equipment remote fault diagnosis method based on a large model described above. Please refer to... Figure 7 , Figure 7 This is a schematic diagram of the structure of the remote fault diagnosis device for equipment provided in the embodiments of this application.

[0089] The multimodal data fusion module is used to acquire real-time sensor data, historical text data, and on-site multimedia data of the equipment, and to align and fuse the real-time sensor data, historical text data, and on-site multimedia data based on a unified time reference to generate multimodal feature vectors.

[0090] The intelligent reasoning module is used to input multimodal feature vectors into a multimodal reasoning model based on an attention mechanism for processing, and to obtain initial diagnostic results that include the cause of the fault and preliminary handling suggestions.

[0091] The interactive verification module is used to respond to verification requests for the initial diagnostic results, provide an interactive interface to receive verification instructions, and generate the final diagnostic results based on the verification instructions.

[0092] The assurance decision module is used to generate resource scheduling instructions based on the final diagnosis results in order to execute remote assurance operations.

[0093] It is understood that the detailed functional implementation of each of the above units / modules can be found in the description in the foregoing method embodiments, and will not be repeated here.

[0094] It should be understood that the above-described device is used to execute the methods in the above embodiments. The implementation principle and technical effect of the corresponding program modules in the device are similar to those described in the above methods. The working process of the device can be referred to the corresponding process in the above methods, and will not be repeated here.

[0095] Based on the methods in the above embodiments, this application provides an electronic device, with reference to... Figure 8 The electronic device may include a processor, a communications interface, a memory, and a communication bus, wherein the processor, communications interface, and memory communicate with each other via the communication bus. The processor can invoke logical instructions stored in the memory to execute the methods described in the above embodiments.

[0096] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0097] Based on the methods in the above embodiments, this application provides a computer-readable storage medium storing a computer program that, when run on a processor, causes the processor to execute the methods in the above embodiments.

[0098] Based on the methods in the above embodiments, this application provides a computer program product that, when run on a processor, causes the processor to execute the methods in the above embodiments.

[0099] It is understood that the processor in the embodiments of this application can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor can be a microprocessor or any conventional processor.

[0100] The method steps in this application embodiment can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can reside in an ASIC.

[0101] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0102] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for remote equipment fault diagnosis based on a large model, characterized in that, include: The system acquires real-time sensor data, historical text data, and on-site multimedia data of the equipment, aligns the real-time sensor data and on-site multimedia data with the historical text data based on a unified time reference, and generates a multimodal feature vector. The multimodal feature vectors are input into a multimodal inference model based on an attention mechanism for processing to obtain an initial diagnostic result containing the cause of the fault and preliminary handling suggestions; In response to a verification request for the initial diagnostic result, an interactive interface is provided to receive verification instructions, and a final diagnostic result is generated based on the verification instructions; Based on the final diagnostic results, resource scheduling instructions are generated to execute remote support operations.

2. The equipment remote fault diagnosis method based on a large model as described in claim 1, characterized in that, Based on a unified time reference, the real-time sensor data and on-site multimedia data are aligned and fused to generate multimodal feature vectors, including: The real-time sensor data and on-site multimedia data are given a unified timestamp and then spatiotemporally aligned. Temporal features of aligned sensor data, visual features of multimedia data, and semantic features of text data are extracted separately. The temporal features, semantic features, and visual features are fused to generate the multimodal feature vector.

3. The equipment remote fault diagnosis method based on a large model as described in claim 2, characterized in that, The multimodal feature vectors are input into an attention-based multimodal inference model for processing, including: The association weights among the temporal features, semantic features, and visual features are calculated using an attention mechanism; The temporal features, semantic features, and visual features are weighted and fused according to the association weights to obtain context-aware fused features; Based on the context-aware fusion features, fault reasoning is performed through the multimodal reasoning model to output the initial diagnostic result.

4. The equipment remote fault diagnosis method based on a large model as described in claim 1, characterized in that, The step of providing an interactive interface to receive verification instructions in response to a verification request for the initial diagnostic results includes: Establish a data channel with the remote user terminal, and display the initial diagnostic results and the reasoning basis of the multimodal reasoning model on the interactive interface; Feedback data from the remote user terminal is received through the data channel; The feedback data is integrated with the initial diagnostic results to generate the final diagnostic result.

5. The equipment remote fault diagnosis method based on a large model as described in claim 4, characterized in that, The feedback data includes confirmation, correction, or supplementary information regarding the initial diagnostic results; The step of integrating feedback data with the initial diagnostic results includes: when the feedback data is corrective or supplementary information, optimizing the relevant parameters of the multimodal inference model using the corrective or supplementary information, and regenerating the diagnostic results as the final diagnostic results.

6. The equipment remote fault diagnosis method based on a large model as described in claim 1, characterized in that, Following the step of automatically generating resource scheduling instructions based on the final diagnostic results, the method further includes: The final diagnosis results, resource scheduling instructions and execution results of this fault will be recorded as a new fault case. The new fault cases are characterized and then added to the fault case knowledge base to achieve self-updating of the fault case knowledge base.

7. The remote equipment fault diagnosis method based on a large model as described in claim 1, characterized in that, Before acquiring real-time sensor data, historical text data, and on-site multimedia data from the equipment, the following steps are also included: Retrieve historical failure cases similar to the current equipment status from the failure case knowledge base; The feature information of the retrieved historical failure cases is used as prior knowledge and input into the multimodal reasoning model to assist in failure reasoning.

8. A device for remote fault diagnosis, characterized in that, The device includes: The multimodal data fusion module is used to acquire real-time sensor data, historical text data and on-site multimedia data of the equipment, and to align and fuse the real-time sensor data and on-site multimedia data based on a unified time reference to generate multimodal feature vectors. The intelligent reasoning module is used to input the multimodal feature vector into the multimodal reasoning model based on the attention mechanism for processing, and to obtain an initial diagnostic result containing the cause of the fault and preliminary handling suggestions; An interactive verification module is used to respond to a verification request for the initial diagnostic result, provide an interactive interface to receive verification instructions, and generate a final diagnostic result based on the verification instructions; The safeguard decision module is used to generate resource scheduling instructions based on the final diagnostic results to execute remote safeguard operations.

9. An electronic device, characterized in that, Includes memory and one or more processors; The memory is coupled to the one or more processors, and the memory is used to store computer program code, the computer program code including computer instructions; The one or more processors invoke the computer instructions to cause the electronic device to perform the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium comprising instructions, characterized in that: When the instructions are executed on an electronic device, the electronic device causes the electronic device to perform the method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Equipment abnormity visual diagnosis method, electronic equipment and storage medium

    CN122090358A

  • Device abnormal visual diagnosis method, electronic device and storage medium

    CN122090358B