Large model enhanced equipment operation and maintenance multi-modal knowledge graph construction method and system and storage medium

By introducing multimodal large models to construct equipment operation and maintenance knowledge graphs, the problems of scarcity of samples and single modality in traditional methods are solved, and high-quality equipment operation and maintenance knowledge graphs are realized, which improves the accuracy and efficiency of operation and maintenance decisions.

CN120450011APending Publication Date: 2025-08-08CHONGQING UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510554938.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Traditional equipment operation and maintenance methods rely on a single text mode, resulting in scarce sample labeling, inability to effectively utilize image modal information, and it is difficult to build a comprehensive and high-quality knowledge map, which affects operation and maintenance efficiency and accuracy, especially in the analysis of new and old parts and complex equipment relationships.

Method used

The multimodal knowledge graph construction method enhanced by large model is adopted, and QWEN-VL is introduced through the dual-stream Transformer architecture to dynamically align images and text, combined with the MT-Transformer module for cross-modal timing fusion, and the multimodal dynamic weight attention module is used to improve the modeling accuracy of entity relationships, and complete the MNER and MRE tasks.

Benefits of technology

It improves the construction quality of equipment operation and maintenance knowledge graphs, enhances the continuity of the relationship between old and new parts and the causal chain embedding accuracy, improves the accuracy and efficiency of operation and maintenance decisions, and solves the problems of incomplete coverage of graph entities and broken causal chains in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120450011A_ABST
    Figure CN120450011A_ABST
Patent Text Reader

Abstract

The invention discloses a large model enhanced equipment operation and maintenance multi-modal knowledge graph construction method, which comprises the following steps of: firstly, introducing QWEN-VL into a double-flow Transform architecture, realizing dynamic alignment of an image and a text, supplementing few-sample semantics, and improving the identification precision of a part and problem entity relationship in the field of operation and maintenance; secondly, an MT-Transform module is provided, cross-modal time sequence embedding fusion is realized through fusion causal, expansion convolution and memory slot mechanisms, and the association continuity of new and old parts and causal chain embedding precision are improved; then, designing a multi-modal dynamic weight attention guiding module, introducing a weight key value to guide attention focus points, fusing schematic diagrams and text features, and improving entity relationship modeling precision of parts and faults; and finally, the marked image embedding and position embedding are fused into RoBERTa feature embedding, MNER and MRE tasks are completed in combination with CRF and SOFTMAX, and multi-modal operation and maintenance atlas construction is realized. The invention further provides a large-model-enhanced equipment operation and maintenance multi-modal knowledge graph construction system and a storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of equipment operation and maintenance technology, and specifically relates to a method, system and storage medium for constructing a large-model enhanced multimodal knowledge graph for equipment operation and maintenance. Background Art

[0002] As a key support link in industrial production systems, equipment operation and maintenance (O&M) focuses on ensuring continuous and stable operation of equipment under optimal conditions through systematic maintenance strategies, thereby maximizing equipment service life, reducing potential production safety hazards, and providing a solid foundation for improving enterprise production efficiency and economic benefits. As modern industrial equipment becomes increasingly complex and integrated, its operating status directly impacts the continuity of the production process, the consistency of product quality, and the safety of the operating environment. Implementing high-quality, intelligent O&M management not only reduces the frequency of unplanned downtime but also prevents production interruptions, quality defects, or major safety incidents caused by equipment failure.

[0003] Traditional equipment operation and maintenance solutions rely primarily on manual decision-making. Even maintenance experts need to study and review extensive maintenance documentation, resulting in low equipment operation and maintenance efficiency. This can easily lead to improper maintenance and missed inspections of critical components. This also hinders the systematic accumulation and standardized reuse of domain expert experience. Furthermore, a large amount of unstructured maintenance data, such as repair records and operation logs, is not effectively converted into reusable domain knowledge, hindering the iterative optimization of the domain knowledge system and potentially leading to data value loss. Consequently, intelligent operation and maintenance methods such as expert systems and knowledge bases have emerged. However, while expert systems and knowledge bases can provide regularized experience storage, their static knowledge representation and one-way reasoning mechanisms are ill-suited to the dynamic associations required by the high-density entities in multi-source data required for equipment operation and maintenance. Unlike traditional relational databases for equipment operation and maintenance, knowledge graphs, as structured knowledge bases, can support more precise associations between maintenance entities and efficient retrieval of maintenance knowledge. Their core characteristic is the use of a graph structure as a data representation, in which nodes represent knowledge entities in related domains, while edges characterize the relationships between these entities. Supported by formal semantics, this graph structure provides computers and models with an efficient and explicit knowledge reasoning mechanism. It can also construct complex contextual semantic networks, significantly enhancing the model's semantic parsing capabilities. Therefore, equipment operation and maintenance methods based on knowledge graphs can concisely and effectively represent equipment and fault data in a structured and semantic manner, while providing intelligent support for equipment operation and maintenance decisions. Equipment operation and maintenance methods based on knowledge graphs are highly dependent on the quality of the constructed graphs, which directly affects the accuracy and coverage of fault reasoning. However, current equipment operation and maintenance graph construction technologies have many problems, particularly because they only use a single textual form and ignore information from image modalities. This can easily lead to incomplete and inadequate information used to construct the graph, significantly affecting the quality of the operation and maintenance knowledge graph, and thus the efficiency and quality of knowledge graph-based operation and maintenance work.

[0004] First, in the field of industrial equipment operation and maintenance, some text entities describing equipment status and problems are poorly labeled, resulting in too few training samples. The rich feature information contained in the corresponding images, such as shape, position, and connection status, is often overlooked. Existing technologies are unable to effectively extract this part of the feature knowledge to construct a graph with few samples, resulting in incomplete entity coverage in the knowledge graph. This makes it easy for graph-based fault diagnosis to miss key status features, reducing the targeted nature of operation and maintenance measures. Second, equipment operation and maintenance is a long-term process. New parts coexist with old parts during the long-term operation of industrial equipment. The equipment has dynamic associations in structure, function, and status at different time sequences. Parts or problem descriptions in related texts are too far apart, making it impossible for existing models to simultaneously consider the relationship between new and old parts. Furthermore, discrete text makes it difficult to restore the continuous association between part entities, resulting in a time-sensitive fault in the constructed graph and a broken causal chain. This weakens the ability to model equipment status issues during operation and maintenance, making it difficult to support dynamic risk deduction and root cause tracing in preventive maintenance. Finally, the relationships between parts in industrial equipment are complex, and there is coupling between parts of different levels and types. The text descriptions of these connection states and part features are lengthy and complex. A description may contain multiple entities and relationships. This information is often more intuitively included in two-dimensional or three-dimensional schematics, such as Figure 1 As shown, relying solely on text is difficult to accurately mine entities and relationships. This can cause the constructed graph to miss part entities or part connection relationships, resulting in the lack of hierarchical assembly relationships and functional coupling logic in the graph. This impacts the systematic risk assessment and maintenance strategy formulation of equipment and increases the risk of misjudgment. Therefore, effectively extracting and integrating fault diagnosis and operation and maintenance knowledge accumulated in image modalities is a top priority for building a high-quality operation and maintenance knowledge graph and performing high-quality operation and maintenance. Summary of the Invention

[0005] In view of this, in order to solve the problems existing in the prior art, the purpose of the present invention is to provide a method, system and storage medium for constructing a large-model enhanced multimodal knowledge graph for equipment operation and maintenance.

[0006] In order to achieve the above object, the present invention provides the following technical solutions:

[0007] A method for constructing a large-model-enhanced multimodal knowledge graph for equipment operation and maintenance includes the following steps:

[0008] Step 1: Multimodal Large Model Embedding Enhancement

[0009] The pre-trained multimodal large model QWEN-VL is introduced. Using a dual-stream Transformer architecture, the input text and image data are initially embedded. This dynamically aligns cross-modal semantics, supplements few-shot semantic information, extracts cross-modal features, and mines implicit device features and complex associations by marking key image regions and text entities.

[0010] Step 2: Cross-modal temporal fusion

[0011] The MT-Transformer module was designed to leverage the local modeling capabilities of TCN to independently model both new and old text and image modal data in the operation and maintenance domain, storing them in multimodal memory slots. The attention mechanism was used to integrate new and old part features, strengthening the causal chain embedding in the temporal dimension.

[0012] Step 3: Dynamic Weight Attention Guidance:

[0013] Through the multimodal dynamic weighted attention module, combined with the cross-modal co-occurrence matrix and the unimodal adjacency matrix, the attention weights of text and images are adaptively adjusted to improve the modeling accuracy of complex part relationships;

[0014] Step 4: Complete the construction of the operation and maintenance map

[0015] MNER task: Combine the image features of the visual Transformer fused with QWEN-VL tags and position embedding with the text features of RoBERTa; use the CRF function to calculate the probability distribution of the label sequence, output the entity label sequence, and complete multimodal named entity recognition;

[0016] MRE task: splicing multimodal features, combining the Softmax function to calculate the relationship probability between specific entity pairs, screening high-confidence associations, and completing multimodal relationship extraction.

[0017] Furthermore, in step 1, the method steps for embedding and enhancing the multimodal large model are as follows:

[0018] 11) Use a two-stream architecture for data processing, using the RoBERTa model to extract semantic features from text, and ViT to extract local image features from images through the visual Transformer;

[0019] 12) Introducing the multimodal large language model QWEN-VL as an external knowledge base to optimize the embedding representations of RoBERTa and ViT;

[0020] 13) The embedding generated by the multimodal large model is combined with the dual-stream embedding through a weighted fusion strategy. The formula is:

[0021] E fused =α·E QWEN +β·(w t ·X t +w v ·X v w V )

[0022] Where: E fusedrepresents the two-dimensional embedding generated by the multimodal large model; X t and X v Represent the embeddings generated by RoBERTa and ViT respectively; α is the weight parameter of the large model embedding, which is used to control its contribution to the final fusion result; β is the weight parameter of the two-stream model embedding; w t and w v are the weight parameters of text and image features, w V It is the projection matrix corresponding to the image modality, which aims to embed the image into a two-dimensional space.

[0023] Furthermore, in step 2, the method steps for cross-modal time series fusion are as follows:

[0024] 21) Construct image modal memory slot M V and text modal memory slot M t , bimodal features are mapped to a unified key-value space through a key-value projection matrix to achieve data memory function;

[0025] 22) Integrate causal convolution and dilated convolution in the memory module, pass the input sequence layer by layer through multi-level temporal convolution units, and enhance the local perception ability of the memory module:

[0026]

[0027] Among them: H l is the hidden layer output of the lth layer of temporal convolution; k is the convolution kernel size; d is the expansion factor; s is the current time step index; σ is the nonlinear activation function ReLU; W i is the convolution kernel weight matrix; b is the bias term;

[0028] 23) A learnable expansion factor allocation strategy is used to dynamically adjust the expansion rate of each layer of temporal convolution through a gating mechanism:

[0029] d l =Sigmoid(W n H l-1 +b n )·d max

[0030] Where: d l is the expansion rate of the lth layer of temporal convolution; H l-1 is the hidden layer output of the l-1th layer of temporal convolution; d max is the preset maximum expansion rate, W n is the gating weight matrix, b n Gate bias

[0031] 24) Through the mapping matrix of the two modalities, the image and text features are mapped to the attention key-value space to obtain the K and V values of a single modality respectively;

[0032] 25) For the feature knowledge of the new input, the projection matrix of the graph data part is spatially compressed and encoded, and the query vector is generated respectively with the embedding of the sequence of the text data part; the historical memory and the current input features are fused through the multi-head attention mechanism to output the cross-modal temporal embedding.

[0033] Furthermore, in step 24), the K and V values of a single mode are obtained as follows:

[0034]

[0035] K fusion =W k [K v ⊙K t ]、

[0036] Among them: K v and K t Represents the K value of image and text modality respectively; V v and V t Represents the V value of image and text mode respectively; as well as Represents the image and text modal key-value projection matrices respectively; K fusion With V fusion is the cross-modal fusion vector; W k and W v are the weight ratios of K and V key values under the current task respectively; ⊙ represents the element-by-element product between modalities to strengthen the interaction of complementary features; Represents channel dimension concatenation to preserve modality specificity.

[0037] Furthermore, in step 25), the method of fusing historical memory and current input features through a multi-head attention mechanism to output a cross-modal temporal embedding is:

[0038] MultiHead(Q,K fusion ,V fusion )=Concat(head1,…,head h )W o

[0039]

[0040] Where: head represents the output of different attention heads; B represents the bias when processing the fusion vector; O represents the embedding result of the final output; Q is the query vector; h represents the number of heads of multi-head attention; d represents the total dimension of the input vector; W ° Represents the output projection matrix.

[0041] Furthermore, in step 3, the multimodal dynamic weight attention module performs the following method steps:

[0042] 31) Given the image modality embedding and text modality embedding, construct the co-occurrence matrix C based on cross-modal co-occurrence statistics co and the relative position matrix R within the modal;

[0043] 32) Define the visual head and text head attention mechanism, respectively for the visual head and text head:

[0044]

[0045] in: and They are text header and visual header respectively; and is the linear transformation matrix of query, key and value; R t and R v are the relative position matrices within the text modality embedding and image modality respectively; x t and x v They are text modality embedding and image modality embedding respectively;

[0046] 33) Through residual connection, the original feature data of a single modality is imported to guide the dynamic weight key value to obtain the predicted classification

[0047] 34) Introducing a learnable external parameter dynamic weight key, by classifying the prediction Gradient descent is performed with the data label through the cross entropy function, and the value vector in the attention is interactively weighted to optimize the entity relationship modeling.

[0048] Furthermore, in step 33), the prediction classification for:

[0049]

[0050] Where: H' T is the text feature guided by the dynamic weight key value; H' V is a visual feature guided by dynamic weight key value; H fusion is the image-text fusion feature; W clf is the weight matrix of the classifier;

[0051] In step 34), calculate L CE For K learnabel Gradient:

[0052]

[0053] Where: LCE is the loss function; N is the total number of samples, C is the number of categories, y ij is the true label of the i-th sample belonging to the j-th class; η is the learning rate; K learnable are learnable external parameters;

[0054] The interactive modality weighting method for the value vector in the attention is:

[0055]

[0056] Where: H is the multimodal bias of the V vector; head* is the updated attention head.

[0057] Furthermore, the applications of the multimodal large model QWEN-VL include:

[0058] Visually mark the position and connection status of parts in the equipment schematic;

[0059] Parse the entities and relations in the fault log text and generate triple prompts;

[0060] The labeled visual embedding and position embedding are integrated into the text features of RoBERTa, and CRF and Softmax are combined to complete the MNER and MRE tasks.

[0061] The present invention also proposes a large-model enhanced equipment operation and maintenance multimodal knowledge graph construction system for executing the above-mentioned equipment operation and maintenance multimodal knowledge graph construction method, which is characterized by comprising:

[0062] A multimodal data input module for receiving text and image data;

[0063] QWEN-VL large model interface for cross-modal semantic alignment and few-shot enhancement;

[0064] MT-Transformer temporal fusion module, used for associative modeling of new and old parts;

[0065] MDWAM attention guidance module for complex relationship analysis;

[0066] The knowledge graph output module is used to generate structured equipment operation and maintenance graphs.

[0067] The present invention also proposes a storage medium storing a computer program, which, when executed by a processor, implements the steps of the method for constructing a large-model-enhanced multimodal knowledge graph for equipment operation and maintenance as described above.

[0068] The beneficial effects of the present invention are:

[0069] Industrial equipment operation and maintenance is a core component of ensuring production safety and efficiency, and urgently requires intelligent technology support. Knowledge graphs, which represent equipment and faults using a graph structure, support efficient association and rapid search of operation and maintenance knowledge and are widely used in intelligent equipment operation and maintenance decision-making. Traditional knowledge graph construction methods rely on a single text modality and face challenges such as scarce sample annotations, difficulty in dynamically associating new and old equipment, and insufficient analysis of complex equipment relationships. These issues lead to missing graph entities and broken causal chains, which in turn affect decision-making quality. Therefore, this paper proposes a large-model-enhanced multimodal knowledge graph construction method for equipment operation and maintenance. By introducing a large multimodal model, we fully utilize the multimodal knowledge in the operation and maintenance domain, effectively improve the understanding and modeling of complex operation and maintenance knowledge, and enhance the quality of knowledge graph construction. First, we introduce QWEN-VL into the two-stream Transformer architecture to achieve dynamic alignment of images and text and supplement few-shot semantics, thereby improving the accuracy of part and problem entity relationship recognition in the operation and maintenance domain. Second, we propose the MT-Transformer module, which integrates causal, dilated convolution, and memory slot mechanisms to achieve cross-modal temporal embedding fusion, improving the continuity of associations between new and old parts and the accuracy of causal chain embedding. Then, we design a multimodal dynamic weighted attention guidance module, introduce weighted keys to guide attention focus, and fuse schematic and text features to improve the accuracy of part and fault entity relationship modeling. Finally, to fully utilize the multimodal understanding capabilities of the large multimodal model, its labeled image embedding and position embedding are integrated into the feature embedding of RoBERTa. Combined with CRF and SOFTMAX, this module performs MNER and MRE tasks, thus realizing the construction of a multimodal operation and maintenance graph. The multimodal operation and maintenance graph constructed by the present invention can be effectively used in equipment operation and maintenance. During the graph reasoning process, the multimodal graph reasoning effect is significantly better than that of a single modality, which proves the superiority of the multimodal operation and maintenance graph constructed by the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] In order to make the purpose, technical solutions and beneficial effects of the present invention more clear, the present invention provides the following drawings for illustration:

[0071] Figure 1 Comparison of text and visual representation of equipment failure under two different semantics;

[0072] Figure 2 The overall architecture of the method for constructing a multimodal knowledge graph for equipment operation and maintenance enhanced by the large model of the present invention;

[0073] Figure 3 This is a schematic diagram of the structure of MT-Transformer;

[0074] Figure 4 Schematic diagram of the multimodal dynamic weight-guided attention module;

[0075] Figure 5is the accuracy and loss value of MllmDA-KGC in MNER and MRE;

[0076] Figure 6 For text-image pairs and the instructions used;

[0077] Figure 7 Schematic diagram of two prediction results of MEGA, EEGA and MllmDA-KGC;

[0078] Figure 8 Visualization of multimodal graphs and single modal graphs in the operation and maintenance field. DETAILED DESCRIPTION

[0079] The present invention will be further described below with reference to the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it. However, the embodiments are not intended to limit the present invention.

[0080] like Figure 2 As shown, this is the overall architecture of the method for constructing a multimodal knowledge graph for equipment operation and maintenance proposed in this embodiment. The overall architecture of this embodiment is divided into four stages. First, the data input layer inputs pictures and texts in the operation and maintenance field, and simultaneously performs multimodal named entity recognition (MNER) and multimodal relationship extraction (MRE) tasks; secondly, in order to solve the problem of few samples in the operation and maintenance field, we combine the dual-stream Transformer architecture and introduce a multimodal large model QWEN-VL, which can combine text information and corresponding image information, and mark the key information therein, and assist the framework in deeply exploring the key features implicit in the operation and maintenance process and the complex correlation between devices; thirdly, we propose a multimodal sequence convolutional attention module MT-Transformer, which integrates the local modeling capabilities of TCN to respectively analyze the new and old text modalities in the operation and maintenance field. After independent modeling and storage with the image modal data, attention-based feature fusion is performed to establish a multimodal memory slot (Memory_Bank) module to strengthen the continuity of the association between new and old parts and improve the embedding accuracy of the causal chain between parts and problem entities in the time dimension; then, a multimodal dynamic weight guided attention module MDWAM is proposed to establish cross-modal co-occurrence and unimodal adjacency analysis, introduce dynamic weight attention guidance keys to guide the direction of attention, accurately model complex parts and their associated structures, and improve the accuracy of graph construction; finally, at the task output stage, in order to better utilize the multimodal information processing capabilities of the multimodal large model, the image features fused with QWEN-VL tags in the visual Transformer are The position embedding Token Embedding is integrated into the feature embedding of the final output of RoBERTa, and then the probability distribution of the label sequence is calculated using the CRF function, and the entity label sequence is output to complete the MNER task. Similarly, in order to ensure a comprehensive understanding and combination of the output of different modalities, the image features are spliced. With text features Then, combined with the Softmax function, the relationship probability between specific entity pairs is calculated to complete the MRE task and finally complete the construction of the operation and maintenance map.

[0081] In the field of industrial equipment operation and maintenance, due to factors such as enterprise data privacy protection and the non-transferability of experience caused by differences in equipment models, there is little high-quality labeled data available for model training. This greatly limits the generalization ability of the model in vertical fields. On the one hand, core data in operation and maintenance scenarios, such as equipment fault logs and maintenance work orders, are mostly stored in a decentralized manner in the form of a mixture of unstructured text and images, and involve commercial secrets that are difficult to openly access; on the other hand, the professionalism and long-tail characteristics of equipment mechanism knowledge, such as the wear threshold of a specific model of bearing, lead to high labeling costs. Traditional pre-training models are difficult to directly capture domain-specific entity relationships and failure modes. At the same time, in the process of industrial equipment operation and maintenance, the relevant operation and maintenance processes have a long cycle, the entity density in the operation and maintenance knowledge is high, and the entities in different time series are highly discrete, resulting in the possibility that previous knowledge will be forgotten and cannot interact and integrate well with the latest knowledge. This will cause the embedding loss of parts or problem features, which in turn affects a series of subsequent operations. To address the above issues, this embodiment proposes a multimodal temporal Transformer architecture MT-Transformer, which is inspired by TCN. The multimodal memory slot constructed has both local perception characteristics and global dependency modeling capabilities of integrated attention. By introducing a hybrid hierarchical structure of causal convolution and dilated convolution, a hidden layer module with temporal memory enhancement characteristics is constructed to adapt to the needs of joint embedding of historical knowledge features and new knowledge features in the process of device state evolution. The MT-Transformer structure is as follows: Figure 3 shown.

[0082] Specifically, the method for constructing a multimodal knowledge graph for equipment operation and maintenance enhanced by a large model in this embodiment includes the following steps.

[0083] Step 1: Multimodal Large Model Embedding Enhancement

[0084] The pre-trained multimodal large model QWEN-VL is introduced, and the input text data and image data are preliminarily embedded respectively through the dual-stream Transformer architecture, dynamically aligning cross-modal semantics, supplementing few-sample semantic information, extracting cross-modal features, and mining the implicit features and complex associations of devices by marking key image areas and text entities.

[0085] Specifically, large language models refer to Transformer language models with hundreds of billions of parameters or more. These models have become powerful tools in the field of natural language processing. After pre-training, they can have certain data parsing capabilities in different fields. In order to fully utilize the knowledge guidance capabilities of large language models, we designed a mechanism of multimodal embedding and weighted fusion. For the selection of multimodal large models, we chose the locally deployed QWEN-VL. In the analysis of industrial operation and maintenance data, the multimodal large language model QWEN-VL can automatically extract key entities and relationship concepts from pre-processed text data according to prompts. By parsing text data such as large-scale industrial fault descriptions, common entity and relationship concepts such as "components", "parts", "faults", "causes", "solutions", and "connected to" can be identified, and data embedding guidance can be provided in combination with relevant pictures. On this basis, the preliminary extracted concepts are fused with image features to ensure the accuracy and relevance of each concept and finally the relevant data is labeled. Specifically, the method steps for multimodal large model embedding enhancement are as follows:

[0086] 11) Preliminary embedding of text and images: A two-stream architecture is used for data processing. The RoBERTa model is used to extract semantic features from text, and ViT uses a visual Transformer to extract local image features from images.

[0087] Adjusting the high-level structure of the language model can more effectively apply knowledge to downstream tasks. Therefore, this example uses a two-stream architecture for data processing, using RoBERTa and ViT for preliminary data embedding of text and image data, respectively. RoBERTa extracts semantic information from unstructured text such as equipment fault logs and maintenance work orders, while ViT uses a visual transformer to capture visual features from equipment schematics or photos. This step lays the foundation for subsequent multimodal large-scale model embedding fusion.

[0088] 12) Large model-assisted embedding generation: Introducing the multimodal large language model QWEN-VL as an external knowledge base to optimize the embedding representation of RoBERTa and ViT.

[0089] Specifically, based on the preliminary feature extraction, this embodiment introduces the multimodal large language model QWEN-VL as an external knowledge base to assist in the generation of the final embedding. Specifically, QWEN-VL analyzes the text and corresponding image data under the predefined concept prompts such as "component", "part", "location", and "connected to", helping to optimize the embedding representation of RoBERTa and ViT. For example, when the text mentions "bearing wear", QWEN-VL can supplement the relevant information embedding such as location, shape, etc. based on its knowledge base and the corresponding image, thereby enhancing the contextual relevance of the text embedding.

[0090] 13) Weighted fusion strategy for large models and dual-stream frameworks: The embeddings generated by the multimodal large model are combined with the dual-stream embeddings through a weighted fusion strategy.

[0091] Specifically, based on the introduction of a large model as an external knowledge base, this embodiment integrates the information collected by the dual-stream architecture as the final embedded representation. Specifically, the input image is embedded as a three-dimensional tensor Where H and W represent the height and width of the image respectively, d represents the embedding dimension, and the text data is represented as a two-dimensional matrix Here, l is the sequence length, and d represents the embedding dimension. To ensure that the content extracted by the multimodal large model can be effectively combined with the content extracted by the ViT and RoBERTa models, the module adopts a weighted fusion strategy and dynamically adjusts the contribution ratio of the two. The specific formula is as follows:

[0092] E fused =α·E QWEN +β·(w t ·X t +w v ·X v w V )

[0093] Where: E fused represents the two-dimensional embedding generated by the multimodal large model; X t and X v Represent the embeddings generated by RoBERTa and ViT respectively; α is the weight parameter of the large model embedding, which is used to control its contribution to the final fusion result; β is the weight parameter of the two-stream model embedding; w t and w v are the weight parameters of text and image features, w V It is the projection matrix corresponding to the image modality, which aims to embed the image into a two-dimensional space.

[0094] The dynamic adjustment of weight parameters is based on task requirements and data characteristics. For tasks that require high expertise, the large model is given a higher weight α; for tasks that rely on low-level features, the weight β of the ordinary model is increased, and w is dynamically adjusted according to the correlation between text and image. t and w v After this step, the data of the two modalities are spliced with the fusion knowledge respectively: image data and the corresponding fusion knowledge E fused Combined, the final output is Similarly, the text representation is enhanced to X′ t =X t ·E fused .

[0095] Step 2: Cross-modal temporal fusion

[0096] The MT-Transformer module is designed to leverage the local modeling capabilities of TCN to independently model the new and old text modal and image modal data in the operation and maintenance field, and store them in multimodal memory slots. The new and old part features are integrated through the attention mechanism to strengthen the causal chain embedding in the time dimension.

[0097] Specifically, a temporal convolutional network consists of causal convolution, dilated convolution, and residual blocks. Causal convolution is characterized by the fact that the output at the current moment depends only on the current and previous inputs. This design effectively avoids data leakage. However, to capture more information, causal convolution often requires a more complex network structure. In contrast, dilated convolution compensates for this deficiency by simplifying the structure while expanding the receptive field through interval sampling.

[0098] In this embodiment, the method steps for cross-modal time series fusion are as follows:

[0099] 21) Construct image modal memory slot M V and text modal memory slot M t , the bimodal features are mapped to a unified key-value space through the key-value projection matrix to realize the memory function of the data.

[0100] To effectively memorize complex equipment knowledge during long-term operation and maintenance, this embodiment proposes a multimodal knowledge storage mechanism based on sequential convolution, inspired by the TCN network. By explicitly constructing memory banks for image and text modalities, it achieves persistent storage and dynamic call of cross-modal prior knowledge. The specific definition is:

[0101]

[0102] Where: Mv is the image modality memory slot; Mt is the text modality memory slot; N is the memory slot depth. To address the problem of different shapes of dual-modality memory slots, this embodiment defines the dual-modality key value projection matrix W v 、W t , mapping image and text features into a key-value space of uniform shape Realize the memory function of data.

[0103] 22) The causal convolution and dilated convolution are integrated into the memory module, and the input sequence is passed layer by layer through multi-level temporal convolution units to enhance the local perception ability of the memory module.

[0104] In order to effectively extract and embed the comprehensive features of the historical knowledge of the two modalities in the operation and maintenance field, this embodiment integrates causal convolution and dilated convolution in the memory module to enhance the local perception ability of the memory module. Specifically, without considering the sequence length, for a given input sequence Where T is the time step, d is the feature dimension, and the hidden layer output feature The multi-level temporal convolution unit is passed layer by layer, and its mathematical expression is:

[0105]

[0106] Among them: H l is the hidden layer output of the lth layer of temporal convolution; k is the convolution kernel size; d is the expansion factor; s is the current time step index; σ is the nonlinear activation function ReLU; W i is the convolution kernel weight matrix; b is the bias term.

[0107] By dynamically adjusting k and d, the model can adaptively build multi-granularity time resolution analysis capabilities. That is, smaller k and d are used to capture historical local equipment or problem characteristics within a shorter period, while larger parameter combinations are extended to different equipment or problem characteristics within a longer period to model cross-period historical patterns.

[0108] 23) A learnable expansion factor allocation strategy is adopted to dynamically adjust the expansion rate of each layer of temporal convolution through a gating mechanism.

[0109] Furthermore, to avoid the problem of feature scale uniformity caused by the fixed expansion coefficient in traditional TCN, this embodiment designs a learnable expansion factor allocation strategy, which dynamically adjusts the expansion rate d of each layer through a gating mechanism. l , its calculation process can be expressed as:

[0110] d l =Sigmoid(W n H l-1 +b n )·d max

[0111] Where: d l is the expansion rate of the lth layer of temporal convolution; H l-1 is the hidden layer output of the l-1th layer of temporal convolution; d max is the preset maximum expansion rate, W n is the gating weight matrix, b n is the gate bias.

[0112] This strategy enables the model to autonomously optimize the receptive field range based on the temporal characteristics of the input sequence, thereby enhancing the long-term causal modeling capability while retaining local details.

[0113] 24) After performing the above operations on the content in the memory slot, the image and text features are mapped to the attention key-value space through the mapping matrix of the two modalities, and the K and V values of each modality are obtained respectively, which are specifically expressed as:

[0114]

[0115] K fusion =W k [K v ⊙K t ]、

[0116] Among them: K v and K t Represents the K value of image and text modality respectively; V v and V t Represents the V value of image and text mode respectively; as well as Represents the image and text modal key-value projection matrices respectively; K fusion With V fusion is the cross-modal fusion vector; W k and W v are the weight ratios of K and V key values under the current task respectively; ⊙ represents the element-by-element product between modalities to strengthen the interaction of complementary features; Representation channel dimension splicing to preserve modality specificity. These two operations explicitly model nonlinear interactions between modalities and preserve independent features, respectively. This ensures cross-modal semantic alignment while avoiding oversmoothing of modality-specific information, thereby adapting to the complex correlation characteristics of multimodal data in industrial scenarios.

[0117] 25) For the feature knowledge of the new input, the projection matrix of the graph data part is spatially compressed and encoded, and the query vector is generated respectively with the embedding of the sequence of the text data part; the historical memory and the current input features are fused through the multi-head attention mechanism to output the cross-modal temporal embedding.

[0118] For the newly input feature knowledge, this embodiment divides it into two parts and processes them separately. The projection matrix W of the graph data part is v After spatial compression encoding, the query vectors are generated by embedding the partial sequence of text data. and On this basis, we use the multi-head attention scaled dot product form to obtain the attention score, and finally use residual connection and layer normalization to enhance the stability of gradient propagation. Specifically, the method of fusing historical memory and current input features through the multi-head attention mechanism to output cross-modal temporal embedding is as follows:

[0119] MultiHead(Q,K fusion ,V fusion )=Concat(head1,...,head h )W o

[0120]

[0121] Where: head represents the output of different attention heads; B represents the bias when processing the fusion vector; O represents the embedding result of the final output; Q is the query vector; h represents the number of heads of multi-head attention; d represents the total dimension of the input vector; W ° Represents the output projection matrix.

[0122] Step 3: Dynamic Weighted Attention Guidance

[0123] Through the multimodal dynamic weight attention module, combined with the cross-modal co-occurrence matrix and the unimodal adjacency matrix, the attention weights of text and images are adaptively adjusted to improve the modeling accuracy of complex part relationships.

[0124] The operation and maintenance of industrial equipment involves a large number of complex devices and issues. These devices are typically composed of multiple highly coupled parts, and their operating status is affected by a combination of factors. Therefore, the resulting operation and maintenance knowledge is often characterized by high entity density and complexity. Long text-based maintenance logs or fault analysis reports contain a large number of descriptions of equipment parts and their interrelationships. The association structure between these entities is often very complex, involving multi-level causal chains, functional dependencies, and spatial constraints. In addition, image data such as equipment schematics or on-site photos contain rich assembly constraints and functional semantic information, but this information is often difficult to fully mine and utilize through traditional text modal methods.

[0125] In response to the above challenges, this embodiment proposes a multimodal dynamic weighted attention module (MultimodalDynamic Weighted Attention), which aims to integrate device schematics and text features, analyze cross-modal co-occurrence analysis and unimodal adjacency relationships and dynamic weight values, and combine the dynamic weight value adjustment mechanism to accurately model complex coupled part entities and their association relationships in long texts. Specifically, the module first performs deep feature extraction on text and image data, and uses cross-modal co-occurrence analysis to mine implicit associations between the two modalities, such as the "motor" mentioned in the text and the corresponding part position in the image, and captures local structural information within the same modality through unimodal adjacency analysis, such as the entity context relationship in the text or the spatial layout in the image. On this basis, this embodiment introduces an external dynamic weight key (Dynamic Weight Key) to adaptively adjust the contribution ratio of different modalities in the attention mechanism, thereby alleviating the problem of dynamic imbalance in inter-modal learning, extracting and modeling entity features in complex texts and images, and providing a data basis for subsequent MNER and MRE tasks. The multimodal dynamic weighted attention module is shown in the figure below. Figure 4 Specifically, the multimodal dynamic weight attention module performs the following method steps:

[0126] 31) MDWAM constructs the co-occurrence matrix C based on cross-modal co-occurrence statistics under the premise of given image modality embedding and text modality embedding co And the relative position matrix R within the modal. The specific calculation methods of the two matrices are as follows:

[0127]

[0128] Where: i represents the index of the i-th visual embedding element; j represents the index of the j-th text embedding element; D is the number of contexts; δ is an indicator function that returns 1 if the visual embedding element i and the corresponding text element appear in the context D, otherwise it returns 0. The final C co Each element in the matrix represents the number of times the corresponding visual and textual elements co-occur.

[0129] R t / v (i,j)=|pos(i)-pos(j)|

[0130] Where: pos(i) represents the position of element i; pos(j) represents the position of element j; i and j represent the position index of each element. By calculating the distance between the two index embeddings in the index space, the relative position of the embedded content in a single modality is determined.

[0131] 32) This embodiment considers the relative distance matrix within a single modality and the contribution matrix within different modalities, and defines the visual head and text head attention mechanisms, which are visual head and text head respectively.

[0132]

[0133]

[0134] in: and They are text header and visual header respectively; and is the linear transformation matrix of query, key and value; R t and R v are the relative position matrices within the text modality embedding and image modality respectively; x t and x v They are text modality embedding and image modality embedding respectively.

[0135] Where R t / v The larger the value is, the farther the two are from each other in the text or image. In this case, the model is more likely to notice the correlation between them. On the contrary, the model pays more attention to local features. co The larger the value, the more the model focuses on the multimodal fusion information at this time. Conversely, the model focuses more on the unimodal features.

[0136] 33) Through residual connection, the original feature data of a single modality is imported to guide the dynamic weight key value to obtain the predicted classification

[0137] In the attention mechanism, the value vector is a very important component, which carries the actual information content of the input data. In order to better process multimodal complex data, this embodiment introduces a learnable external parameter dynamic weight key and imports the original feature data H of a single modality through residual connection. (i) Provide guidance on dynamic weight keys by classifying predictions The gradient descent is performed with the data label through the cross entropy function, the value vector in the attention is weighted in an interactive mode, and the selection of the value vector in a single mode is optimized, so that the model can better pay attention to the complex entity relationship in the operation and maintenance process. Specifically, the prediction classification for:

[0138]

[0139] Where: H' T is the text feature guided by the dynamic weight key value; H' V is a visual feature guided by dynamic weight key value; H fusion is the image-text fusion feature; Wclf is the weight matrix of the classifier.

[0140] 34) Introducing a learnable external parameter dynamic weight key, by classifying the prediction Gradient descent is performed with the data label through the cross entropy function, and the value vector in the attention is interactively weighted to optimize the entity relationship modeling.

[0141] Specifically, by fusing features, the model can simultaneously consider the detailed features of the text and the corresponding image, thereby enhancing the model's feature extraction capability. Based on the chain rule, this embodiment calculates L CE For K learnable Gradient:

[0142]

[0143] Where: L CE is the loss function; N is the total number of samples, C is the number of categories, y ij is the true label of the i-th sample belonging to the j-th class; η is the learning rate; K learnable are learnable external parameters.

[0144] Afterwards, through K learnable The V vector is weighted to guide the attention head to update, so that the attention head of this embodiment can make the most of the multimodal information. Specifically, the interactive modality weighting method for the value vector in the attention is:

[0145]

[0146] Where: H is the multimodal bias of the V vector; head* is the updated attention head.

[0147] The feature content of the visual header is then embedded with the text header to complete feature extraction and provide data conditions for the completion of the subsequent graph construction task.

[0148] Step 4: Complete the construction of the operation and maintenance map

[0149] MNER task: Combine the image features of QWEN-VL tags and position embedding in the visual Transformer with the text features of RoBERTa; use the CRF function to calculate the probability distribution of the label sequence, output the entity label sequence, and complete multimodal named entity recognition.

[0150] MRE task: splicing multimodal features, combining the Softmax function to calculate the relationship probability between specific entity pairs, screening high-confidence associations, and completing multimodal relationship extraction.

[0151] Specifically, in this embodiment, the application of the multimodal large model QWEN-VL includes: visually marking the position and connection status of parts in the equipment schematic; parsing the entities and relationships in the fault log text to generate triple prompts; integrating the marked visual embedding and position embedding into the text features of RoBERTa, and combining CRF and Softmax to complete MNER and MRE tasks.

[0152] This embodiment also proposes a large-model enhanced equipment operation and maintenance multimodal knowledge graph construction system, which is used to execute the equipment operation and maintenance multimodal knowledge graph construction method described above in this embodiment, including:

[0153] A multimodal data input module for receiving text and image data;

[0154] QWEN-VL large model interface for cross-modal semantic alignment and few-shot enhancement;

[0155] MT-Transformer temporal fusion module, used for associative modeling of new and old parts;

[0156] MDWAM attention guidance module for complex relationship analysis;

[0157] The knowledge graph output module is used to generate structured equipment operation and maintenance graphs.

[0158] This embodiment also proposes a storage medium storing a computer program, which, when executed by a processor, implements the steps of the method for constructing a large-model-enhanced multimodal knowledge graph for equipment operation and maintenance as described above in this embodiment.

[0159] The specific effects of the large-model enhanced equipment operation and maintenance multimodal knowledge graph construction method of the present invention are described in detail below in combination with specific experiments.

[0160] To demonstrate the superiority of the method framework proposed in this embodiment, we used the automobile operation and maintenance dataset to verify our multimodal graph construction method, and conducted experimental analysis from multiple perspectives. The final results showed that the MllmDA-KGC model proposed in this embodiment outperformed other baseline methods in all the experiments shown.

[0161] (1) Experimental Dataset

[0162] The data set used in this embodiment comes from the automobile operation and maintenance records of a certain automobile company, which records the vehicle failures and maintenance experience from 2022 to 2024. It includes 440 fault handling processes, and each handling process includes in detail the vehicle components (components) that appeared in the fault, the parts in the components (parts), the cause of the fault (cause), the result caused by the fault (result), the technology used for maintenance (technology), and the characteristics of the components, causes and technologies. To eliminate recurring problems, we counted the specific faults with the highest proportion in different system parts in the data, the number of fault entries and the proportion in the operation and maintenance records, as shown in Table 1. In the schematic diagram of the operation and maintenance records in the table, the left side is the component schematic diagram and the right side is the fault phenomenon diagram. At the same time, by summarizing the various components and problems that appeared in the operation and maintenance data, we defined 14 operation and maintenance relationships. All operation and maintenance relationships are shown in Table 2. For each relationship in the table, we define three elements. The first two are entity categories and the third is a relationship. This relationship is a directed relationship, from the first entity to the second.

[0163] Table 1 The most common faulty components and schematic diagrams in vehicle operation and maintenance records

[0164]

[0165] Table 2 Schematic diagram of vehicle operation and maintenance relationship

[0166]

[0167] (2) Evaluation indicators

[0168] In the field of graph construction, model evaluation usually relies on the following indicators: Precision, Recall, and F1 score. These are key indicators for evaluating whether the model can accurately extract entities and relationships. Their calculation method is as follows:

[0169]

[0170] TP represents entities or relationships that were correctly identified and added to the knowledge graph; FP represents entities or relationships that were incorrectly identified as correct but were incorrect or did not exist; and FN represents entities or relationships that were not identified but actually existed. Of these three metrics, precision focuses on ensuring that the added information is correct, while recall strives to ensure that all required information is correctly identified. The F1 value provides a balance between these two, helping to comprehensively evaluate model performance. Therefore, the F1 value is used as the primary evaluation metric, with precision and recall serving as auxiliary evaluation indicators.

[0171] (3) Model overall experiment

[0172] In order to verify the effectiveness and superiority of the framework proposed in this embodiment, an overall experiment was first conducted on it. In the experiment, the PyTorch architecture was used to implement the MllmDA-KGC model. Specifically, the input text data was limited to a single piece of data 100-200Tokens, the corresponding input image resolution was unified to 224×224, the network hidden layers of the dual-stream architecture were all set to 12 layers, the hidden layer size was set to 768×768, and the memory slot depth in MT-Transformer was set to 100. The AdamW algorithm was selected as the optimization algorithm for the neural network. In addition, in order to alleviate the overfitting problem, we used a dropout layer, and the dropout rate was defined as 0.2. During the training process, the number of epochs was set to 100, the batch size was set to 16, and the learning rate was set to 3e-5. The model accuracy and training loss are shown in Figure 2. Figure 5 shown.

[0173] As can be seen from the model training diagram, the model can achieve a high F1 value under the dataset. Although there are certain accuracy fluctuations during the training process, especially in the MNER task, it can ultimately achieve good convergence and high accuracy, which means that it has very good results in multimodal operation and maintenance knowledge extraction.

[0174] To validate the performance advantages of the multimodal knowledge graph construction framework MllmDA-KGC proposed in this example in the field of equipment operation and maintenance, several cutting-edge general-purpose multimodal knowledge graph construction models were selected as baseline methods for comparative analysis: MEGA optimizes cross-modal relationship reasoning through graph structure alignment, UMGF implements entity extraction using a fine-grained graph-text alignment strategy, TMR establishes a cross-modal translation mechanism for semantic mapping, and EEGA jointly models entities and relationships through edge-enhanced graph networks. Comparisons with the state-of-the-art model and MllmDA-KGC validated the effectiveness of the proposed model, as shown in Table 3.

[0175] Table 3 Performance comparison of different models for NER and RE

[0176]

[0177]

[0178] In the comparative experimental results, the superior performance indicators are retained for ease of observation. Because the MEGA and UMGF models are not traditional joint extraction models, each model's specific task focus makes it difficult to effectively understand and integrate the features extracted for the MNER and MRE tasks. Among the four compared models, MEGA achieved relatively low F1 scores of 70.23% and 73.06% for the MNER and MRE tasks, respectively. This is due to the model's heavy reliance on image annotations and auxiliary images, which are scarce in the field of equipment operation and maintenance and fail to meet the model's training requirements. Similarly, UMGF achieved scores of 71.33% and 73.31% for the MNER and MRE tasks, respectively, 17.07% and 20.48% lower than MllmDA-KGC. Experimental results for multimodal joint extraction models show that TMR and EEGA achieved F1 scores of 77.93%, 79.05%, 80.64%, and 82.59% for the MNER and MRE tasks, respectively. These scores lag significantly behind the F1 scores of 88.40% and 93.79%, respectively, achieved by MllmDA-KGC on the same task. This is due to the inability of the TMR and EEGA models to accurately model the entities and relations in complex text within the O&M domain, resulting in the inability to correctly extract all relevant entities and relations. Although the EEGA model incorporates an edge-enhanced graph alignment network, it cannot guarantee the precise parsing of all corresponding relations and fails to address the problem of knowledge forgetting when processing long text inputs. This significant performance gap effectively demonstrates the superiority of the proposed method.

[0179] (4) Multimodal large model optimization

[0180] To determine the most appropriate large-scale multimodal language model, this embodiment conducted large-scale multimodal language model analysis and testing in the field of equipment operation and maintenance to verify the multimodal information understanding capabilities of different large-scale multimodal language models. In the experiment, relatively mature large-scale multimodal language models from the industry were selected for testing. At the same time, to ensure the privacy of enterprise data, the models were deployed locally. Specifically, each large-scale multimodal language model was provided with images and text from the corresponding dataset, and the large-scale model was required to parse the images and their corresponding text. To demonstrate the versatility of the model and reduce its reliance on domain data, the model was not fine-tuned.

[0181] For ease of observation, instead of outputting data in an embedded format, the entities and relationships related to equipment operation and maintenance are displayed directly. The same prompts are provided to different large-scale models to guide the output of entities and potentially related triplets based on the fusion of graphical and textual information. The accuracy of triplet extraction is analyzed to assess the model's attention to multimodal information in operation and maintenance.

[0182] For our dataset, the following instructions are given: "Perform in-depth analysis of images and texts about industrial equipment operation and maintenance, and output the entities and relationships emphasized in the text. Entities include components, causes, technologies, and characteristics; relationships include connections to, components, fault phenomena, fault causes, and characteristics. To a certain extent, entity relationships are supplemented, and triples are output." The text and image pairs input in the experiment are as follows: Figure 6 The specific experimental results are shown in Table 3.

[0183] Table 3. Triplet output of different large models

[0184]

[0185] For easier observation, we mark the results that do not meet the prompt requirements with dotted lines. Mark incorrect output results with a dot-dash line Mark the content output by parsing the text with a solid line Draw a double solid line to mark the output content of the fused graphic and text information Experiments on the parsing of multimodal knowledge related to industrial equipment operation and maintenance using large multimodal language models show that LLaVA fails to fully understand the meaning of prompts containing industrial entities and relationships, and the output differs from our expectations. This is undesirable when guiding the model to embed the target. Under the guidance of prompts, the remaining large models can parse the text content and effectively output the multiple triples present in the operation and maintenance text. However, in miniGPT4o, some relationships do not meet our requirements, which may introduce some errors in the construction of the overall equipment operation and maintenance map. In experiments with InternVL and QWEN-VL, both can accurately parse the operation and maintenance text data according to prompt instructions. However, when parsing multimodal knowledge related to operation and maintenance, QWEN-VL can effectively incorporate entity location information in the image to supplement the text content, indicating that the model can effectively attend to the entities and relationships in the industrial equipment operation and maintenance process under prompt instructions.

[0186] (5) Module contribution quantification experiment

[0187] In this embodiment, the model consists of three main parts: the MT-Transformer data embedding part driven by the multimodal large model after data input, the MDWAM feature extraction part, and the CRF\SOFTMAX classification part. In order to verify the effectiveness of several parts of the model proposed in this embodiment and quantify the contribution of each module, we designed four sets of ablation experiments. In the experiment, since it is necessary to process data of two modalities, we retain the dual-stream architecture as the baseline. At the same time, we distinguish whether MT-Transformer is integrated into the large model, and divide it into MT-Transformer* (removing the multimodal large model) and MT-Transformer (combining the multimodal large model) to better verify the necessity of multimodal large model driving. The experiments were designed with consistent dataset selection and model parameter settings. Four model module configurations were designed: (1) dual-stream framework + MT-Transformer + MDWAM + CRF / SOFTMAX; (2) dual-stream framework + MT-Transformer* + MDWAM + CRF / SOFTMAX; (3) dual-stream framework + MDWAM + CRF / SOFTMAX; and (4) dual-stream framework + MT-Transformer + CRF / SOFTMAX. The experimental results are shown in Tables 4 and 5.

[0188] Table 4 Ablation test results of the model in the MNER task

[0189]

[0190] Table 5 Ablation experiment results of the model in the MRE task

[0191]

[0192] It can be seen that in the MNER and MRE tasks, the most significant change in the model's evaluation indicators is whether or not a multimodal large language model is used as an external knowledge base. In the MRE task, without a multimodal large language model, the model's F1 score is only 78.69%, a 15.10% drop compared to the complete MllmDA-KGC. Similarly, the F1 score in the MNER task also dropped by 9.05%. This shows that after the multimodal large model strengthens the embedding of the input content, the model can make good use of the external knowledge brought by the multimodal large model for accurate extraction of relationships. It can effectively utilize the external knowledge provided by the multimodal large model to accurately extract relationships and entities, effectively solving the problems of scarce on-site samples and insufficient resolution of text images.

[0193] Secondly, in the two tasks, the effect brought by the MT-Transformer and Memory_Bank mechanism is second only to the multimodal large model. As can be seen from Tables 4 and 5, for the MNER and MRE tasks, the model achieved an improvement of 5.3% and 7.46% respectively, indicating that the model can better remember and utilize previously appeared knowledge, and can effectively solve problems such as long field equipment operation and maintenance cycle, forgetting of original knowledge, and low utilization rate.

[0194] The third module, MDWA, brought the model an improvement of up to 3.49% in the MRE task and a 2.39% improvement in the MNER task, indicating that the model can obtain more hierarchical and detailed component features and effectively identify and extract entities and relationships from complex operation and maintenance data.

[0195] (6) Visualization of model effects

[0196] In order to evaluate the effectiveness of the proposed model, Figure 7 Two examples and their prediction results are presented in Figure 2, showing predicted entities and relationships related to the field of equipment operation and maintenance. In experiment (a), challenges in addressing the modality gap between text and images, particularly the incomplete alignment of visual and textual information, led to TMR extracting incorrect information. This passive misalignment led to errors in entity recognition and hindered the effective integration of MNER and MRE data. In the experiment, the entity type "tear" was misidentified as "cause," which affected the determination of entity relationships and ultimately led to incorrect predictions by the TMR model. Furthermore, the "Chassis protection plate\Failure cause" entity and its relationships were not recognized by either the TMR or EEGA models, indicating that the memory integration of image and text data was not effectively performed, resulting in poor recognition and analysis of distant entities, omission of knowledge that arises during the equipment operation and maintenance process, and ultimately missing graph nodes. In experiment (b), "deposite" was inaccurately extracted as an entity by TMR, and the "Chamber\Failure phenomenon" entity and its relationships were ignored in the prediction. EEGA and TMR also failed to predict the entity and relationship "Carbondeposit\Has attribute," demonstrating that graph construction in the complex textual environment of equipment operation and maintenance is difficult, especially when the content and images are highly correlated. In contrast, MllmDA-KGC correctly predicted the correct answer in both predictions, demonstrating the model's effectiveness and accuracy in mining multimodal knowledge in the field of equipment operation and maintenance.

[0197] (7) Application effect verification

[0198] In order to verify the superiority of the multimodal equipment operation and maintenance knowledge graph proposed in this embodiment, a set of additional comparative experiments were designed. In these experiments, the complete dataset in Section 4.1 was used for graph construction, and a multimodal knowledge graph and a unimodal knowledge graph were constructed respectively, targeting the field of equipment operation and maintenance. The two knowledge graphs generated during the experiment are partially shown in Figure 8 , and Table 6 lists the specific parameters of these two maps in detail.

[0199] Table 6 Specific parameters of the two maps

[0200]

[0201] It can be clearly seen that with the enhancement of image modality, the knowledge graph on the left can view more detailed fault problems and fault phenomena in the operation and maintenance field, including fault location, shape, etc., indicating that the information richness of the multimodal knowledge graph is significantly stronger than that of the single-modal knowledge graph on the right.

[0202] After constructing the graph, we used the TransR algorithm, a well-established algorithm in the field of knowledge graph reasoning and completion, to make inference decisions about equipment operation and maintenance based on the knowledge graph. To verify the superiority of different graphs, we calculated their MRR and Hit@1 scores during the inference process. The detailed inference results are shown in Table 7.

[0203] Table 7 Comparison of MRR and Hit@1 scores of different modal knowledge graphs

[0204]

[0205] Experimental results demonstrate that multimodal knowledge graphs achieve significantly higher performance than unimodal knowledge graphs, outperforming them by 12.21% in MRR and 8.81% in Hits@1. This demonstrates that reasoning based on multimodal knowledge graphs can more accurately identify entities and their relationships, and demonstrate greater precision and reliability in knowledge graph prediction and completion. By integrating textual and visual information, multimodal approaches have been shown to capture richer semantic details, thereby improving the model's understanding of complex data. Further validation confirms the potential value of multimodal data fusion in enhancing knowledge graph construction, advancing information extraction tasks, and improving the performance of downstream applications.

[0206] (8) Conclusion

[0207] Industrial equipment operation and maintenance is a key link in production safety and efficiency. Using knowledge graph technology for equipment operation and maintenance can achieve graph structure representation and efficient knowledge association between equipment and faults. However, traditional graph methods are limited to a single text modality and face challenges such as sample scarcity and difficulty in dynamic association modeling, resulting in missing graph entities and broken causal chains, which restrict the accuracy of operation and maintenance decisions. Therefore, this embodiment proposes a large-model-enhanced multimodal knowledge graph construction method for equipment operation and maintenance based on a multimodal large-model-enhanced two-stage attention model (MllmDA-KGC) for constructing an equipment operation and maintenance knowledge graph. MllmDA-KGC first introduces QWEN-VL into the two-stream Transformer architecture to achieve dynamic alignment of images and text and supplement few-shot semantics, improving the accuracy of entity relationship recognition for parts and problems in the operation and maintenance domain. Secondly, it proposes the MT-Transformer module, which integrates causal relationships, dilated convolutions, and a memory slot mechanism to achieve cross-modal temporal embedding fusion, improving the continuity of associations between new and old parts and the accuracy of causal chain embedding. Then, it designs a multimodal dynamic weighted attention guidance module, introducing weighted key values to guide attention points, integrating schematic and text features, and improving the accuracy of entity relationship modeling for parts and faults. Finally, to fully utilize the multimodal understanding capabilities of the multimodal large model, its labeled image embedding and position embedding are integrated into the feature embedding of RoBERTa. Combined with CRF and SOFTMAX, it performs MNER and MRE tasks, achieving the construction of a multimodal operation and maintenance graph. This example uses a car operation and maintenance dataset from a certain automobile company to verify the proposed method. The model achieves an F1 score of 88.40% and 93.79% in the MNER and MRE tasks, respectively, demonstrating its effectiveness in constructing a multimodal knowledge graph for equipment operation and maintenance. Furthermore, during graph reasoning, the multimodal graph reasoning effect is significantly better than that of a single modality, demonstrating the superiority of the multimodal graph.

[0208] The main contributions of this embodiment are as follows:

[0209] 1. We propose a multimodal embedding approach based on a pre-trained large visual language model and a dual-stream architecture. By leveraging its cross-modal semantic association capabilities, we perform fine-grained analysis of local visual features in device images and dynamically align them with textual entity semantics. By incorporating the large model's domain-learning prior knowledge, we construct a cross-modal semantic mapping of device status, enriching semantic richness with few samples and improving the recognition accuracy of entities and relationships within the O&M knowledge system.

[0210] 2. We propose a multimodal sequential convolutional attention module, MT-Transformer, which combines the local modeling capabilities of TCN with the global modeling capabilities of Transformer to perform joint embedding training on images and accompanying text generated during long-term equipment operation and maintenance. We establish an operation and maintenance knowledge memory slot mechanism, which complements semantics through cross-modal temporal feature fusion. This improves the continuous embedding of knowledge associations between new and old parts and the embedding accuracy of entity causal chains in the temporal dimension.

[0211] 3. A multimodal dynamic weighted attention module (Multimodal Dynamic WeightedAttention) is proposed, which integrates equipment schematics and text features, introduces attention weight keys, and accurately models complex coupled part entities and associated structures in long operation and maintenance texts through cross-modal co-occurrence analysis and unimodal adjacency analysis. It constructs a hierarchical equipment knowledge graph and enhances the accuracy and completeness of entity connection relationship feature modeling in image-text data.

[0212] The above embodiments are merely preferred embodiments for the purpose of fully illustrating the present invention, and the scope of protection of the present invention is not limited thereto. Equivalent substitutions or modifications made by those skilled in the art based on the present invention are within the scope of protection of the present invention. The scope of protection of the present invention shall be subject to the claims.

Claims

1. A method for constructing a large-model-enhanced multimodal knowledge graph for equipment operation and maintenance, characterized by: The steps include: Step 1: Multimodal Large Model Embedding Enhancement The pre-trained multimodal large model QWEN-VL is introduced. Using a dual-stream Transformer architecture, the input text and image data are initially embedded. This dynamically aligns cross-modal semantics, supplements few-shot semantic information, extracts cross-modal features, and mines implicit device features and complex associations by marking key image regions and text entities. Step 2: Cross-modal temporal fusion Design the MT-Transformer module, which leverages the local modeling capabilities of TCN to independently model the new and old text modal and image modal data in the operation and maintenance domain, and store them in multimodal memory slots. The attention mechanism is used to fuse the features of new and old parts, strengthening the causal chain embedding in the time dimension; Step 3: Dynamic Weight Attention Guidance: Through the multimodal dynamic weighted attention module, combined with the cross-modal co-occurrence matrix and the unimodal adjacency matrix, the attention weights of text and images are adaptively adjusted to improve the modeling accuracy of complex part relationships; Step 4: Complete the construction of the operation and maintenance map MNER task: Combine the image features of the visual Transformer fused with QWEN-VL tags and position embedding with the text features of RoBERTa; use the CRF function to calculate the probability distribution of the label sequence, output the entity label sequence, and complete multimodal named entity recognition; MRE task: splicing multimodal features, combining the Softmax function to calculate the relationship probability between specific entity pairs, screening high-confidence associations, and completing multimodal relationship extraction.

2. The method for constructing a large-model-enhanced multimodal knowledge graph for equipment operation and maintenance according to claim 1 is characterized by: In step 1, the method steps for embedding and enhancing the multimodal large model are as follows: 11) Use a two-stream architecture for data processing, using the RoBERTa model to extract semantic features from text, and ViT to extract local image features from images through the visual Transformer; 12) Introducing the multimodal large language model QWEN-VL as an external knowledge base to optimize the embedding representations of RoBERTa and ViT; 13) The embedding generated by the multimodal large model is combined with the dual-stream embedding through a weighted fusion strategy. The formula is: E fused =α·E QWEN +β·(w t ·X t +w v ·X v ·w V ) Where: E fused represents the two-dimensional embedding generated by the multimodal large model; X t and X v Represent the embeddings generated by RoBERTa and ViT respectively; α is the weight parameter of the large model embedding, which is used to control its contribution to the final fusion result; β is the weight parameter of the two-stream model embedding; w t and w v are the weight parameters of text and image features, w V It is the projection matrix corresponding to the image modality, which aims to embed the image into a two-dimensional space.

3. The method for constructing a large-model-enhanced multimodal knowledge graph for equipment operation and maintenance according to claim 1 is characterized by: In step 2, the method steps for cross-modal time series fusion are as follows: 21) Construct image modal memory slot M V and text modal memory slot M t , bimodal features are mapped to a unified key-value space through a key-value projection matrix to achieve data memory function; 22) Integrate causal convolution and dilated convolution in the memory module, pass the input sequence layer by layer through multi-level temporal convolution units, and enhance the local perception ability of the memory module: Among them: H l is the hidden layer output of the lth layer of temporal convolution; k is the convolution kernel size; d is the expansion factor; s is the current time step index; σ is the nonlinear activation function ReLU; W i is the convolution kernel weight matrix; b is the bias term; 23) A learnable expansion factor allocation strategy is used to dynamically adjust the expansion rate of each layer of temporal convolution through a gating mechanism: d l =Sigmoid(W n H l-1 +b n )·d max Where: d l is the expansion rate of the lth layer of temporal convolution; H l-1 is the hidden layer output of the l-1th layer of temporal convolution; d max is the preset maximum expansion rate, W n is the gating weight matrix, b n Gate bias 24) Through the mapping matrix of the two modalities, the image and text features are mapped to the attention key-value space to obtain the K and V values of each modality respectively; 25) For the feature knowledge of the new input, the projection matrix of the graph data part is spatially compressed and encoded, and the query vector is generated respectively with the embedding of the sequence of the text data part; the historical memory and the current input features are fused through the multi-head attention mechanism to output the cross-modal temporal embedding.

4. The method for constructing a large-model-enhanced multimodal knowledge graph for equipment operation and maintenance according to claim 3 is characterized by: In step 24), the K and V values of a single mode are obtained as follows: K fusion ×W k [K v ⊙K t ]、 Among them: K v and K t Represents the K value of image and text modality respectively; V v and V t Represents the V value of image and text mode respectively; as well as Represents the image and text modal key-value projection matrices respectively; K fusioh With V fusion is the cross-modal fusion vector; W k and W v are the weight ratios of K and V key values under the current task respectively; ⊙ represents the element-by-element product between modalities to strengthen the interaction of complementary features; Represents channel dimension concatenation to preserve modality specificity.

5. The method for constructing a large-model-enhanced multimodal knowledge graph for equipment operation and maintenance according to claim 4 is characterized by: In step 25), the method of fusing historical memory and current input features through a multi-head attention mechanism to output a cross-modal temporal embedding is: MultiHead(Q,K fusion ,V fusion )=Concat(head1,...,head h )W o Where: head represents the output of different attention heads; B represents the bias when processing the fusion vector; O represents the embedding result of the final output; Q is the query vector; h represents the number of heads of multi-head attention; d represents the total dimension of the input vector; W ° Represents the output projection matrix.

6. The method for constructing a large-model-enhanced multimodal knowledge graph for equipment operation and maintenance according to claim 1 is characterized by: In step 3, the multimodal dynamic weight attention module performs the following method steps: 31) Given the image modality embedding and text modality embedding, construct the co-occurrence matrix C based on cross-modal co-occurrence statistics co and the relative position matrix R within the modal; 32) Define the visual head and text head attention mechanism, respectively for the visual head and text head: in: and They are text header and visual header respectively; and is the linear transformation matrix of query, key and value; R t and R v are the relative position matrices within the text modality embedding and image modality respectively; x t and x v They are text modality embedding and image modality embedding respectively; 33) Through residual connection, the original feature data of a single modality is imported to guide the dynamic weight key value to obtain the predicted classification 34) Introducing a learnable external parameter dynamic weight key, by classifying the prediction Gradient descent is performed with the data label through the cross entropy function, and the value vector in the attention is interactively weighted to optimize the entity relationship modeling.

7. The method for constructing a large-model-enhanced multimodal knowledge graph for equipment operation and maintenance according to claim 6 is characterized by: In step 33), the predicted classification for: H fusion =H' V +H' T , Where: H' T is the text feature guided by the dynamic weight key value; H' V is a visual feature guided by dynamic weight key value; H fusion is the image-text fusion feature; W clf is the weight matrix of the classifier; In step 34), calculate L CE For K learnabel Gradient: Where: L CE is the loss function; N is the total number of samples, C is the number of categories, y ij is the true label of the i-th sample belonging to the j-th class; η is the learning rate; K learnable are learnable external parameters; The interactive modality weighting method for the value vector in the attention is: Where: H is the multimodal bias of the V vector; head* is the updated attention head.

8. The method for constructing a large-model-enhanced multimodal knowledge graph for equipment operation and maintenance according to claim 1 is characterized by: Applications of the multimodal large model QWEN-VL include: Visually mark the position and connection status of parts in the equipment schematic; Parse the entities and relations in the fault log text and generate triple prompts; The labeled visual embedding and position embedding are integrated into the text features of RoBERTa, and CRF and Softmax are combined to complete the MNER and MRE tasks.

9. A large-model enhanced multimodal knowledge graph construction system for equipment operation and maintenance, used to execute the equipment operation and maintenance multimodal knowledge graph construction method according to any one of claims 1 to 8, characterized in that: include: A multimodal data input module for receiving text and image data; QWEN-VL large model interface for cross-modal semantic alignment and few-shot enhancement; MT-Transformer temporal fusion module, used for associative modeling of new and old parts; MDWAM attention guidance module for complex relationship analysis; The knowledge graph output module is used to generate structured equipment operation and maintenance graphs.

10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for constructing a large-model enhanced equipment operation and maintenance multimodal knowledge graph as described in any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Operation and maintenance method, device and equipment of present network system, medium and program product

    CN120875851A

  • Network system operation and maintenance method, device, equipment, medium and program product

    CN120875851B

  • Artificial intelligence-based three-dimensional model structure semantic extraction method and system

    CN121147903A

  • Satellite-ground system operation and maintenance method and system based on cooperation of large and small models

    CN121660081A