Multimodal semantic model training method and system based on knowledge enhancement
By building a multimodal data fusion module and a knowledge graph embedding layer, combining knowledge-guided attention mechanism and multi-stage pre-training, the knowledge fusion and noise optimization problems of multimodal models in complex tasks are solved, and the model's understanding and reasoning ability is improved.
Patent Information
- Application Number
- CN202510549227.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-12
AI Technical Summary
The existing multimodal models lack explicit knowledge guidance and are difficult to deal with complex tasks. The learning of single knowledge graph representations has not been effectively integrated with multimodal representations, and the problems of multimodal alignment and knowledge noise optimization have not been solved by the system, affecting the generalization ability and robustness of the model.
Build a multimodal data fusion module, establish a knowledge graph embedding layer, perform feature fusion through a knowledge-guided attention mechanism, and use a multi-stage pre-training strategy to optimize model performance, including cross-modal comparison learning, knowledge adversarial training and logical reasoning verification.
It improves the model's understanding and reasoning ability in complex scenarios, enhances the integration of multimodal information and knowledge graphs, and improves the generalization ability and robustness of the model.
Smart Images

Figure CN120471059A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multimodal semantic model training method and system based on knowledge enhancement, belonging to the technical field of artificial intelligence. Background Art
[0002] With the rapid development of artificial intelligence technology, multimodal learning has become a research hotspot. It achieves more comprehensive information understanding and processing by integrating different modal data such as text, images, and audio.
[0003] However, existing technologies still have the following shortcomings: First, although traditional multimodal models (such as CLIP, Florence, etc.) can process multimodal data, they lack explicit knowledge guidance and are difficult to handle complex tasks that require domain common sense and background knowledge. Second, existing knowledge graph embedding methods (such as TransE, RotatE, etc.) mainly focus on the representation learning of a single knowledge graph and fail to effectively integrate with multimodal representations, resulting in the model being unable to fully utilize the structured information in the knowledge graph to enhance multimodal understanding. Third, existing pre-training strategies often separate multimodal alignment from knowledge fusion, and fail to systematically solve the coupled optimization problem of knowledge noise injection and modal alignment, affecting the generalization ability and robustness of the model.
[0004] Therefore, there is an urgent need for a method that can effectively integrate multimodal information and knowledge graphs and optimize model performance through multi-stage pre-training strategies to improve the model's understanding and reasoning capabilities in complex scenarios. Summary of the Invention
[0005] In order to solve the above problems in the prior art, the present invention proposes a multimodal semantic model training method and system based on knowledge enhancement.
[0006] The technical solutions of the present invention are as follows:
[0007] In one aspect, the present invention proposes a multimodal semantic model training method based on knowledge enhancement, comprising the following steps:
[0008] Construct a multimodal data fusion module to receive and preprocess multimodal heterogeneous data to form a multimodal feature vector;
[0009] Establish a knowledge graph embedding layer, obtain knowledge graphs corresponding to different types of knowledge bases, and extract knowledge vectors corresponding to entities in each knowledge graph;
[0010] Establish a knowledge-guided attention mechanism to fuse the knowledge vectors of each entity with the multimodal feature vectors to generate a fused feature vector;
[0011] The target semantic model is trained based on the fused feature vector.
[0012] As a preferred embodiment, the multimodal heterogeneous data includes:
[0013] Text modality data, image modality data, and voice modality data;
[0014] The step of receiving multimodal heterogeneous data and performing preprocessing to form a multimodal feature vector includes:
[0015] Perform feature extraction on text modality data, image modality data, and voice modality data respectively to obtain text feature vectors, image feature vectors, and voice feature vectors;
[0016] Project text feature vectors, image feature vectors, and speech feature vectors into a unified semantic space through a learnable orthogonal mapping matrix;
[0017] The projected text feature vector, image feature vector and speech feature vector are concatenated in the semantic space to generate a multimodal feature vector.
[0018] As a preferred embodiment, the method for extracting the knowledge vector corresponding to the entity in each knowledge graph is specifically as follows:
[0019] For any knowledge graph, extract all entities and assign an initialization vector to each entity, and assign a corresponding relationship weight matrix to each relationship between entities;
[0020] Each entity is used as a node and input into the graph neural network for feature encoding. The relational attention weight is calculated based on the relational weight matrix and the neighboring nodes of each entity, and the knowledge vector of each entity is encoded through the relational attention weight.
[0021] As a preferred embodiment, the steps of establishing a knowledge-guided attention mechanism and fusing the knowledge vectors of each entity with the multimodal feature vector to generate a fused feature vector are specifically as follows:
[0022] Establish a multi-head attention mechanism, where each attention head corresponds to a knowledge graph;
[0023] By calculating the semantic similarity between the multimodal features and the knowledge vector of each entity, highly relevant entities are screened based on the semantic similarity;
[0024] Attention calculation is performed to obtain interaction weights based on the relationship paths between the selected highly relevant entities and the multimodal feature vectors in the knowledge graph. In the attention calculation, the multimodal feature vector is used as the query vector, and the knowledge vector of the highly relevant entity is used as the value vector and key vector.
[0025] The corresponding highly correlated entities and multimodal feature vectors are weightedly fused through interaction weights to form a fused feature vector.
[0026] As a preferred embodiment, after the step of training the target semantic model based on the fused feature vector, the knowledge vector of each entity is reversely adjusted by the training task loss of the target semantic model.
[0027] As a preferred embodiment, the process of training the target semantic model based on the fused feature vector includes:
[0028] Cross-modal comparative training phase, knowledge adversarial training phase, and logical reasoning verification phase.
[0029] On the other hand, the present invention also proposes a multimodal semantic model training system based on knowledge enhancement, comprising:
[0030] Multimodal data fusion module, used to receive multimodal heterogeneous data and perform preprocessing to form multimodal feature vectors;
[0031] The knowledge extraction module is used to establish a knowledge graph embedding layer, obtain knowledge graphs corresponding to different types of knowledge bases, and extract knowledge vectors corresponding to entities in each knowledge graph;
[0032] The knowledge injection module is used to establish an attention mechanism based on knowledge guidance, and fuse the knowledge vector of each entity with the multimodal feature vector to generate a fused feature vector;
[0033] The model training module trains the target semantic model based on the fused feature vector.
[0034] On the other hand, the present invention also proposes an electronic device, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the program, it implements the knowledge-enhanced multimodal semantic model training method as described in any embodiment of the present invention.
[0035] On the other hand, the present invention also proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the knowledge-enhanced multimodal semantic model training method as described in any embodiment of the present invention.
[0036] In another aspect, the present invention further provides a method for diagnosing a fault in an electric power device, comprising the following steps:
[0037] The multi-dimensional monitoring data of the target power equipment is collected as multi-modal heterogeneous data through a sensor array arranged on the target power equipment;
[0038] Acquire knowledge information related to the target power equipment and construct a knowledge graph of the target power equipment;
[0039] Based on the multimodal heterogeneous data and the knowledge graph of the target power equipment, the target power equipment fault diagnosis semantic model is trained using the knowledge-enhanced multimodal semantic model training method described in any embodiment of the present invention;
[0040] Fault diagnosis of target power equipment is performed through the trained target power equipment fault diagnosis semantic model and the multimodal heterogeneous data of the target power equipment collected in real time.
[0041] Additional aspects and advantages of the present invention will be set forth in the following description, and some of them will be obvious from the description, or may be learned by practicing the present invention. In addition, the various aspects and advantages of the present invention may be realized and obtained by the method steps and combinations particularly pointed out in the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 This is a flow chart of the method according to the first embodiment of the present invention. DETAILED DESCRIPTION
[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0044] It should be understood that the step numbers used herein are only for convenience of description and are not intended to limit the order in which the steps are to be executed.
[0045] It should be understood that the terms used in the present specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, the singular forms "a", "an" and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0046] The terms “include” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0047] The term "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items.
[0048] Example 1:
[0049] See also Figure 1This embodiment proposes a multimodal semantic model training method based on knowledge enhancement, including the following steps:
[0050] S100: Construct a multimodal data fusion module to receive and preprocess multimodal heterogeneous data to form a multimodal feature vector.
[0051] Specifically, the multimodal data fusion module receives heterogeneous data in multiple modalities, including but not limited to text, images, audio, and video. For text data, the BERT pre-trained model is used for encoding, converting the input text sequence into a 768-dimensional text feature vector. For image data, a ResNet-50 network is used to extract a 2048-dimensional visual feature vector. For audio data, a VGGish model is used to extract a 128-dimensional audio feature vector. For video data, a 3D convolutional neural network and a temporal modeling network are combined to extract a 1024-dimensional video feature vector.
[0052] In order to unify the dimensions of different modal features, a feature projection layer is designed to map each modal feature into a unified 512-dimensional feature space. The projection process uses a combination of linear transformation and nonlinear activation function. The specific formula is:
[0053] Fm=ReLU(Wm·Xm+bm)
[0054] Among them, Xm represents the original modal features, Wm and bm are the weight matrix and bias vector of the projection layer respectively, Fm is the modal feature vector after projection, and ReLU is the activation function.
[0055] Furthermore, to address the issue of missing data for different modalities, a missing modality completion mechanism was designed. When data for a particular modality is missing, an approximate representation of the missing modality is generated through a linear combination of features from other modalities, ensuring that the model can still operate effectively despite incomplete data.
[0056] S200. Establish a knowledge graph embedding layer, obtain knowledge graphs corresponding to different types of knowledge bases, and input each knowledge graph into the graph neural network to generate the corresponding knowledge vector.
[0057] Specifically, the knowledge graph embedding layer first extracts structured knowledge from multiple domain knowledge bases, including general knowledge bases (such as Wikidata and ConceptNet) and domain-specific knowledge bases. Each knowledge graph consists of a set of entities E, a set of relations R, and a set of triples T = {(h, r, t) | h, t∈E, r∈R}, where h represents the head entity, r represents the relation, and t represents the tail entity.
[0058] Graph neural networks are used to dynamically embed knowledge graphs, capturing high-level semantic associations through a relational path attention mechanism. Specifically, a relation-aware graph attention network (GRCN) is used to process knowledge graphs, which can distinguish different types of relations and capture complex interactions between entities.
[0059] S300, establishing a knowledge-guided attention mechanism to fuse the knowledge vectors of each entity with the multimodal feature vector to generate a fused feature vector;
[0060] S400: Training a target semantic model based on the fused feature vector.
[0061] The training implements a multi-stage pre-training strategy.
[0062] The first stage performs cross-modal contrastive learning and optimizes inter-modal consistency through negative sampling:
[0063] In this stage, cross-modal contrastive learning is performed. For each N sample in a training batch, the different modal features of each sample are considered as positive sample pairs, while the features of other samples in the same batch are considered as negative samples. For the modal features Fi1 and Fi2 of sample i, the contrastive learning loss is calculated as follows:
[0064]
[0065] Among them, sim() represents the similarity calculation function, τ is the temperature coefficient, and the temperature coefficient is a hyperparameter in the contrast loss function, and its role is analogous to the Boltzmann distribution in statistical mechanics.
[0066] For each N sample in a training batch, the different modal features of each sample are considered as positive sample pairs, while the features of other samples in the same batch are considered as negative samples. To enhance contrastive learning, a hard example mining strategy is adopted, dynamically selecting semantically similar but different class samples as hard negative examples, thereby improving the model's discriminative ability. Specifically, in each training batch, the similarity of all negative sample pairs is calculated, and the top 20% with the highest similarity are selected as hard negative examples, and their weight in the loss function is increased.
[0067] The second stage uses knowledge adversarial training to generate perturbation samples to improve the robustness of the model.
[0068] At this stage, the robustness of the model is enhanced through adversarial training. In the specific implementation, the perturbation in the gradient direction is added to the input features to generate adversarial samples F adv :
[0069]
[0070] Among them, F is the original feature, ε is the perturbation size, which is set to 0.01. is the gradient of the loss function with respect to the input feature, sign(·) is the sign function, y is the true label, and L() is the loss function of knowledge adversarial training.
[0071] The adversarial training loss function is defined as the weighted sum of the original sample loss and the adversarial sample loss:
[0072] L adv =L(F,y)+λ adv ·L(F adv ,y)
[0073] Among them, λ adv To counter the loss weight, it is set to 0.5.
[0074] The third stage introduces knowledge reasoning tasks to verify the model's ability to understand implicit logical relationships.
[0075] In this phase, a knowledge reasoning task is designed to require the model to predict missing relations or entities in the knowledge graph. Specifically, a portion of triples (h, r, t) are randomly masked from the knowledge graph and the model is required to recover these triples based on contextual information.
[0076] The loss function of the knowledge reasoning task adopts cross entropy loss:
[0077]
[0078] Where T' is a set of triples, P(t|h,r) represents the probability of the model predicting the tail entity t given the head entity h and the relation r, and P(h|r,t) represents the probability of the model predicting the head entity h given the relation r and the tail entity t.
[0079] The loss function combination of multi-stage training is:
[0080] L tOTAL =λ1L CL +λ2L Adv +λ3L KR
[0081] Among them, L CL is the contrastive learning loss, L Adv is the adversarial training loss, L KR is the cross entropy loss for knowledge reasoning. λ1, λ2, and λ3 are weight coefficients for each loss term. The weight coefficients can be pre-examined and are set to 1.0, 0.5, and 0.3, respectively, in this embodiment.
[0082] During the training process, a dynamic weight adjustment strategy is adopted to adaptively adjust the weights of each loss term based on the performance of the validation set to ensure the stability and effectiveness of model training. Specifically, after every 5 epochs (rounds) of training, the model performance is evaluated on the validation set, and the weight coefficients are adjusted according to the performance changes:
[0083] λ i =λ i (1+a·ΔPerfi)
[0084] Where ΔPerfi represents the relative change in the performance of the i-th training task compared to the previous evaluation, and a is the adjustment coefficient, which is set to 0.1 in this embodiment.
[0085] In addition, to prevent overfitting, an early stopping strategy and a learning rate decay mechanism are introduced. If the validation set performance does not improve for 5 consecutive epochs, the learning rate is halved; if there is no improvement for 10 consecutive epochs, training is stopped.
[0086] As a preferred implementation of this embodiment, the multimodal heterogeneous data includes:
[0087] Text modality data, image modality data, and voice modality data;
[0088] The step of receiving multimodal heterogeneous data and performing preprocessing to form a multimodal feature vector includes:
[0089] Perform feature extraction on text modality data, image modality data, and voice modality data respectively to obtain text feature vectors, image feature vectors, and voice feature vectors;
[0090] Project text feature vectors, image feature vectors, and speech feature vectors into a unified semantic space through a learnable orthogonal mapping matrix;
[0091] The projected text feature vector, image feature vector and speech feature vector are concatenated in the semantic space to generate a multimodal feature vector.
[0092] As a preferred implementation of this embodiment, the method for extracting the knowledge vector corresponding to the entity in each knowledge graph is specifically as follows:
[0093] For any knowledge graph, extract all entities and assign an initialization vector to each entity, and assign a corresponding relationship weight matrix to each relationship between entities;
[0094] Each entity is used as a node and input into the graph neural network for feature encoding. The relational attention weight is calculated based on the relational weight matrix and the neighboring nodes of each entity, and the knowledge vector of each entity is encoded through the relational attention weight.
[0095] Specifically, in the relation-aware graph attention network, the entity vector update formula is:
[0096]
[0097] in, represents the entity vector of entity e in the l+1th layer, σ represents the nonlinear activation function, r is the index of the relationship between entities, R is the set of all relationships between entities, ′∈N(e) represents that entity ′ is the set of neighbor nodes belonging to entity e, ,′ is the relationship path attention coefficient, W r is the relationship weight matrix corresponding to relationship r, h l ′ Represents the entity vector of node ′ at layer l.
[0098] Among them, the calculation formula of ,′ is:
[0099]
[0100] Among them, h e and h′ represent the original entity vectors of node e and node ′, is the feature embedding vector corresponding to the relation r, For splicing operation.
[0101] Based on the above solution, this embodiment introduces a multi-hop reasoning mechanism to enhance the expressive power of the knowledge graph. By sampling relationship paths of varying lengths, it captures long-range dependencies between entities. Specifically, for each entity e, relationship paths of a set length are randomly sampled, and the entity information along these paths is aggregated to form a more comprehensive entity representation.
[0102] Ultimately, the knowledge vector output by the knowledge graph embedding layer contains the semantic and structural information of the entity.
[0103] As a preferred implementation of this embodiment, the steps of establishing a knowledge-guided attention mechanism and fusing the knowledge vectors of each entity with the multimodal feature vector to generate a fused feature vector are specifically as follows:
[0104] Establish a multi-head attention mechanism, where each attention head corresponds to a knowledge graph;
[0105] By calculating the semantic similarity between the multimodal features and the knowledge vector of each entity, highly relevant entities are screened based on the semantic similarity;
[0106] The relationship paths corresponding to the multimodal feature vectors in the knowledge graph through the filtered highly relevant entities are as follows: patchy shadows → imaging manifestations → pneumonia → common causes → bacterial infection.
[0107] Perform attention calculation to obtain interaction weights. In the attention calculation, the multimodal feature vector is used as the query vector, and the knowledge vector of the highly relevant entity is used as the value vector and key vector. Specifically:
[0108]
[0109] Where Q is the linear transformation of modal characteristics, K e / V is the knowledge vector projection, d k To enhance the expressive power of the attention mechanism, relative position encoding is introduced:
[0110]
[0111] Among them, M is the relative position bias matrix, which is used to capture the positional relationship between the knowledge vector and the modal feature.
[0112] In addition, to handle the heterogeneity between different modalities and knowledge, modality-specific attention heads are designed:
[0113] MultiModalAttn(Fm,Ke)
[0114] =Concat(Attn1(Fm1,Ke),...,AttnM(FmM,Ke))W o
[0115] Among them, Fmi represents the characteristics of the i-th mode, M is the number of modes, and W o is the output projection matrix.
[0116] The corresponding highly correlated entities and multimodal feature vectors are weightedly fused through interaction weights to form a fused feature vector.
[0117] The above scheme realizes dynamic weighted knowledge injection, dynamically adjusts the knowledge dependency strength according to the input data, and can identify the knowledge entities most relevant to the current modal features, thereby improving the efficiency of model training.
[0118] As a preferred implementation of this embodiment, after the step of training the target semantic model based on the fused feature vector, the knowledge vector of each entity is reversely adjusted through the training task loss of the target semantic model.
[0119] Example 2:
[0120] This embodiment proposes a multimodal semantic model training system based on knowledge enhancement, including:
[0121] A multimodal data fusion module, configured to receive and preprocess multimodal heterogeneous data to form a multimodal feature vector; this module is configured to implement the function of step S100 in the first embodiment and will not be described in detail here;
[0122] A knowledge extraction module is used to establish a knowledge graph embedding layer, obtain knowledge graphs corresponding to different types of knowledge bases, and extract knowledge vectors corresponding to entities in each knowledge graph. This module is used to implement the function of step S200 in Example 1 and will not be repeated here.
[0123] A knowledge injection module is used to establish an attention mechanism based on knowledge guidance, and to perform feature fusion on the knowledge vector of each entity and the multimodal feature vector to generate a fused feature vector. This module is used to implement the function of step S300 in the first embodiment and will not be described in detail here.
[0124] The model training module trains the target semantic model based on the fused feature vector; this module is used to implement the function of step S400 in the first embodiment and will not be described in detail here.
[0125] Example 3:
[0126] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method described in any embodiment of the present invention is implemented.
[0127] Example 4:
[0128] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the method described in any embodiment of the present invention is implemented.
[0129] In the embodiments of the present application, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. Among them, A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can represent: a, b, c, a and b, a and c, b and c or a and b and c, where a, b, c can be single or multiple.
[0130] Those skilled in the art will appreciate that the various units and algorithm steps described in the embodiments disclosed herein can be implemented using a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0131] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0132] In the several embodiments provided in this application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of this application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory; hereinafter referred to as: ROM), random access memory (Random Access Memory; hereinafter referred to as: RAM), magnetic disk or optical disk, and other media that can store program code.
[0133] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention's description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A multimodal semantic model training method based on knowledge enhancement, characterized in that: The following steps are involved: Construct a multimodal data fusion module to receive and preprocess multimodal heterogeneous data to form a multimodal feature vector; Establish a knowledge graph embedding layer, obtain knowledge graphs corresponding to different types of knowledge bases, and extract knowledge vectors corresponding to entities in each knowledge graph; Establish a knowledge-guided attention mechanism to fuse the knowledge vectors of each entity with the multimodal feature vectors to generate a fused feature vector; The target semantic model is trained based on the fused feature vector.
2. The multimodal model training method based on knowledge enhancement according to claim 1, characterized in that: The multimodal heterogeneous data includes: Text modality data, image modality data, and voice modality data; The step of receiving multimodal heterogeneous data and performing preprocessing to form a multimodal feature vector includes: Perform feature extraction on text modality data, image modality data, and voice modality data respectively to obtain text feature vectors, image feature vectors, and voice feature vectors; Project text feature vectors, image feature vectors, and speech feature vectors into a unified semantic space through a learnable orthogonal mapping matrix; The projected text feature vector, image feature vector and speech feature vector are concatenated in the semantic space to generate a multimodal feature vector.
3. The multimodal model training method based on knowledge enhancement according to claim 1, characterized in that: The method for extracting the knowledge vector corresponding to the entity in each knowledge graph is specifically as follows: For any knowledge graph, extract all entities and assign an initialization vector to each entity, and assign a corresponding relationship weight matrix to each relationship between entities; Each entity is used as a node and input into the graph neural network for feature encoding. The relational attention weight is calculated based on the relational weight matrix and the neighboring nodes of each entity, and the knowledge vector of each entity is encoded through the relational attention weight.
4. The multimodal model training method based on knowledge enhancement according to claim 1, characterized in that: The steps of establishing a knowledge-guided attention mechanism and fusing the knowledge vectors of each entity with the multimodal feature vector to generate a fused feature vector are as follows: Establish a multi-head attention mechanism, where each attention head corresponds to a knowledge graph; By calculating the semantic similarity between the multimodal features and the knowledge vector of each entity, highly relevant entities are screened based on the semantic similarity; Attention calculation is performed to obtain interaction weights based on the relationship paths between the selected highly relevant entities and the multimodal feature vectors in the knowledge graph. In the attention calculation, the multimodal feature vector is used as the query vector, and the knowledge vector of the highly relevant entity is used as the value vector and key vector. The corresponding highly correlated entities and multimodal feature vectors are weightedly fused through interaction weights to form a fused feature vector.
5. The multimodal model training method based on knowledge enhancement according to claim 4 is characterized in that: After the target semantic model is trained based on the fused feature vector, the knowledge vector of each entity is reversely adjusted through the training task loss of the target semantic model.
6. The multimodal model training method based on knowledge enhancement according to claim 1, characterized in that: The training process of the target semantic model based on the fused feature vector includes: Cross-modal comparative training phase, knowledge adversarial training phase, and logical reasoning verification phase.
7. A multimodal semantic model training system based on knowledge enhancement, characterized in that: include: Multimodal data fusion module, used to receive multimodal heterogeneous data and perform preprocessing to form multimodal feature vectors; The knowledge extraction module is used to establish a knowledge graph embedding layer, obtain knowledge graphs corresponding to different types of knowledge bases, and extract knowledge vectors corresponding to entities in each knowledge graph; The knowledge injection module is used to establish an attention mechanism based on knowledge guidance, and fuse the knowledge vector of each entity with the multimodal feature vector to generate a fused feature vector; The model training module trains the target semantic model based on the fused feature vector.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the multimodal semantic model training method based on knowledge enhancement is implemented as described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the multimodal semantic model training method based on knowledge enhancement as described in any one of claims 1 to 6 is implemented.
10. A method for diagnosing faults in electric power equipment, characterized in that: The following steps are involved: The multi-dimensional monitoring data of the target power equipment is collected as multi-modal heterogeneous data through a sensor array arranged on the target power equipment; Acquire knowledge information related to the target power equipment and construct a knowledge graph of the target power equipment; Based on the multimodal heterogeneous data and the knowledge graph of the target power equipment, the target power equipment fault diagnosis semantic model is trained using the knowledge-enhanced multimodal semantic model training method according to any one of claims 1 to 6; Fault diagnosis of target power equipment is performed through the trained target power equipment fault diagnosis semantic model and the multimodal heterogeneous data of the target power equipment collected in real time.
Citation Information
Cited By
Multi-modal data semantic alignment method and device based on cross-modal attention mechanism
CN120724398A
Multi-modal semantic understanding method and device based on multi-order progressive alignment, computer equipment and storage medium
CN120744143A
Landslide susceptibility prediction method based on knowledge graph and spatial-temporal feature fusion
CN120929974A
Domain knowledge enhancement method and device for infrastructure defect detection large model
CN121031760A
Domain knowledge enhancement methods and devices for large-scale infrastructure defect detection models
CN121031760B