Diagnosis method based on cross-modal charging pile maintenance knowledge base
By building a cross-modal charging pile maintenance knowledge base, using image and text encoders to generate high-dimensional semantic coding, and combining low-dimensional graph embedding and language model, the modal complementarity and hardware cost problems of the charging pile maintenance knowledge base are solved, and rapid and specific diagnostic solutions are achieved.
Patent Information
- Application Number
- CN202510333438.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-08
AI Technical Summary
充电桩维护知识复杂且零散,现有知识库检索生成技术存在模态互补困难和高硬件成本的问题,尤其在移动场景中难以有效生成具体诊断方案。
A cross-modal charging pile maintenance knowledge base is built, and a attention-based image encoder, text encoder and graph neural network are used to integrate visual and text modal information to generate high-dimensional cross-modal semantic coding, and search and match through low-dimensional graph embedding, and a diagnostic solution is generated based on language models.
It realizes the rapid generation of rich and specific charging pile diagnostic solutions at low hardware costs, which are suitable for mobile scenarios, reducing hardware deployment costs and improving retrieval efficiency.
Smart Images

Figure CN120277202A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of charging pile fault diagnosis, and in particular to a diagnosis method based on a cross-modal charging pile maintenance knowledge base. Background Art
[0002] With the long-term outdoor operation of charging facilities, the failure rate of charging facilities shows an increasing trend year by year during operation. The charging pile maintenance knowledge includes a large amount of maintenance operations, software instructions, technical drawings, and fault troubleshooting step procedures. Moreover, the charging pile design has differences, and each manufacturer only has its own independent problem-solving solutions. The technical materials are relatively scattered, which further increases the complexity of the charging pile maintenance knowledge. If maintained manually without the help of tools, it requires high skills for operation and maintenance personnel. Especially for newly recruited employees, they need to master a large amount of technical materials in a short time. Currently, in response to this situation, a charging pile maintenance knowledge base is often constructed to solve the problem. If it is a simple text index, not much additional work is required, and the database itself can provide good support. However, in order to understand the maintenance process more vividly, maintenance content based on a combination of text and pictures or videos is often provided. The introduction of modalities, especially images and videos, brings further difficulties to diagnostic analysis. In addition, simple retrieval based on the knowledge base also has the problem of being unable to handle the complex relationship between problems and solutions. For example, a single problem may retrieve multiple solutions, and many of the results are difficult to directly apply and require maintenance personnel to distinguish and process. With the development of language models, the technology of enhanced retrieval generation combined with language models for knowledge retrieval and generation has gradually developed. There is a method that uses a knowledge graph to cooperate with the reasoning and analysis ability of a language model for enhanced retrieval generation. However, for the mobile scenario of charging pile maintenance, the cost of deploying the language model needs to be low enough, that is, the number of model parameters should be small. However, once the number of model parameters is small and there is a lack of sufficient knowledge, although the content inferred according to the knowledge graph has sufficient relevance, the content itself is often very general, which affects the maintenance guidance.
[0003] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present invention provides a diagnosis method based on a cross-modal charging pile maintenance knowledge base.
[0005] The present invention provides a diagnosis method based on a cross-modal charging pile maintenance knowledge base, including:
[0006] Use an attention-based image encoder, text encoder, and graph neural network to process cross-modal charging pile maintenance knowledge to build a cross-modal charging pile maintenance knowledge base. The vectorized cross-modal charging pile maintenance knowledge in the cross-modal charging pile maintenance knowledge base includes high-dimensional cross-modal semantic encoding and low-dimensional graph embeddings;
[0007] When performing diagnostic processing using a charging pile diagnosis query, first extract the embeddings of entities and relationships in the diagnosis query through a low-dimensional embedding model, and use the similarity between the embeddings of entities and relationships in the diagnosis query and the graph embeddings to match the closest set number of target graph embeddings;
[0008] Obtain the cross-modal semantic encoding that matches the target graph embedding;
[0009] The language model uses the retrieved cross-modal semantic encoding and the diagnosis query to generate a diagnostic processing plan.
[0010] Furthermore, the use of an attention-based image encoder and text encoder, and a graph neural network to process cross-modal charging pile maintenance knowledge to build a cross-modal charging pile maintenance knowledge base includes:
[0011] Sort out cross-modal charging pile maintenance knowledge in the form of pictures, texts, and videos;
[0012] Organize the cross-modal charging pile maintenance knowledge into visually and textually consistent forms;
[0013] Use an attention-based image encoder and text encoder to respectively extract semantic features from the original information of the visually and textually consistent modalities and map them into high-dimensional cross-modal semantic encodings;
[0014] Use an attention-based graph neural network to extract low-dimensional graph embeddings of the knowledge graph of entities involved in charging pile maintenance knowledge, and associate the same charging pile maintenance knowledge graph embedding with its cross-modal semantic encoding.
[0015] Furthermore, for the attention parts involved in the attention-based image encoder, text encoder, and graph neural network, introduce a signal-to-noise ratio adaptive coefficient γ for noise robustness optimization, as follows:
[0016]
[0017] q,k i ,k j ,v i are the query vector, the i-th key vector, the j-th key vector, and the i-th value vector respectively. γ is the signal-to-noise ratio adaptive coefficient. The greater the noise, the smaller the γ value, and the smaller the noise, the greater the γ value.
[0018] Furthermore, the attention-based image encoder adopts a specified layer of the Vision Transformer model, followed by a first mapping fully-connected layer; the attention-based text encoder adopts a specified layer of the language model of the Transformer architecture, followed by a second mapping fully-connected layer; the first mapping fully-connected layer and the second mapping fully-connected layer output semantic features with the same dimension;
[0019] The extracted cross-modal semantic features are mapped into cross-modal semantic encodings through semantic fusion.
[0020] Furthermore, the attention-based image encoder and text encoder are constrained by an alignment loss to achieve fine-grained alignment of semantic features. The formula of the alignment loss is as follows:
[0021]
[0022] L g (VisionEn(v i ),TextEn(t i )) represents the global image-text contrast loss obtained by analyzing the visual modality charging pile maintenance knowledge v i and the text modality charging pile maintenance knowledge t i through the encoders; L Region (r j,v ,r j,t ) is the matching loss of the corresponding words r i in the regional semantics r j,v of the visual modality charging pile maintenance knowledge v i and the text modality charging pile maintenance knowledge t j,t . λ1 and λ2 are the dynamic weights of the global image-text contrast loss and the local regional word pair matching loss respectively. r j,v and r j,t are randomly sampled.
[0023] Furthermore, the ways of the semantic fusion include:
[0024] Pool and concatenate the semantic features of the visual modality and the text modality, and then fuse them through a fusion multi-layer perceptron containing a non-linear activation function;
[0025] Perform cross-attention on the semantic features of the visual modality and the text modality after pooling, and then fuse them through a fusion multi-layer perceptron containing a non-linear activation function;
[0026] Perform gated adaptive weighting on the semantic features of the visual modality and the text modality after pooling, and then fuse them through a fusion multi-layer perceptron containing a non-linear activation function.
[0027] Furthermore, extracting the low-dimensional graph embeddings of the knowledge graph of entities involved in the charging pile maintenance knowledge using an attention-based graph neural network includes:
[0028] Extract the entities and the relationships between entities in the charging pile maintenance knowledge, use the entities as nodes, and use the relationships between entities as edges to construct a knowledge graph;
[0029] For any node and edge in the knowledge graph, use a low-dimensional embedding model to extract low-dimensional embedding features respectively;
[0030] The graph neural network aggregates the information from neighbor entities for an entity to obtain a graph embedding according to the adjacency relationship determined by the knowledge graph, where the information includes the entity relationship with the neighbor entity and the low-dimensional embedding features of the neighbor entity.
[0031] Furthermore, in the graphic and text-based charging pile maintenance knowledge, the original information of the visual modality and the text modality appears in pairs, and the semantics between different modalities are consistent;
[0032] In the video-based charging pile maintenance knowledge, the original information of the visual modality and the speech modality appears aligned along the time axis. Processing the video-based charging pile maintenance knowledge into a unified form with the graphic and text-based knowledge includes: converting the speech of the video-based charging pile maintenance knowledge into temporally aligned subtitles, that is, extracting the original information of the text modality from the video; grouping the frames of the video-based charging pile maintenance knowledge according to the semantic relationship of the subtitles and the time axis of the subtitles to obtain a sub-frame sequence, and sampling representative frames from the sub-frame sequence for each frame in the sub-frame sequence, and the paired sampling frames and subtitle semantics are consistent.
[0033] Furthermore, generating a diagnostic processing plan using the retrieved cross-modal semantic encoding and diagnostic query includes: embedding the cross-modal semantic encoding and diagnostic query as prompt words of a language model, and controlling the language model to generate a diagnostic processing plan based on the high-dimensional cross-modal semantic encoding.
[0034] Furthermore, in order to enable the language model to generate a diagnostic processing plan based on the cross-modal semantic encoding, a mapping layer for the cross-modal semantic encoding is provided for the language model, and the mapping layer is trained to generate an encoding supported by the language model from the cross-modal semantic encoding.
[0035] The above technical solutions provided by the embodiments of the present invention have the following advantages compared with the prior art:
[0036] This application uses an attention-based image encoder and text encoder to integrate the global information of the visual modality and the keyword information of the text to form a cross-modal semantic encoding, realizing cross-modal complementarity. The cross-modal complementarity provides richer information compared with the graph embedding, and can effectively provide richer features in the task of generating a diagnostic processing plan.
[0037] This application constructs a low-dimensional graph embedding for retrieval using an attention-based graph neural network. When performing knowledge retrieval based on a diagnostic query, retrieval matching is carried out through the low-dimensional diagnostic query and the graph embedding, with a small computational amount and high speed, being suitable for low hardware costs.
[0038] In this application, a language model that supports processing cross-modal semantic encoding uses the retrieved cross-modal semantic encoding and the diagnostic query to generate a diagnostic processing plan. When generating the diagnostic processing plan, the language model actually mainly refers to the cross-modal semantic encoding to generate content, avoiding the situation where the language model simply uses graph class embeddings to generate highly generalized content or generates hard-to-understand content by supplementing the generalized graph class embedding framework based on its own hallucinations. Even a language model with a small number of parameters can achieve good diagnostic effects, reducing the deployment hardware cost and being suitable for mobile scenarios of charging pile maintenance. Brief Description of the Drawings
[0039] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.
[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0041] Figure 1 It is a flowchart of a diagnostic method based on a cross-modal charging pile maintenance knowledge base provided by an open embodiment of the present invention.
[0042] Figure 2 It is a flowchart of using an attention-based image encoder, a text encoder, and a graph neural network to process cross-modal charging pile maintenance knowledge to construct a cross-modal charging pile maintenance knowledge base provided by an open embodiment of the present invention.
[0043] Figure 3 It is a flowchart of using a graph neural network to extract low-dimensional graph embeddings of a knowledge graph of entities involved in charging pile maintenance knowledge provided by an open embodiment of the present invention.
[0044] Figure 4 It is a schematic diagram of the joint training of a low-dimensional embedding model and a graph neural network provided by an open embodiment of the present invention.
[0045] Figure 5 It is a schematic diagram of a language model and a mapping layer provided by an open embodiment of the present invention.
[0046] Figure 6Schematic diagram of a diagnostic device based on a cross-modal charging pile maintenance knowledge base provided by an exemplary embodiment of the present invention. Detailed implementation manners
[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0048] It should be noted that in this document, the terms "include", "comprise", or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the phrase "including one..." does not exclude the presence of additional identical elements in the process, method, article, or device including the element.
[0049] Embodiment 1
[0050] Refer to Figure 1 As shown, an embodiment of the present invention provides a diagnostic method based on a cross-modal charging pile maintenance knowledge base, including:
[0051] S100. Use an attention-based image encoder, a text encoder, and a graph neural network to process cross-modal charging pile maintenance knowledge to construct a cross-modal charging pile maintenance knowledge base. The vectorized cross-modal charging pile maintenance knowledge in the cross-modal charging pile maintenance knowledge base includes high-dimensional cross-modal semantic encoding and low-dimensional graph embeddings.
[0052] During the specific implementation process, as Figure 2 shown, the process of step S100 includes:
[0053] S101. Sort out cross-modal charging pile maintenance knowledge in the form of graphics and texts and videos.
[0054] S102. Organize the cross-modal charging pile maintenance knowledge into a visually and textually consistent form.
[0055] In the graphic and text-based charging pile maintenance knowledge, the original information of the visual modality and the text modality appears in pairs, and the semantics are consistent between different modalities. The original information of the visual modality contains a large amount of spatial details related to the components to be maintained, but also contains a lot of redundant information irrelevant to the maintenance tasks. The original information of the text modality contains words with sparse semantics or low-correlation contexts.
[0056] In the video-based charging pile maintenance knowledge, the original information of the visual modality and the speech modality appears aligned along the time axis. Processing the video-based charging pile maintenance knowledge into a unified form with the graphic and text-based type involves: converting the speech in the video-based charging pile maintenance knowledge into temporally aligned subtitles, that is, extracting the original information of the text modality from the video; grouping the frames of the video-based charging pile maintenance knowledge according to the relationship between the subtitle semantics and the subtitle time axis to obtain a sub-frame sequence, and sampling representative sampling frames from the sub-frame sequence for each frame in the sub-frame sequence. The paired sampling frames and subtitle semantics are consistent.
[0057] S103, Use an attention-based image encoder and a text encoder to extract semantic features from the original information of the visually and textually consistent modalities respectively, and map them into high-dimensional cross-modal semantic encodings. Integrate the global information of the visual modality and the keyword information of the text to form a cross-modal semantic encoding, achieving cross-modal complementarity. This cross-modal complementarity provides richer information compared to graph embeddings and can effectively provide more abundant features in the generation task of the diagnostic processing solution. In this way, when the language model generates the diagnostic processing solution, it actually mainly refers to the cross-modal semantic encoding to generate content, avoiding the language model simply using graph-based embeddings to generate highly generalized content or obtaining hard-to-understand content by supplementing the generalized graph-based embedding framework based on its own hallucinations.
[0058] In the specific implementation process, the process includes:
[0059] Extract semantic features from the visual modality charging pile maintenance knowledge vi through an attention-based image encoder: Among them, the attention-based image encoder uses a specified layer of the Vision Transformer model, followed by a first mapping fully connected layer, and the first mapping fully connected layer outputs semantic features;
[0060] Extract semantic features from the text modality charging pile maintenance knowledge t i corresponding to the visual modality charging pile maintenance knowledge v i through an attention-based text encoder: Among them, the attention-based text encoder uses a specified layer of the language model of the Transformer architecture, followed by a second mapping fully connected layer, and the second mapping fully connected layer outputs semantic features with the same dimension as the first mapping fully connected layer.
[0061] During the semantic encoding process, the image encoder and text encoder based on the alignment loss constraint attention are used to achieve fine-grained alignment of semantic features. The formula for the alignment loss is as follows:
[0062]
[0063] L g (VisionEn(v i ),TextEn(t i )) represents the global image-text contrast loss obtained by analyzing the visual modality charging pile maintenance knowledge v i and the text modality charging pile maintenance knowledge t i through the encoder; L Region (r j,v ,r j,t ) is the matching loss between the regional semantics r i of the visual modality charging pile maintenance knowledge v j,v and the corresponding word r i in the text modality charging pile maintenance knowledge t j,t . λ1 and λ2 are the dynamic weights of the global image-text contrast loss and the local regional word pair matching loss respectively. r j,v and r j,t are randomly sampled. The random sampling of the alignment loss local is used to calculate the matching loss, which can effectively reduce the computational amount while meeting the alignment constraint.
[0064] The extracted cross-modal semantic features are mapped into cross-modal semantic encoding through semantic fusion.
[0065] The ways of the semantic fusion include:
[0066] Pool the semantic features of the visual modality and the text modality and then splice them, and then fuse them through a fusion multi-layer perceptron containing a non-linear activation function.
[0067] An example is as follows: F multi =MLP a (Pool1(F v )⊕Pool2(F t )) where Pool1() is multi-scale average pooling, which extracts the overall and local attention content from the semantic features of the visual modality through multi-scale average pooling, so as to exclude the redundancy irrelevant to the task; Pool2() is multi-scale max pooling; multi-scale max pooling captures the most significant features in the text modality semantics from different scales, such as relationships and entities, and ignores redundant context; MLP a() is a fusion multi-layer perceptron with a non-linear activation function; the concatenated features are mapped to a high-dimensional semantic space through a multi-layer perceptron with a non-linear activation function, and the non-linear learning of the complex interaction relationships of cross-modal features brought by the non-linear activation function is used to improve the representation ability, combine the visual modality and the text modality to achieve semantic enhancement, and convert sparse multi-modal features into compact and highly discriminative cross-modal semantic encodings.
[0068] The semantic features of the visual modality and the text modality are pooled and then cross-attention is performed, and then fused through a fusion multi-layer perceptron with a non-linear activation function.
[0069] An example is as follows: F multi = MLP a (crossAttention(Pool1(F v ), Pool2(F t ))), where Pool1() is multi-scale average pooling, which extracts global and local attention content from the semantic features of the visual modality at multiple scales, thereby excluding task-irrelevant redundancies; Pool2() is multi-scale max pooling; multi-scale max pooling captures the most significant features in the text modality semantics, such as relationships and entities, from different scales and ignores redundant context; crossAttention() is cross-attention, which replaces concatenation through cross-attention and enables dynamic interaction between the visual and text modality semantic features, and MLP a () is a fusion multi-layer perceptron with a non-linear activation function; the concatenated features are mapped to a high-dimensional semantic space through a multi-layer perceptron with a non-linear activation function, and the non-linear learning of the complex interaction relationships of cross-modal features brought by the non-linear activation function is used to improve the representation ability, combine the visual modality and the text modality to achieve semantic enhancement, and convert sparse multi-modal features into compact and highly discriminative cross-modal semantic encodings.
[0070] The semantic features of the visual modality and the text modality are pooled and then gated adaptive weighting is performed, and then fused through a fusion multi-layer perceptron with a non-linear activation function.
[0071] An example is as follows: F multi = μ·Pool1(F v )+(1 - μ)Pool2(F t ), Among them, Pool1() is multi-scale average pooling, which extracts global and local attention content from the semantic features of the visual modality at multiple scales through multi-scale average pooling, thereby excluding task-irrelevant redundancy; Pool2() is multi-scale max pooling; multi-scale max pooling captures the most significant features in the semantic of the text modality from different scales, such as relationships and entities, and ignores redundant context; μ is the gating weight, which automatically adjusts the modality contribution through the gating weight, W μ is the parameter matrix for calculating the gating weight, and σ() is the non-linear activation function of the gating weight.
[0072] In the above process, the pooling operation of the visual modality can also adopt spatial attention weighted pooling, and the pooling operation of the text modality can also adopt self-attention weighted pooling.
[0073] For the part of attention involved in the process, a signal-to-noise ratio adaptive coefficient γ is introduced for noise robustness optimization, and the method is as follows:
[0074]
[0075] q, k i , k j , v i are the query vector, the i-th key vector, the j-th key vector, and the i-th value vector respectively. γ is the signal-to-noise ratio adaptive coefficient. The greater the noise, the smaller the γ value, and the smaller the noise, the greater the γ value. The greater the noise, the smaller the γ value, and the smaller γ value makes the weight distribution more uniform, reducing the concentrated attention on the highly correlated keys that may be contaminated by noise, and resisting the influence of noise through more extensive attention; while when the noise is smaller, the γ value is larger, and the larger γ value, the weights will be concentrated on the highly correlated keys, improving accuracy; realizing that when facing noise, stable performance can still be maintained and robustness can be improved.
[0076] Through the above process, the charging pile maintenance knowledge is encoded into cross-modal semantic encoding and stored in the cross-modal charging pile maintenance knowledge base.
[0077] S104. Use the graph neural network to extract the low-dimensional graph embedding of the knowledge graph of the entities involved in the charging pile maintenance knowledge, and associate the same charging pile maintenance knowledge graph embedding with its cross-modal semantic encoding.
[0078] Among them, as Figure 3 shown, using the graph neural network to extract the low-dimensional graph embedding of the knowledge graph of the entities involved in the charging pile maintenance knowledge includes:
[0079] Extract the entities and the relationships between entities in the charging pile maintenance knowledge, use the entities as nodes, and use the relationships between entities as edges to construct a knowledge graph;
[0080] For any node and edge in the knowledge graph, use the low-dimensional embedding model to extract low-dimensional embedding features respectively;
[0081] The graph neural network determines the adjacency relationship according to the knowledge graph, aggregates information from neighbor entities for an entity to obtain a graph embedding, where the information includes the entity relationship between the neighbor entities and the low-dimensional embedding features of the neighbor entities.
[0082] The large-scale charging pile maintenance knowledge makes the overall scale of the cross-modal charging pile maintenance knowledge base relatively large. If similarity retrieval is performed on high-dimensional cross-modal speech coding, during the diagnosis process, the amount of calculation is large, the retrieval time is long, and the hardware specification requirements are high, which does not conform to the mobile application scenario of charging pile maintenance. To address this problem, in step S103, a knowledge graph is constructed, the low-dimensional embedding model extracts the low-dimensional embedding features of the nodes and edges of the knowledge graph, and the graph neural network aggregates based on the spatial structure, node features, and edge features of the knowledge graph to obtain a low-dimensional graph embedding. During the retrieval process, the similarity of the low-dimensional graph embedding is used for matching, which can greatly improve the retrieval efficiency.
[0083] When performing diagnostic processing using the charging pile diagnostic query, first, the embedding of the entity and relationship in the diagnostic query is extracted through the low-dimensional embedding model, and the most similar set number of target graph embeddings are matched using the similarity between the embedding of the entity and relationship in the diagnostic query and the graph embedding. As Figure 4 shown, the low-dimensional embedding model and the graph neural network need to be jointly trained. Specifically, a projection layer is set after the low-dimensional embedding model, the similarity between the output results of the low-dimensional embedding model and the projection layer and the graph embedding output by the graph neural network is calculated, and a similarity matrix of the graph embedding and the embedding of the entity and relationship in the diagnostic query is obtained. The similarity parameters in the similarity matrix conform to the similarity parameters between the diagnostic query and the corresponding knowledge graph of the graph embedding.
[0084] Obtain the cross-modal semantic coding that matches the target graph embedding.
[0085] The language model that supports processing cross-modal semantic coding uses the retrieved cross-modal semantic coding and the diagnostic query to generate a diagnostic processing plan. The language model that supports processing cross-modal semantic coding uses the retrieved cross-modal semantic coding and the diagnostic query to generate a diagnostic processing plan, including: using the cross-modal semantic coding and the diagnostic query as the prompt words of the language model, and controlling the language model to generate a diagnostic processing plan based on the high-dimensional cross-modal semantic coding. As Figure 5As shown in the figure, in order to enable the language model to generate a diagnostic processing solution based on cross-modal semantic encoding, a mapping layer for cross-modal semantic encoding is provided for the language model. The mapping layer is set in front of the Embedding model of the language model and is used to convert the cross-modal semantic encoding for the Embedding model to process. To achieve the above purpose, it is necessary to train the mapping layer to generate an encoding supported by the language model from the cross-modal semantic encoding. During the training process, the parameters of the language model are configured to be frozen, and the parameters of the mapping layer are configured to be adjustable. The language model receives the cross-modal semantic encoding through the mapping layer. It is expected that the content output by the language model is close to the rich charging pile maintenance knowledge corresponding to the cross-modal semantic encoding. The cross-entropy loss between the output content and the rich charging pile maintenance knowledge corresponding to the cross-modal semantic encoding is used as the loss function, and the parameters of the mapping layer are adjusted with the aim of minimizing the loss function.
[0086] Embodiment 2
[0087] Referring to Figure 6 As shown in the figure, the present invention provides a diagnostic device based on a cross-modal charging pile maintenance knowledge base, including: a sending end and a receiving end. Both the sending end and the receiving end include a processing unit, a storage unit, and a communication unit interconnected via a bus. The communication units of the sending end and the receiving end are connected. Among them, the storage unit stores a computer program. When the processing unit reads and executes the computer program, the diagnostic method based on the cross-modal charging pile maintenance knowledge base is implemented, including:
[0088] Using an attention-based image encoder, text encoder, and graph neural network to process cross-modal charging pile maintenance knowledge to construct a cross-modal charging pile maintenance knowledge base. Among them, the vectorized cross-modal charging pile maintenance knowledge in the cross-modal charging pile maintenance knowledge base includes high-dimensional cross-modal semantic encoding and low-dimensional graph embeddings;
[0089] When performing diagnostic processing using a charging pile diagnostic query, first extract the embeddings of entities and relationships in the diagnostic query through a low-dimensional embedding model;
[0090] Use the similarity between the embeddings of entities and relationships in the diagnostic query and the graph embeddings to match the closest set number of target graph embeddings;
[0091] Obtain the cross-modal semantic encoding that matches the target graph embedding;
[0092] A language model that supports processing cross-modal semantic encoding uses the retrieved cross-modal semantic encoding and the diagnostic query to generate a diagnostic processing solution.
[0093] Certainly, the computer program stored in the storage unit of a diagnostic device based on a cross-modal charging pile maintenance knowledge base provided by an embodiment of the present invention is not limited to the method operations described above, and can also execute related operations in a diagnostic method based on a cross-modal charging pile maintenance knowledge base provided by any embodiment of the present invention.
[0094] Embodiment 3
[0095] An embodiment of the present invention provides a computer-readable storage medium, which stores computer instructions, and when the computer instructions are executed by a processor, the diagnostic method based on a cross-modal charging pile maintenance knowledge base is implemented, including:
[0096] Using an attention-based image encoder, text encoder, and graph neural network to process cross-modal charging pile maintenance knowledge to construct a cross-modal charging pile maintenance knowledge base, wherein the vectorized cross-modal charging pile maintenance knowledge in the cross-modal charging pile maintenance knowledge base includes high-dimensional cross-modal semantic encoding and low-dimensional graph embedding;
[0097] When performing diagnostic processing using a charging pile diagnostic query, first extract the embeddings of entities and relationships in the diagnostic query through a low-dimensional embedding model;
[0098] Use the similarity between the embeddings of entities and relationships in the diagnostic query and the graph embedding to match the closest set number of target graph embeddings;
[0099] Obtain the cross-modal semantic encoding that matches the target graph embedding;
[0100] A language model that supports processing cross-modal semantic encoding uses the retrieved cross-modal semantic encoding and the diagnostic query to generate a diagnostic processing solution.
[0101] Certainly, the computer program stored in a computer-readable storage medium provided by an embodiment of the present invention is not limited to the method operations described above, and can also execute related operations in a diagnostic method based on a cross-modal charging pile maintenance knowledge base provided by any embodiment of the present invention.
[0102] In the embodiments provided by the present invention, it should be understood that the disclosed structure and method can be implemented in other ways. For example, the structural embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of structures or units can be in electrical, mechanical or other forms.
[0103] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0104] In addition, each functional unit in various embodiments of the present invention may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0105] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A diagnostic method based on a cross-modal charging pile maintenance knowledge base, characterized in that Including: Using an attention-based image encoder, text encoder, and graph neural network to process cross-modal charging pile maintenance knowledge to construct a cross-modal charging pile maintenance knowledge base. Among them, the vectorized cross-modal charging pile maintenance knowledge in the cross-modal charging pile maintenance knowledge base includes high-dimensional cross-modal semantic encoding and low-dimensional graph embedding; When performing diagnostic processing using a charging pile diagnostic query, first extract the embeddings of entities and relationships in the diagnostic query through a low-dimensional embedding model; Use the similarity between the embeddings of entities and relationships in the diagnostic query and the graph embedding to match the closest set number of target graph embeddings; Obtain the cross-modal semantic encoding that matches the target graph embedding; A language model that supports processing cross-modal semantic encoding uses the retrieved cross-modal semantic encoding and the diagnostic query to generate a diagnostic processing plan.
2. The diagnostic method based on the cross-modal charging pile maintenance knowledge base according to claim 1, wherein The use of an attention-based image encoder, text encoder, and graph neural network to process cross-modal charging pile maintenance knowledge to construct a cross-modal charging pile maintenance knowledge base includes: Organize cross-modal charging pile maintenance knowledge in the form of pictures, texts, and videos; Organize the cross-modal charging pile maintenance knowledge into a visually and textually consistent form; Use an attention-based image encoder and text encoder to extract semantic features from the original information of the visually and textually consistent modalities respectively, and map them into high-dimensional cross-modal semantic encoding; Use an attention-based graph neural network to extract the low-dimensional graph embedding of the knowledge graph of entities involved in the charging pile maintenance knowledge, and associate the same charging pile maintenance knowledge graph embedding with its cross-modal semantic encoding.
3. The diagnostic method based on the cross-modal charging pile maintenance knowledge base according to claim 2, wherein, For the attention parts involved in the attention-based image encoder, text encoder, and graph neural network, introduce a signal-to-noise ratio adaptive coefficient γ for noise robustness optimization, and the method is as follows: q, k i , k j , v i are the query vector, the i-th key vector, the j-th key vector, and the i-th value vector respectively. γ is the signal-to-noise ratio adaptive coefficient. The greater the noise, the smaller the value of γ, and the smaller the noise, the greater the value of γ.
4. The diagnostic method based on the cross-modal charging pile maintenance knowledge base according to claim 2, wherein The attention-based image encoder adopts a specified layer of the Vision Transformer model, followed by a first mapping fully connected layer; the attention-based text encoder adopts a specified layer of the language model of the Transformer architecture, followed by a second mapping fully connected layer; the first mapping fully connected layer and the second mapping fully connected layer output semantic features with the same dimension; Map the extracted cross-modal semantic features into cross-modal semantic encoding through semantic fusion.
5. The diagnostic method based on the cross-modal charging pile maintenance knowledge base according to claim 4, characterized in that, Based on the alignment loss, constrain the attention-based image encoder and text encoder to achieve fine-grained alignment of semantic features. Among them, the formula of the alignment loss is as follows: L g (VisionEn(v i ),TextEn(t i )) represents the maintenance knowledge of the visual modality charging pile v i and the maintenance knowledge of the text modality charging pile t i The global graph-text contrast loss obtained by encoder analysis; L Region (r j,v ,r j,t ) is the regional semantics r i of the visual modality charging pile maintenance knowledge v j,v and the text modality charging pile maintenance knowledge t i The matching loss of the corresponding word r j,t in them. λ1 and λ2 are the dynamic weights of the global graph-text contrast loss and the local regional word pair matching loss respectively. rj ,v and rj ,t are selected by random sampling.
6. The diagnostic method based on the cross-modal charging pile maintenance knowledge base according to claim 4, characterized in that, The methods of the semantic fusion include: Pool and splice the semantic features of the visual modality and the text modality, and then fuse them through a fusion multi-layer perceptron containing a non-linear activation function; Perform cross-attention on the semantic features of the visual modality and the text modality after pooling, and then fuse them through a fusion multi-layer perceptron containing a non-linear activation function; Perform gated adaptive weighting on the semantic features of the visual modality and the text modality after pooling, and then fuse them through a fusion multi-layer perceptron containing a non-linear activation function.
7. The diagnostic method based on the cross-modal charging pile maintenance knowledge base according to claim 2, wherein Using an attention-based graph neural network to extract the low-dimensional graph embedding of the knowledge graph of entities involved in the charging pile maintenance knowledge includes: Extract entities and relationships between entities in the charging pile maintenance knowledge, use the entities as nodes, and use the relationships between entities as edges to construct a knowledge graph; For any node and edge in the knowledge graph, use a low-dimensional embedding model to extract low-dimensional embedding features respectively; Based on the adjacency relationship determined by the knowledge graph, the graph neural network aggregates information from neighbor entities for the entity to obtain a graph embedding, where the information includes: the entity relationship between the neighbor entities and the low-dimensional embedding features of the neighbor entities.
8. The diagnostic method based on the cross-modal charging pile maintenance knowledge base according to claim 2, wherein, In the graphic and text-based charging pile maintenance knowledge, the original information of the visual modality and the text modality appears in pairs, and the semantics between different modalities are consistent; In the video-based charging pile maintenance knowledge, the original information of the visual modality and the speech modality appears aligned along the time axis. Process the video-based charging pile maintenance knowledge into the same form as the graphic and text-based one. The process includes: converting the speech of the video-based charging pile maintenance knowledge into subtitles aligned in time sequence; grouping the frames of the video-based charging pile maintenance knowledge according to the semantic relationship between the subtitles and the time axis of the subtitles to obtain a sub-frame sequence, and sampling representative frames from the sub-frame sequence for each frame in the sub-frame sequence. The paired sampled frames and subtitle semantics are consistent.
9. The diagnostic method based on a cross-modal charging pile maintenance knowledge base according to claim 1, characterized in that Using the retrieved cross-modal semantic encoding and diagnostic query to generate a diagnostic processing plan includes: embedding the cross-modal semantic encoding and diagnostic query as prompts for the language model, and controlling the language model to generate a diagnostic processing plan based on the high-dimensional cross-modal semantic encoding.
10. The diagnostic method based on a cross-modal charging pile maintenance knowledge base according to claim 1, characterized in that, In order to enable the language model to generate a diagnostic processing plan based on the cross-modal semantic encoding, provide a mapping layer for the cross-modal semantic encoding to the language model, and train the mapping layer to generate an encoding supported by the language model from the cross-modal semantic encoding.
Citation Information
Cited By
Charging equipment cloud call fault diagnosis method and device and computer equipment
CN120748397A
New energy automobile charging state test method based on electromagnetic compatibility
CN121253921A