Power cross-modal knowledge migration method and system, electronic equipment and storage medium

Through fine-grained data division and adaptive calibration, fine-grained relationship node sets are generated, and feature encoding and fine-grained feature extraction are performed on object nodes and structural nodes, which solves the problem that the existing technology is difficult to apply to complex scenarios, and realizes the capture of fine-grained features and the improvement of knowledge transfer.

CN120014508APending Publication Date: 2025-05-16CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510049847.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing multimodal knowledge migration method of power systems is difficult to apply to complex scenarios containing multiple objects, and it is impossible to better capture fine-grained features.

Method used

A cross-modal knowledge migration method for power is proposed. By obtaining inspection videos and text of the power system, fine-grained data division and adaptive calibration are performed, fine-grained relational node sets are generated, and feature encoding and fine-grained feature extraction are performed on object nodes and structural nodes to complete knowledge migration.

Benefits of technology

In complex scenarios containing multiple objects, fine-grained features can be better captured and the accuracy and efficiency of knowledge transfer can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014508A_ABST
    Figure CN120014508A_ABST
Patent Text Reader

Abstract

The invention belongs to a knowledge migration method, and provides an electric power cross-modal knowledge migration method and system, electronic equipment and a storage medium for solving the problems that an existing electric power system multi-modal knowledge migration method is not suitable for a complex scene containing a plurality of objects and cannot better capture fine granularity features. The method comprises the following steps: performing fine granularity data division on response output of an inspection video of a power system, adaptively calibrating feature relevance among different frames in the response output of the inspection video to obtain a calibrated video frame, generating a fine granularity relation node set according to the calibrated video frame, and generating a fine granularity relation node set; then feature coding is carried out on object nodes in the object node set and structure nodes in the structure node set, and knowledge migration of the inspection video is completed; meanwhile, fine-grained feature extraction is conducted on the object path and the structure path of the inspection text, object path features and structure path features are obtained, and knowledge migration of the inspection text is completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to a knowledge transfer method, and specifically to an electric power cross-modal knowledge transfer method, system, electronic device and storage medium. Background Art

[0002] Cross-modal knowledge transfer in power refers to the process of applying knowledge learned in one mode to another mode. In the power sector, it usually involves the fusion of multimodal data such as text, images, and sounds in order to more effectively utilize various information, improve operational efficiency, and enhance service quality. The application prospects of cross-modal knowledge transfer in the power sector are broad, bringing revolutionary improvements to business scenarios. Taking intelligent inspection as an example, this technology can accurately capture the status of power equipment, such as transmission lines and transformers, through image recognition. At the same time, combined with the equipment maintenance history and parameter information in the text data, it can greatly improve the accuracy and efficiency of fault diagnosis. In addition, the application of sound recognition technology also enables abnormal sounds of power equipment to be identified in a timely manner, which complements the equipment features described in the text and further enhances the accuracy of fault detection. The integrated application of these technologies not only optimizes the operation and maintenance management of power equipment, but also provides a strong guarantee for the safe and stable operation of the power system.

[0003] However, existing methods are mostly devoted to learning coarse-grained information of each modality, which is suitable for simple scenes containing only one object. For complex scenes containing multiple objects, it is not possible to capture fine-grained features well. Summary of the invention

[0004] This application aims to address the technical problems of existing multimodal knowledge transfer methods for power systems, which are not applicable to complex scenarios involving multiple objects and cannot capture fine-grained features well, and to provide a power cross-modal knowledge transfer method, system, electronic device and storage medium.

[0005] In order to achieve the above objectives, this application adopts the following technical solutions: In a first aspect, the present application proposes a method for cross-modal knowledge transfer in electric power, comprising: Obtain inspection videos and inspection texts of the power system respectively; Perform fine-grained data division on the response output of the inspection video, adaptively calibrate the feature correlation between different frames in the response output of the inspection video, and obtain calibrated video frames; According to the calibrated video frame, a fine-grained relationship node set including an object node set and a structure node set is generated; Feature encoding is performed on the object nodes in the object node set and the structure nodes in the structure node set respectively to complete the knowledge transfer of the inspection video; Fine-grained feature extraction is performed on the object path and structure path of the inspection text respectively to obtain object path features and structure path features, thus completing the knowledge transfer of the inspection text.

[0006] In the second aspect, the present application proposes a power cross-modal knowledge transfer system, comprising: A data acquisition module, used to respectively acquire inspection videos and inspection texts of the power system; A calibration module is used to perform fine-grained data division on the response output of the inspection video, adaptively calibrate the feature correlation between different frames in the response output of the inspection video, and obtain a calibrated video frame; A fine-grained relationship module, used for generating a fine-grained relationship node set including an object node set and a structure node set according to the calibrated video frame; A video feature encoding module is used to perform feature encoding on object nodes in the object node set and structure nodes in the structure node set, respectively, to complete knowledge transfer of inspection videos; The text feature extraction module is used to perform fine-grained feature extraction on the object path and structure path of the inspection text respectively, obtain object path features and structure path features, and complete the knowledge transfer of the inspection text.

[0007] In a third aspect, the present application proposes an electronic device, comprising: a memory, and one or more processors; the memory is coupled to the processor; wherein computer program code is stored in the memory, and the computer program code includes computer instructions, and when the computer instructions are executed by the processor, the electronic device performs the steps of the above-mentioned power cross-modal knowledge transfer method.

[0008] In a fourth aspect, the present application proposes a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned power cross-modal knowledge transfer method are implemented.

[0009] Compared with the prior art, this application has the following beneficial effects: This application proposes a method for cross-modal knowledge transfer of electric power, which performs fine-grained data division for the response output of the inspection video of the power system, adaptively calibrates the feature correlation between different frames in the response output of the inspection video, obtains the calibrated video frame, and then generates a fine-grained relationship node set based on the calibrated video frame, and then performs feature encoding on the object nodes in the object node set and the structure nodes in the structure node set respectively, to complete the knowledge transfer of the inspection video; at the same time, fine-grained feature extraction is performed on the object path and structure path of the inspection text respectively, to obtain object path features and structure path features, to complete the knowledge transfer of the inspection text, and to migrate both the inspection text and the inspection video to the vector space. In this application, a fine-grained adaptive frame calibration method is proposed for the inspection video, which can fuse multiple video frames to extract key frames, and then extracts the fine-grained relationship features of the key frames respectively, including object features at the semantic level and structural features between multiple objects. Next, fine-grained relationship nodes are generated, which can further enhance the representation ability of features. At the same time, for the inspection text, semantic paths and structural paths are used to capture complete dependencies to ensure that the fine-grained relationship of the text is effectively extracted. Therefore, the method of the present application is applicable to complex scenes containing multiple objects, and can also better capture fine-grained features in complex scenes.

[0010] The present application also proposes a power cross-modal knowledge transfer system, an electronic device and a computer storage medium, which have all the advantages of the above-mentioned power cross-modal knowledge transfer method. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0012] Figure 1 This is a first flow chart of the cross-modal knowledge transfer method for electric power in this application; Figure 2 This is a second flow chart of the cross-modal knowledge transfer method for electric power in this application; Figure 3 A connection diagram of the power cross-modal knowledge transfer system of this application. DETAILED DESCRIPTION

[0013] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations.

[0014] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application for which protection is sought, but merely represents selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.

[0015] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.

[0016] In the description of the embodiments of the present application, it should be noted that if the terms "upper", "lower", "horizontal", "inner", etc. appear, the orientation or position relationship indicated is based on the orientation or position relationship shown in the drawings, or the orientation or position relationship in which the invented product is usually placed when used. It is only for the convenience of describing the present application and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first", "second", etc. are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.

[0017] In addition, if the term "horizontal" appears, it does not mean that the component must be absolutely horizontal, but can be slightly tilted. For example, "horizontal" only means that its direction is more horizontal than "vertical", which does not mean that the structure must be completely horizontal, but can be slightly tilted.

[0018] In the description of the embodiments of the present application, it is also necessary to explain that, unless otherwise clearly specified and limited, the terms "set", "install", "connect", and "connect" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal connection of two components. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.

[0019] Cross-modal knowledge transfer in power refers to the process of applying knowledge learned in one modality to another modality. In the power sector, it usually involves fusing multimodal data such as text, images, and sounds to make more effective use of various information. Cross-modal knowledge transfer in power has broad application potential in business scenarios in the power sector. It can significantly improve operational efficiency, effectively reduce costs, and improve service quality. More importantly, it brings innovative solutions to the power industry. With the continuous advancement of technology, the application of cross-modal knowledge transfer in the power sector is expected to become more extensive and in-depth, further promoting the transformation and upgrading of the power industry.

[0020] Based on the above situation, the present application proposes a method, system, electronic device and storage medium for power cross-modal knowledge transfer. The present application is described in detail below in conjunction with embodiments and drawings.

[0021] like Figure 1 As shown, it is a first flow chart of the method for cross-modal knowledge transfer of electric power in this application, which may include: S101, respectively obtaining inspection videos and inspection texts of the power system.

[0022] Through the monitoring cameras or drone inspection equipment of the power system, the inspection video of the power system can be collected in real time or regularly. The video content can cover the key components of the power system, such as transformers, switchgear, transmission lines, etc., as well as their operating status and surrounding environment. The inspection text can include descriptive information related to the inspection video extracted from the power system's operation and maintenance records, maintenance reports, fault logs and other text materials. The text content can include detailed information such as the name, model, location, operating status, fault type, maintenance record, etc. of the equipment.

[0023] S102, performing fine-grained data division on the response output of the inspection video, adaptively calibrating the feature correlation between different frames in the response output of the inspection video, and obtaining calibrated video frames.

[0024] In practical applications, the inspection video can be divided into multiple frames, each of which represents a time point in the video. The feature correlation between different frames is analyzed, such as the object's motion trajectory, shape change, color change, etc. Adaptive algorithms are used to calibrate based on the feature correlation between frames to eliminate the inter-frame differences caused by camera movement, lighting changes, etc. The calibrated video frames are obtained, which have higher temporal and spatial consistency and accuracy.

[0025] S103: Generate a fine-grained relationship node set including an object node set and a structure node set according to the calibrated video frame.

[0026] The components or parts of the power system are identified from the calibrated video frames as object nodes. Each object node can then be identified and classified, such as transformers, switchgear, transmission lines, etc. The spatial relationships and connection relationships between object nodes are then analyzed, such as adjacent relationships, inclusion relationships, connection relationships, etc. Structural nodes can be generated based on these relationships to represent the structural and topological information between power system components. Finally, object nodes and structural nodes are combined to form a fine-grained relationship node set, which contains detailed information and structural features of the power system.

[0027] S104, feature encoding is performed on the object nodes in the object node set and the structure nodes in the structure node set respectively, so as to complete the knowledge transfer of the inspection video.

[0028] In practical applications, feature extraction is performed on each object node, such as shape features, texture features, color features, etc., and appropriate encoding methods, such as convolutional neural networks (CNNs) in deep learning, are used to encode the features of object nodes into high-dimensional vectors. Feature extraction can also be performed on each structural node, such as the connection relationship, distance relationship, and direction relationship between nodes, and the features of structural nodes can be encoded into vector form using graph neural networks (GNNs) or similar methods. Through feature encoding, the knowledge in the inspection video can be transferred to the vector space to form a feature representation that can be used for subsequent analysis and processing.

[0029] S105 , performing fine-grained feature extraction on the object path and the structure path of the inspection text respectively, obtaining object path features and structure path features, and completing knowledge transfer of the inspection text.

[0030] In practical applications, the object path described in the inspection text can be analyzed, such as the maintenance history of the equipment, the order of fault occurrence, etc., to extract key information in the object path, such as the equipment name, fault type, maintenance time, etc., and encode this information into vector form to form object path features. The structural path described in the inspection text can also be analyzed, such as the topological structure of the power system, the connection relationship between devices, etc., to extract key information in the structural path, such as the connection method and transmission path between devices, etc., and encode this information into vector form to form structural path features. Through fine-grained feature extraction, the knowledge in the inspection text is transferred to the vector space to form a feature representation that can be used for subsequent analysis and processing.

[0031] like Figure 2 As shown, it is a second flow chart of the electric power cross-modal knowledge transfer method of the present application, which may include: S201, fine-grained relational feature learning for video streams.

[0032] (1) Fine-grained adaptive video frame calibration.

[0033] Input Video ,Include Frame, denoted as ,in .for , Indicates Frame, the response activation of a convolutional layer can be regarded as a 3D tensor , the three dimensions refer to height, width and depth. Based on this, the video The response output can be expressed as It should be noted that each frame in the video will generate a response activation when it passes through a convolutional layer in the convolutional neural network. The response output is a set of response activations for all frames, containing a 3D tensor produced after each frame of the video passes through the convolutional layer.

[0034] To enhance the video The feature representation ability of the video The response output is divided into fine-grained data and the feature correlation between different video frames is adaptively calibrated. Specifically: First, a global average pooling operation is used to learn a set of frame-level statistics. ,for ,make

[0035] in, Indicates height, Indicates width, Indicates depth, Represents a 3D tensor.

[0036] This set of frame-level statistics reflects the overall activation level of each frame after the convolution layer. In addition, global average pooling helps the network learn a more global and robust feature representation, and can also reduce the dimension of the feature map, thereby reducing computational complexity and memory consumption.

[0037] Then, in order to enhance the model's expressiveness and learn high-level abstract features, A nonlinear transformation operation is performed:

[0038] in, represents the attention score vector, represents the relu activation function, represents the sigmoid activation function, represents the weight matrix of the first fully connected layer, Represents the weight matrix of the second fully connected layer. Specifically, when performing nonlinear transformation, the frame-level statistics are first passed through the first fully connected layer to obtain the intermediate feature representation, and then the ReLU activation function is applied to the intermediate feature representation to introduce nonlinearity, and then passed through the second fully connected layer, and the sigmoid activation function is applied to obtain the attention score vector. Through this process, the model can learn high-level abstract features, enhance the model's expressiveness, and enable the model to better handle complex data and tasks.

[0039] At the same time, for video In the response output of Frame response output , perform a flatten operation on it to obtain a one-dimensional ,but , Represents a set of frame feature vectors. It should be noted that the flatten operation refers to a flattening operation, which is usually used to convert a multidimensional array into a one-dimensional array.

[0040] Finally, in order to achieve effective complementarity between frames and obtain robust video features, each frame is weighted and aggregated according to its importance to obtain a calibrated video frame. :

[0041] , in, represents a diagonal matrix, Indicates The attention score vector of the frame, is the calibration frame after attention processing. It should be noted that weighted aggregation can make full use of the information of each frame, highlight the contribution of important frames, and weaken the influence of unimportant frames, thereby processing noise and redundant information in the video, improving the robustness and discrimination of video features, and improving the accuracy of subsequent target detection.

[0042] (2) Calibration frame object-structure feature generation.

[0043] For the calibrated video frames , the scene graph generation method is used to generate a fine-grained relationship node set, including an object node set and a structure node set, denoted as .in, Represents the set of object nodes, corresponding to the noun part of each calibration frame, that is, the specific objects or entities appearing in the video frame. These objects can be people, animals, objects, scene elements, etc., which constitute the basis of the video content. Represents a set of structural nodes, corresponding to the connecting part between two nouns, that is, the relationship or interaction between objects. These relationships can be spatial relationships (such as "on the left of...", "on...", etc.), temporal relationships (such as "before", "after", etc.), or semantic relationships (such as "hold", "wear", etc.).

[0044] It should be noted that scene graph is a structured data representation method that can present objects and their relationships in video frames in the form of a graph. In a scene graph, nodes represent objects or structural elements, and edges represent connections or relationships between objects. By using scene graph generation methods to generate fine-grained relationship nodes, we can have a deeper understanding of the objects and their relationships in the video content.

[0045] (3) Calibration frame object-structure feature encoding The pre-trained Swin Transformer model (a Transformer-based visual model) is used to train the object node set and structure node sets Visual feature extraction is performed, and the corresponding visual feature extraction results are recorded as and ,in, and Respectively represent object node features and It should be noted that the pre-trained Swin Transformer model is a Swin Transformer model that has been pre-trained on a large-scale dataset and has learned rich visual features, which can be directly used for feature extraction or as a basis for fine-tuning.

[0046] Then, auxiliary label features of object nodes and structure nodes are extracted and .in, and They are the unique hot encoding of object nodes and the unique hot encoding of structure nodes, which represent object nodes respectively. and structure nodes Tags, represents the auxiliary label weight matrix of the object node, Represents the auxiliary label weight matrix of the structural nodes.

[0047] Next, the visual features and auxiliary label features of object nodes and structure nodes are jointly learned. That is, the joint learning result of object nodes for:

[0048] Joint learning results of structural nodes for:

[0049] in, represents the visual-auxiliary label joint feature weight matrix.

[0050] In order to further improve the representation ability of features, a graph neural network is designed to update object nodes and structure nodes. By promoting information exchange between neighboring nodes, information exchange between different nodes is achieved. Among them, object nodes are updated by themselves, and structure nodes are updated by aggregation of neighboring object nodes. Specifically: The initialization characteristics when updating object nodes and structure nodes are expressed as:

[0051]

[0052] Then, let , .

[0053] in, Indicates The object node feature vector obtained after the update is Indicates The object node feature vector obtained after the update is Indicates The structural node feature vector obtained after the update is Indicates The feature vector of the object node at one end of the graph neural network edge obtained after the update, Indicates The structural node feature vector obtained after the update is Indicates The feature vector of the object node at the other end of the graph neural network edge obtained after the update, represents a fully connected layer, followed by a tanh activation function. Finally, the output of the graph neural network is the object node feature vector and the structural node feature vector .

[0054] S202, learning fine-grained relational features of text streams.

[0055] (1) Text object-structure feature generation Similar to the video modality, the text modality also contains rich fine-grained information. Traditional processing methods often rely on text segmentation tools to identify and extract nouns and verbs. However, this method often disrupts the internal logical order of the text, resulting in a lack of contextual information. In order to better retain text features and capture contextual information, this application adopts two paths to represent objects and structural information respectively: the first path is the object path, which captures scene fragments and reflects the contextual order relationship of all tokens in the text. The second path is the structural path, which shows the dependencies between the subject, predicate, and object within the sentence. These dependencies can be constructed through triples parsed by the SPICE (Simulation Program with Integrated Circuit Emphasis, a simulation program focusing on integrated circuits) method.

[0056] Specifically, enter a text that consists of words and The object path is composed of three tuples (subject, predicate, and object). Then the object path is 1, and the length is , expressed as , the structure path is The length of the bar is 3, which is expressed as .

[0057] (2) Text object-structure feature encoding For object path Since the self-attention mechanism has some restrictions on the length of text, we can first set the text length to the maximum length supported by the model. , for more than The part with length less than The [pad] operation is used to fill the part to ensure that the length of all paths is consistent. Then, considering that a text may contain different scene fragments, the object path can be segmented, which is recorded as the scene segmentation result. :

[0058] in, There are 4 markers for scene segmentation, which are used to mark the beginning and end of two scene segments in the whole sentence. and Represents the length of two scene clips respectively. represents the semantic feature of the first position, Indicates The semantic features of the position, Indicates The semantic features of the position, Indicates The semantic features of the position, Indicates The semantic features of the position, Indicates Semantic features of a position.

[0059] In addition, set To represent valid tags and padding tags ([pad]), where 1 represents a valid tag and 0 represents an invalid tag.

[0060] Finally, and The concatenation is performed and sent as input to the pre-trained BERT model for feature extraction.

[0061] Similarly, for the structure path , the pre-trained BERT model is also used for feature extraction.

[0062] Finally, the output of the pre-trained BERT is the object path feature and structural path characteristics .

[0063] This application also designs the boundary loss function of the BERT model, which is defined as:

[0064] in, and Represent the video and text in the training set respectively, , , They represent that the video and text selected in a batch are not matched, the video and text selected in a batch are matched, and the text and video selected in a batch are not matched. The optimization goal is to improve ,reduce and , so that the scores of matching pairs are as high as possible, and the scores of mismatching pairs are as low as possible. This application can spontaneously extract fine-grained information from coarse-grained samples. Video-text pairs are composed of object-structure fine-grained relationship blocks. Therefore, the loss function optimizes video-text pairs, which is essentially optimizing fine-grained relationship blocks.

[0065] S203, fine-grained similarity scoring mechanism.

[0066] In order to accurately evaluate the similarity between the inspection video and inspection text modalities of the power system, this application also proposes a refined similarity scoring system. Unlike the coarse-grained similarity calculation performed at the overall video-text level by traditional methods, this application innovatively introduces a fine-grained similarity calculation mechanism, focusing on two detailed levels of objects and structures for measurement. This design is inspired by a deep understanding of the nature of video-text matching: if the objects or structures mentioned in the inspection text are fully and matched in the video, they should receive a higher matching score. Taking video structure-text structure matching as an example, the scoring formula is defined as follows:

[0067] The above formula means to traverse all video structure nodes and text structure nodes , ensuring that every element of the two-dimensional matrix is ​​greater than or equal to 0. Among them, express and The matching score between .

[0068] Considering the equivalence of the physical meaning of summation and maximum value, the simplified formula is as follows:

[0069] This formula means that for all text structure nodes, the video structure node that best matches them is retained.

[0070] Similarly, for video object-text object matching, the scoring formula is defined as follows:

[0071] in, Represents a video object node and text structure nodes The matching score between .

[0072] Next, we sum the video object-text object matches and the video structure-text structure matches to get a fine-grained similarity score, which is defined as follows:

[0073] in, Represents a fine-grained similarity score.

[0074] This application uses two typical modal data, video and text, as examples to perform cross-modal knowledge transfer. For the video modality, a fine-grained adaptive frame calibration method is proposed, which can fuse multiple video frames to extract key frames; then the fine-grained relationship features of the key frames are extracted separately, including object features at the semantic level and structural features between multiple objects; next, the scene graph generation method is used to generate fine-grained relationship nodes. In order to further improve the representation ability of the features, a graph neural network is designed to update the two types of nodes, and by promoting information exchange between neighboring nodes, information exchange between different nodes is achieved. At the same time, for the text modality, semantic paths and structural paths are used to capture complete scene fragments and triple dependencies, respectively, to ensure that fine-grained text relationships are effectively extracted. Finally, a fine-grained similarity scoring mechanism is designed to align the fine-grained relationship features of the two modalities to achieve cross-modal knowledge transfer.

[0075] like Figure 3 As shown, it is a connection diagram of the electric power cross-modal knowledge transfer system of the present application, which may include: A data acquisition module, used to respectively acquire inspection videos and inspection texts of the power system; A calibration module is used to perform fine-grained data division on the response output of the inspection video, adaptively calibrate the feature correlation between different frames in the response output of the inspection video, and obtain a calibrated video frame; A fine-grained relationship module, used for generating a fine-grained relationship node set including an object node set and a structure node set according to the calibrated video frame; A video feature encoding module is used to perform feature encoding on object nodes in the object node set and structure nodes in the structure node set, respectively, to complete knowledge transfer of inspection videos; The text feature extraction module is used to perform fine-grained feature extraction on the object path and structure path of the inspection text respectively, obtain object path features and structure path features, and complete the knowledge transfer of the inspection text.

[0076] In some embodiments of the electric power cross-modal knowledge transfer system of the present application, in the calibration module, fine-grained data division is performed on the response output of the inspection video, and the feature correlation between different frames in the response output of the inspection video is adaptively calibrated, including: The global average pooling operation is used to learn the statistics of the frame level corresponding to the response output of the inspection video; Perform nonlinear transformation on the frame-level statistics to obtain the corresponding attention score vector; Flatten the response output of the inspection video to obtain a set of frame feature vectors; The attention score vector and frame feature vector of each frame are weighted and aggregated to obtain a calibrated video frame.

[0077] In some embodiments of the power cross-modal knowledge transfer system of the present application, in the video feature encoding module, feature encoding is performed on object nodes in the object node set and structure nodes in the structure node set, including: Extracting visual features from the relationship node set and the structure node set respectively to obtain a relationship visual feature extraction result and a structure visual feature extraction result; Extract corresponding auxiliary label features from the visual feature extraction results and the structural visual feature extraction results respectively, and record them as auxiliary label features of object nodes and auxiliary label features of structural nodes; For object nodes and structure nodes, the visual feature extraction results and the auxiliary label features are jointly learned respectively to obtain the joint learning results of the object nodes and the joint learning results of the structure nodes.

[0078] In some embodiments of the power cross-modal knowledge transfer system of the present application, after obtaining the joint learning results of the object nodes and the joint learning results of the structure nodes in the video feature encoding module, the following further comprises: The joint learning results of object nodes and structure nodes are updated separately through graph neural networks.

[0079] In some embodiments of the electric power cross-modal knowledge transfer system of the present application, the text feature extraction module performs fine-grained feature extraction on the object path and the structure path of the inspection text respectively, including: Perform scene segmentation on the object path and the structure path respectively to obtain scene segmentation results of the object path and the structure path; The object path and the structure path are truncated or filled respectively to obtain the object path and the structure path after length processing; Respectively obtain the tag sets of the object path and the structure path after length processing; the tag sets include multiple valid tags and multiple filling tags; splicing the scene segmentation result of the object path and the tag set corresponding to the object path, and splicing the scene segmentation result of the structure path and the tag set corresponding to the structure path to obtain an object splicing result and a structure splicing result; The object splicing results and the structure splicing results are respectively input into the pre-trained BERT model for feature extraction to obtain object path features and structure path features.

[0080] In some embodiments of the power cross-modal knowledge transfer system of the present application, the boundary loss function of the pre-trained BERT model is for:

[0081] in, and Represent the video and text in the training set respectively, , , They respectively represent that the video and text selected in a batch are unmatched pairs, the video and text selected in a batch are matched pairs, and the text and video selected in a batch are unmatched pairs.

[0082] It should be noted that in the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the system embodiments described above are merely schematic. For example, the division of each module is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules can be combined or integrated into another device, or some features can be ignored or not executed. The module described as a separate component may or may not be physically separated. The component displayed as a module may be a physical unit or multiple physical units, that is, it may be located in one place, or it may be distributed in multiple different places. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.

[0083] In addition, each module in each embodiment of the present invention may be integrated into a processing unit, each module may exist physically separately, or two or more modules may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0084] An embodiment of the present application also provides an electronic device, which may include one or more processors, a memory, and a communication interface.

[0085] The memory, the communication interface and the processor are coupled, for example, the memory, the communication interface and the processor may be coupled together via a bus.

[0086] The communication interface is used for data transmission with other devices. The memory stores computer program code. The computer program code includes computer instructions, and when the computer instructions are executed by the processor, the electronic device executes the steps of the above-mentioned power cross-modal knowledge transfer method.

[0087] Wherein, the processor can be a processor or a controller, for example, a central processing unit (CPU), a general processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the present disclosure. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of DSP and microprocessors, and the like. The processor can be used to support electronic devices to execute the method steps provided in the above embodiments.

[0088] The bus may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The above bus may be divided into an address bus, a data bus, a control bus, etc.

[0089] An embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned power cross-modal knowledge transfer method are implemented.

[0090] The computer-readable storage medium involved in the present application includes random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in the technical field.

[0091] The above are only preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for cross-modal knowledge transfer in electric power, characterized in that: include: Obtain inspection videos and inspection texts of the power system respectively; Perform fine-grained data division on the response output of the inspection video, adaptively calibrate the feature correlation between different frames in the response output of the inspection video, and obtain calibrated video frames; According to the calibrated video frame, a fine-grained relationship node set including an object node set and a structure node set is generated; Feature encoding is performed on the object nodes in the object node set and the structure nodes in the structure node set respectively to complete the knowledge transfer of the inspection video; Fine-grained feature extraction is performed on the object path and structure path of the inspection text respectively to obtain object path features and structure path features, thus completing the knowledge transfer of the inspection text.

2. The method for cross-modal knowledge transfer of electric power according to claim 1, characterized in that: The step of dividing the response output of the inspection video into fine-grained data and adaptively calibrating the feature correlation between different frames in the response output of the inspection video includes: The global average pooling operation is used to learn the statistics of the frame level corresponding to the response output of the inspection video; Perform nonlinear transformation on the frame-level statistics to obtain the corresponding attention score vector; Flatten the response output of the inspection video to obtain a set of frame feature vectors; The attention score vector and frame feature vector of each frame are weighted and aggregated to obtain a calibrated video frame.

3. The method for cross-modal knowledge transfer of electric power according to claim 1, characterized in that: The generating of fine-grained relationship nodes including object nodes and structure nodes includes: generating the fine-grained relationship nodes by using a scene graph generating method.

4. The method for cross-modal knowledge transfer of electric power according to claim 1, characterized in that: The feature encoding of the object nodes in the object node set and the structure nodes in the structure node set includes: Extracting visual features from the relationship node set and the structure node set respectively to obtain a relationship visual feature extraction result and a structure visual feature extraction result; Extract corresponding auxiliary label features from the visual feature extraction results and the structural visual feature extraction results respectively, and record them as auxiliary label features of object nodes and auxiliary label features of structural nodes; For object nodes and structure nodes, the visual feature extraction results and the auxiliary label features are jointly learned respectively to obtain the joint learning results of the object nodes and the joint learning results of the structure nodes.

5. The method for cross-modal knowledge transfer of electric power according to claim 4, characterized in that: After obtaining the joint learning results of the object nodes and the joint learning results of the structure nodes, the method further includes: The joint learning results of object nodes and structure nodes are updated separately through graph neural networks.

6. The method for cross-modal knowledge transfer of electric power according to claim 1, characterized in that: The fine-grained feature extraction is performed on the object path and the structure path of the inspection text respectively, including: Perform scene segmentation on the object path and the structure path respectively to obtain scene segmentation results of the object path and the structure path; The object path and the structure path are truncated or filled respectively to obtain the object path and the structure path after length processing; Respectively obtain the tag sets of the object path and the structure path after length processing; the tag sets include multiple valid tags and multiple filling tags; splicing the scene segmentation result of the object path and the tag set corresponding to the object path, and splicing the scene segmentation result of the structure path and the tag set corresponding to the structure path to obtain an object splicing result and a structure splicing result; The object splicing results and the structure splicing results are respectively input into the pre-trained BERT model for feature extraction to obtain object path features and structure path features.

7. The method for cross-modal knowledge transfer of electric power according to claim 6, characterized in that: Boundary loss function of the pre-trained BERT model for: in, and Represent the video and text in the training set respectively, , , They respectively represent that the video and text selected in a batch are unmatched pairs, the video and text selected in a batch are matched pairs, and the text and video selected in a batch are unmatched pairs.

8. A power cross-modal knowledge transfer system, characterized in that: include: A data acquisition module, used to respectively acquire inspection videos and inspection texts of the power system; A calibration module is used to perform fine-grained data division on the response output of the inspection video, adaptively calibrate the feature correlation between different frames in the response output of the inspection video, and obtain a calibrated video frame; A fine-grained relationship module, used for generating a fine-grained relationship node set including an object node set and a structure node set according to the calibrated video frame; A video feature encoding module is used to perform feature encoding on object nodes in the object node set and structure nodes in the structure node set, respectively, to complete knowledge transfer of inspection videos; The text feature extraction module is used to perform fine-grained feature extraction on the object path and structure path of the inspection text respectively, obtain object path features and structure path features, and complete the knowledge transfer of the inspection text.

9. The power cross-modal knowledge transfer system according to claim 8, characterized in that: In the calibration module, fine-grained data division is performed on the response output of the inspection video, and feature correlation between different frames in the response output of the inspection video is adaptively calibrated, including: The global average pooling operation is used to learn the statistics of the frame level corresponding to the response output of the inspection video; Perform nonlinear transformation on the frame-level statistics to obtain the corresponding attention score vector; Flatten the response output of the inspection video to obtain a set of frame feature vectors; The attention score vector and frame feature vector of each frame are weighted and aggregated to obtain a calibrated video frame.

10. The power cross-modal knowledge transfer system according to claim 8, characterized in that: In the video feature encoding module, feature encoding is performed on the object nodes in the object node set and the structure nodes in the structure node set, including: Extracting visual features from the relationship node set and the structure node set respectively to obtain a relationship visual feature extraction result and a structure visual feature extraction result; Extract corresponding auxiliary label features from the visual feature extraction results and the structural visual feature extraction results respectively, and record them as auxiliary label features of object nodes and auxiliary label features of structural nodes; For object nodes and structure nodes, the visual feature extraction results and the auxiliary label features are jointly learned respectively to obtain the joint learning results of the object nodes and the joint learning results of the structure nodes.

11. The power cross-modal knowledge transfer system according to claim 10, characterized in that: In the video feature encoding module, after obtaining the joint learning results of the object nodes and the joint learning results of the structure nodes, the module further includes: The joint learning results of object nodes and structure nodes are updated separately through graph neural networks.

12. The power cross-modal knowledge transfer system according to claim 8, characterized in that: In the text feature extraction module, fine-grained feature extraction is performed on the object path and structure path of the inspection text respectively, including: Perform scene segmentation on the object path and the structure path respectively to obtain scene segmentation results of the object path and the structure path; The object path and the structure path are truncated or filled respectively to obtain the object path and the structure path after length processing; Respectively obtain the tag sets of the object path and the structure path after length processing; the tag sets include multiple valid tags and multiple filling tags; splicing the scene segmentation result of the object path and the tag set corresponding to the object path, and splicing the scene segmentation result of the structure path and the tag set corresponding to the structure path to obtain an object splicing result and a structure splicing result; The object splicing results and the structure splicing results are respectively input into the pre-trained BERT model for feature extraction to obtain object path features and structure path features.

13. The power cross-modal knowledge transfer system according to claim 12, characterized in that: Boundary loss function of the pre-trained BERT model for: in, and Represent the video and text in the training set respectively, , , They respectively represent that the video and text selected in a batch are unmatched pairs, the video and text selected in a batch are matched pairs, and the text and video selected in a batch are unmatched pairs.

14. An electronic device, characterized in that: include: A memory and one or more processors; the memory is coupled to the processor; wherein the memory stores computer program code, the computer program code includes computer instructions, and when the computer instructions are executed by the processor, the electronic device executes the steps of the power cross-modal knowledge transfer method as described in any one of claims 1-7.

15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the power cross-modal knowledge transfer method according to any one of claims 1 to 7 are implemented.