Multi-modal power sample feature migration method and system based on dual cross-modal information decoupling, electronic equipment and storage medium
Through fine-grained image-text alignment and dual-stream attention mechanism, combined with visual semantic space alignment and model pruning algorithm, the alignment difficulties and feature representation differences in multimodal data processing in power scenarios are solved, the positioning accuracy and generalization ability of the model are improved, and the efficient migration and understanding of multimodal features are achieved.
Patent Information
- Application Number
- CN202510735635.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-12
AI Technical Summary
In the power scenario, existing multimodal data processing methods have the following problems: ignoring the fine-grained semantic relationship between images and text, difficulty in multimodal data alignment, large differences in feature representation, and noise problems, resulting in low information utilization, unstable model performance and insufficient generalization ability.
A fine-grained image-text alignment strategy and a dual-stream attention mechanism are adopted to capture multimodal representations through object position and context alignment, perform cross-modal information decoupling processing, and combine visual semantic space alignment and model pruning algorithms to generate diverse text samples and trim redundant parameters.
It improves the ability to understand and transfer multimodal features, enhances the model's positioning accuracy and understanding ability in power scenarios, reduces redundant data, and improves the model's generalization ability and computational efficiency.
Smart Images

Figure CN120632776A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of smart grid technology, and in particular to a multimodal power sample feature migration method, system, electronic device and storage medium based on dual cross-modal information decoupling. Background Art
[0002] With the rapid development of smart grids and the deepening of information technology in the power industry, power system operation data is showing a rapid growth trend. With the rapid development of artificial intelligence technology, advanced technical means to efficiently utilize and intelligently perceive power industry data have become mainstream.
[0003] The existing technologies are summarized as follows: (1) By acquiring multimodal data and using the federated learning framework for distributed training, a cross-regional power question-answering model is constructed; (2) By combining text, image and sensor data and using multimodal fusion technology, the question-answering model's ability to understand power scenarios is improved; (3) Through model pruning technology, the scale of the question-answering model is compressed, computing resource consumption is reduced, and the deployment requirements of power equipment are adapted; (4) Through the knowledge distillation method, the knowledge of large pre-trained models is transferred to lightweight models, improving the efficiency of question-answering and meeting real-time requirements.
[0004] The problems existing in the existing technologies are summarized as follows: (1) In the existing multimodal data processing methods based on deep learning, in the power scenario, the implicit alignment method ignores the fine-grained semantic relationship between images and texts, resulting in low information utilization and difficulty in being widely used in downstream tasks; (2) In the existing explicit alignment methods, in the power scenario, there is a lack of efficient multimodal data alignment mechanism, making it difficult to learn discriminative multimodal representations, resulting in difficulty in aligning data between modalities, complex fusion methods, and possible information redundancy or loss; (3) In the existing multimodal sample feature transfer methods, in the power scenario, the feature representations of different modal data are quite different, and there are noise and missing value problems in the multimodal data, resulting in unstable model performance and difficulty in adapting to the needs of complex power scenarios; (4) In the existing cross-modal mutual information evaluation methods, although the model performance is improved by maximizing or minimizing mutual information in the power scenario, in actual applications, the fine-grained alignment of cross-modal information still faces challenges, and it is difficult to ensure the semantic consistency between different modalities, which affects the generalization ability and real-time performance of the model.
[0005] To this end, how to provide a multimodal power sample feature migration method, system, device and storage medium based on dual cross-modal information decoupling that can effectively solve the problems existing in the above-mentioned existing technologies, further improve the understanding and migration capabilities of multimodal features, and realize the efficient operation of power system tasks is a problem that technical personnel in this field urgently need to solve. Summary of the Invention
[0006] In view of this, the present invention proposes a multimodal power sample feature migration method, system, electronic device and storage medium based on dual cross-modal information decoupling.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] A multimodal power sample feature transfer method based on dual cross-modal information decoupling includes:
[0009] A fine-grained image-text alignment strategy is used to hierarchically align visual and linguistic features to capture discriminative multimodal representations. This strategy includes object position alignment and context alignment, which are used to combine visual and text features to align electrical objects and their spatial locations, and to combine visual features with linguistic context features to capture the global relationship between image and text.
[0010] Based on the dual-stream attention mechanism, enhanced feature extraction is performed on multimodal data; the multimodal data includes visual features and language features processed by a fine-grained image-text alignment strategy. The dual-stream attention mechanism includes an image-adaptive attention mechanism and a text-guided attention mechanism.
[0011] Cross-modal information decoupling is performed on multimodal data to separate semantic features representing general semantics and modality-specific features representing modal characteristics. The semantic features and modality-specific features are then fine-grainedly aligned and recombined to obtain cross-modal fusion features. The fine-grained alignment and recombination process includes establishing fine-grained associations between semantic features and modality-specific features through an attention mechanism and adaptively weighting the features based on the association weights.
[0012] Migrate the deeply fused multimodal power sample features to downstream tasks of the power system.
[0013] Optionally, based on the object position alignment strategy, the power objects and their spatial positions are aligned by combining visual features and text features. Specifically:
[0014] The object position alignment strategy includes two branches: power object and spatial position;
[0015] First, the power object branch fuses visual features with text features through a cross-attention mechanism to extract features related to power objects;
[0016] Secondly, the spatial location branch captures the specific location information of the power object in the image through spatial prior guidance;
[0017] Finally, the outputs of the power object branch and the spatial position branch are fused through matrix multiplication to generate a refined feature representation that takes both the power object and the spatial position into consideration.
[0018] Optionally, based on the contextual alignment strategy, visual features are combined with language context features to capture the global relationship between image and text. Specifically:
[0019] Taking visual features as queries and language context features as keys and values, feature fusion is achieved through the pixel attention mechanism, and the fused features are adjusted through Tanh gating to enhance feature expression capabilities.
[0020] An optional fine-grained image-text alignment strategy also includes: introducing a channel modulation operation to recalibrate multimodal features through channel dependencies, specifically:
[0021] Through channel contraction and expansion operations, channel weights are generated and used to recalibrate multimodal features.
[0022] Optionally, based on the dual-stream attention mechanism, enhanced feature extraction is performed on multimodal data, specifically:
[0023] The image adaptive attention mechanism extracts key area information in the image by self-adjusting visual features;
[0024] The text-guided attention mechanism guides the attention allocation of visual features through text features.
[0025] Optionally, before migrating the deeply fused multimodal power sample features to downstream tasks of the power system, it also includes: using image-text sample augmentation technology based on visual semantic space alignment to map image features and text features to a unified visual semantic space to generate diverse text samples, and using a model pruning algorithm based on information resolution to evaluate the importance of parameters and trim redundant parameters.
[0026] Optionally, we can use image-text sample augmentation technology based on visual semantic space alignment to map image features and text features into a unified visual semantic space to generate diverse text samples. Specifically:
[0027] Extract image features through the visual encoder and map them into the semantic space;
[0028] Leverage pre-trained language models to generate diverse text descriptions that match image semantics.
[0029] Optionally, use an information resolution-based model pruning algorithm to evaluate the importance of parameters and prune redundant parameters, specifically:
[0030] The central server uses LoRA to initialize the trainable parameters and sends the initialized parameters to each data source.
[0031] The data source uses private data to fine-tune the LoRA trainable parameters and calculates the Fisher matrix of each parameter module during the fine-tuning process to evaluate the importance of the parameters.
[0032] The central server aggregates the Fisher information of each parameter module from multiple data sources, performs weighted summation, sorts the Fisher information of each parameter module, and selects the parameter module with the largest information value for importance marking;
[0033] The data source cuts out the remaining unimportant parameter modules based on the marking information of the central server and only retains the important parameter modules.
[0034] The present invention also provides a multimodal power sample feature migration system based on dual cross-modal information decoupling, which utilizes a multimodal power sample feature migration method based on dual cross-modal information decoupling, comprising:
[0035] Multimodal Representation Capture Module: This module uses a fine-grained image-text alignment strategy to hierarchically align visual and linguistic features to capture discriminative multimodal representations. The fine-grained image-text alignment strategy includes an object position alignment strategy and a context alignment strategy, which are used to combine visual and text features to align electrical objects and their spatial locations, and to combine visual features with linguistic context features to capture the global relationship between image and text.
[0036] Feature Representation Enhancement Module: This module is used to enhance feature extraction of multimodal data based on a dual-stream attention mechanism. The multimodal data includes visual features and language features processed by a fine-grained image-text alignment strategy. The dual-stream attention mechanism includes an image-adaptive attention mechanism and a text-guided attention mechanism.
[0037] Cross-modal deep fusion module: This module is used to decouple cross-modal information from multimodal data, separating semantic features that represent general semantics from modality-specific features that represent modality characteristics. It then performs fine-grained alignment and recombination processing on the semantic features and modality-specific features to obtain cross-modal fusion features. This fine-grained alignment and recombination processing includes establishing fine-grained associations between semantic features and modality-specific features through an attention mechanism, and adaptively weighting the features based on the association weights.
[0038] Multimodal power sample feature migration module: used to migrate the multimodal power sample features after deep fusion to downstream tasks of the power system.
[0039] The present invention further provides an electronic device, comprising:
[0040] Memory for storing computer programs;
[0041] A processor is used to implement the steps of a multimodal power sample feature migration method based on dual cross-modal information decoupling when executing a computer program.
[0042] The present invention also provides a computer-readable storage medium, characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of a multimodal power sample feature migration method based on dual cross-modal information decoupling are implemented.
[0043] Through the above technical solutions, it can be seen that compared with the existing technology, the present invention proposes a multimodal power sample feature migration method, system, electronic device and storage medium based on dual cross-modal information decoupling. The present invention introduces a new multimodal fusion method from the perspective of fine-grained alignment, namely a fine-grained image-text alignment strategy, to capture more discriminative representations, effectively improving the positioning accuracy, understanding ability and discrimination ability of the model; at the same time, the extracted multimodal features are captured through the image adaptive attention mechanism and the text-guided attention mechanism to capture important information, enhance the framework-level features and improve the semantic alignment in cross-modal analysis through weak semantic data, effectively model the correlation between different modal data, reduce redundant data, and improve the fusion effect of multimodal features; finally, by performing cross-modal information decoupling processing on the multimodal data, the semantic features representing the general semantics and the modal specific features representing the modal characteristics are separated, and the semantic features and the modal specific features are fine-grained aligned and reorganized to obtain cross-modal fusion features, which effectively improve the understanding and migration capabilities of the model in the power scenario. In addition, the present invention also introduces a visual semantic space alignment-based image-text sample augmentation technology and an information resolution-based model pruning algorithm, which are used to map image features and text features to a unified visual semantic space, generate diverse text samples, enhance the generalization ability of the model, and evaluate the importance of parameters, trim redundant parameters, and reduce the computational complexity of the model. The combination of the two technologies enables the model to further enhance the expression ability of multimodal features while maintaining its lightweight, thereby improving the generalization ability of the model. Based on this, the present invention performs well in power scene description, event location, and knowledge question-answering tasks. Compared with baseline models such as LLaVA and Qwen, it has significant improvements in indicators such as image description and location description, with the highest improvement reaching 12.8%, achieving further improvement in the understanding and transfer capabilities of multimodal features and the efficient operation of power system tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0045] Figure 1 Schematic diagram of the method of the present invention.
[0046] Figure 2 Schematic diagram of the multimodal data feature alignment framework for power systems of the present invention.
[0047] Figure 3 Schematic diagram of the fine-grained image-text alignment strategy of the present invention. DETAILED DESCRIPTION
[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0049] Example 1:
[0050] Embodiment 1 of the present invention discloses a multimodal power sample feature migration method based on dual cross-modal information decoupling, such as Figure 1 Shown, including:
[0051] A fine-grained image-text alignment strategy is adopted to perform hierarchical alignment of visual features and language features to capture discriminative multimodal representations; among them, the fine-grained image-text alignment strategy includes: object position alignment strategy and context alignment strategy, which are used to combine visual features with text features to align power objects and their spatial positions, and to combine visual features with language context features to capture the global relationship between image and text.
[0052] Based on the object position alignment strategy, the power objects and their spatial positions are aligned by combining visual features and text features. Specifically:
[0053] The object position alignment strategy includes two branches: power object and spatial position;
[0054] First, the power object branch fuses visual features with text features through a cross-attention mechanism to extract features related to power objects;
[0055] Secondly, the spatial location branch captures the specific location information of the power object in the image through spatial prior guidance;
[0056] Finally, the outputs of the power object branch and the spatial location branch are fused through matrix multiplication to generate a refined feature representation that considers both the power object and the spatial location. This strategy performs well in power inspection and fault detection tasks, effectively improving the model's positioning accuracy.
[0057] Optionally, based on the contextual alignment strategy, visual features are combined with language context features to capture the global relationship between image and text. Specifically:
[0058] Using visual features as queries and language context features as keys and values, this strategy fuses features using a pixel-based attention mechanism. Tanh gating is then used to adjust the fused features, further enhancing their expressiveness. This strategy performs well in power scene description tasks, generating descriptions that closely match power scenes and improving the model's comprehension capabilities.
[0059] The fine-grained image-text alignment strategy also includes: introducing channel modulation operations to recalibrate multimodal features through channel dependencies, enhancing cross-channel information exchange, and further improving the model's discriminative ability. Specifically:
[0060] Channel contraction and expansion operations generate channel weights, which are then used to recalibrate multimodal features. This operation effectively improves the model's discriminative capabilities, particularly in cross-modal retrieval tasks. Through channel modulation, the model can better capture semantic consistency across modalities, improving the reliability of information retrieval.
[0061] Based on the dual-stream attention mechanism, enhanced feature extraction is performed on multimodal data; wherein, the multimodal data includes: visual features and language features processed by a fine-grained image-text alignment strategy. The dual-stream attention mechanism includes: image adaptive attention mechanism and text-guided attention mechanism.
[0062] Based on the dual-stream attention mechanism, enhanced feature extraction is performed on multimodal data, specifically:
[0063] The image adaptive attention mechanism extracts key area information in the image by self-adjusting visual features;
[0064] The text-guided attention mechanism guides the attention allocation of visual features through text features, ensuring that the model can focus on image areas that are semantically related to the text.
[0065] This mechanism effectively models the correlation between different modal data, reduces redundant data, and improves the fusion effect of multimodal features, especially in power scenarios, and can better capture equipment status and abnormal information.
[0066] The multimodal data feature alignment framework for power systems of the present invention is as follows: Figure 2 As shown in Figure 2, we extract and match multimodal features from two perspectives. First, before the fine-grained image-text alignment strategy performs hierarchical alignment of visual features and language features to capture discriminative multimodal representations, we propose a dual cross-modal information disentanglement module (i.e., Figure 2 The feature extraction backbone network in
[15] is designed to extract fine-grained semantic information and separate it from the corresponding modality-specific information within each modality. Secondly, the extracted semantic features and modality features are aligned through a fine-grained image-text alignment strategy to capture more discriminative representations. Simultaneously, the extracted features are subjected to image-adaptive attention mechanisms and text-guided attention mechanisms to capture important information, enhance frame-level features, and improve semantic alignment in cross-modal analysis using weak semantic data.
[0067] Different from the traditional image-text alignment in previous methods, this paper introduces a new multimodal fusion approach from the perspective of fine-grained alignment, namely the fine-grained image-text alignment strategy, to capture more discriminative representations. Specifically, Figure 3 As shown, given the image visual features Language features An image-text alignment strategy is introduced to perform deep intersection between these visual and language features. Where C, H, and W represent the number of channels, height, and width of the visual feature map; D is the dimension of word embedding; N C 、N G 、N S Represents the length of the context, ground-truth target, and spatial location expression.
[0068] The core components of the fine-grained image-text alignment strategy are object position alignment strategy, context alignment strategy and channel modulation. The present invention proposes an object position alignment strategy to perform deep intersection with visual representation from power objects and spatial positions. This strategy achieves accurate alignment of object-related and spatial features and allows the model to capture more accurate relationships between objects and their positions in the image, thereby enhancing the reference segmentation performance. Specifically, the dual-branch structure is constructed by a power object branch and a spatial position branch. The main part of the power object branch is a power object cross attention that can integrate visual features F I and text features F G The present invention will F I As query and F G As the key and value to achieve feature fusion.
[0069] The multimodal data is subjected to cross-modal information decoupling processing to separate the semantic features representing the general semantics and the modal-specific features representing the modal characteristics, and the semantic features and modal-specific features are fine-grained aligned and recombined to obtain cross-modal fusion features. The fine-grained alignment and recombination processing includes: establishing fine-grained associations between semantic features and modal-specific features through the attention mechanism, and adaptively weighted fusion of features according to the association weights.
[0070] By decoupling visual and textual features, the model can separately process modality-specific information (such as texture and color in images) and cross-modal shared semantic information (such as device status and fault type). The decoupled features are recombined using a fine-grained alignment strategy to achieve deep cross-modal fusion, improving the model's understanding and transferability in power scenarios, particularly in equipment fault detection and knowledge question-answering tasks.
[0071] Step 4: Migrate the deeply fused multimodal power sample features to downstream tasks of the power system, such as power scenario description (CSD), event location (CELC), and knowledge question answering (CKQ). By comprehensively utilizing the information of paired text images, efficient operation of power system tasks can be achieved.
[0072] By leveraging paired text and image information, the model can efficiently process multimodal data in power scenarios. In scene description tasks, the model can generate image descriptions that closely match the power scenario. In event location tasks, the model can accurately pinpoint equipment fault areas. In knowledge question-answering tasks, the model combines image information and textual knowledge to provide specialized diagnoses and recommendations.
[0073] Before migrating the deeply fused multimodal power sample features to the downstream tasks of the power system, it also includes: using the image-text sample augmentation technology based on visual semantic space alignment to map image features and text features to a unified visual semantic space, generating diverse text samples, enhancing the generalization ability of the model, and using the information resolution-based model pruning algorithm to evaluate the importance of parameters, prune redundant parameters, and reduce the computational complexity of the model.
[0074] We use image-text sample augmentation technology based on visual semantic space alignment to map image and text features into a unified visual semantic space, generate diverse text samples, and enhance the generalization ability of the model. Specifically:
[0075] Extract image features through the visual encoder and map them into the semantic space;
[0076] Leverage pre-trained language models to generate diverse text descriptions that match image semantics.
[0077] In this way, the model can generate a large number of high-quality image-text samples for training and optimizing multimodal alignment tasks. This technology not only improves the model's performance in power scene description tasks, but also enhances the model's adaptability to unseen power imagery.
[0078] The information resolution-based model pruning algorithm is used to evaluate the importance of parameters, prune redundant parameters, and reduce the computational complexity of the model. Specifically:
[0079] The central server uses LoRA to initialize the trainable parameters and sends the initialized parameters to each data source.
[0080] The data source uses private data to fine-tune the LoRA trainable parameters and calculates the Fisher matrix of each parameter module during the fine-tuning process to evaluate the importance of the parameters.
[0081] The central server aggregates the Fisher information of each parameter module from multiple data sources, performs weighted summation, sorts the Fisher information of each parameter module, and selects the parameter module with the largest information value for importance marking;
[0082] The data source cuts out the remaining unimportant parameter modules based on the marking information of the central server and only retains the important parameter modules.
[0083] As shown in Table 1, the results of a quantitative evaluation of the performance of a multimodal power sample feature transfer method based on dual cross-modal information decoupling proposed in the present invention on the power transmission-substation-safety supervision dataset. The experiment uses four key indicators, IC, GC, REC, and REG, to measure the performance of the model on different tasks. By comparing the baseline methods (such as LLava, Qwen, etc.) with the method proposed in the present invention, it is observed that the decoupled feature transfer strategy significantly improves the generalization ability of the model on multimodal tasks, especially in scene understanding, equipment fault detection, and knowledge question-answering tasks, which are better than existing methods.
[0084] Table 1 Evaluation results of CSD, CELC and CKQ tasks on the power dataset
[0085]
[0086]
[0087] Experiments show that the method proposed in this invention performs excellently in power scene description, event location and knowledge question-answering tasks. Compared with baseline models such as LLaVA and Qwen, it has significant improvements in indicators such as image description and location description, with the highest improvement reaching 12.8%.
[0088] Example 2:
[0089] Embodiment 2 of the present invention discloses a multimodal power sample feature migration system based on dual cross-modal information decoupling using a multimodal power sample feature migration method based on dual cross-modal information decoupling, comprising:
[0090] Multimodal Representation Capture Module: This module uses a fine-grained image-text alignment strategy to hierarchically align visual and linguistic features to capture discriminative multimodal representations. The fine-grained image-text alignment strategy includes an object position alignment strategy and a context alignment strategy, which are used to combine visual and text features to align electrical objects and their spatial locations, and to combine visual features with linguistic context features to capture the global relationship between image and text.
[0091] Feature Representation Enhancement Module: This module is used to enhance feature extraction of multimodal data based on a dual-stream attention mechanism. The multimodal data includes visual features and language features processed by a fine-grained image-text alignment strategy. The dual-stream attention mechanism includes an image-adaptive attention mechanism and a text-guided attention mechanism.
[0092] Cross-modal deep fusion module: This module is used to decouple cross-modal information from multimodal data, separating semantic features that represent general semantics from modality-specific features that represent modality characteristics. It then performs fine-grained alignment and recombination processing on the semantic features and modality-specific features to obtain cross-modal fusion features. This fine-grained alignment and recombination processing includes establishing fine-grained associations between semantic features and modality-specific features through an attention mechanism, and adaptively weighting the features based on the association weights.
[0093] Multimodal power sample feature migration module: used to migrate the multimodal power sample features after deep fusion to downstream tasks of the power system.
[0094] Example 3:
[0095] Embodiment 3 of the present invention discloses an electronic device, including:
[0096] Memory for storing computer programs;
[0097] A processor is used to implement the steps of a multimodal power sample feature migration method based on dual cross-modal information decoupling when executing a computer program.
[0098] Example 4:
[0099] Embodiment 4 of the present invention discloses a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program implements the steps of a multimodal power sample feature migration method based on dual cross-modal information decoupling.
[0100] The embodiment of the present invention discloses a multimodal power sample feature migration method, system, electronic device and storage medium based on dual cross-modal information decoupling. The present invention introduces a new multimodal fusion method from the perspective of fine-grained alignment, namely a fine-grained image-text alignment strategy, to capture more discriminative representations, effectively improving the positioning accuracy, understanding ability and discrimination ability of the model; at the same time, the extracted multimodal features are captured through the image adaptive attention mechanism and the text-guided attention mechanism to capture important information, enhance the framework-level features and improve the semantic alignment in cross-modal analysis through weak semantic data, effectively model the correlation between different modal data, reduce redundant data, and improve the fusion effect of multimodal features; finally, by performing cross-modal information decoupling processing on the multimodal data, the semantic features representing the general semantics and the modal specific features representing the modal characteristics are separated, and the semantic features and the modal specific features are fine-grained aligned and reorganized to obtain cross-modal fusion features, which effectively improves the understanding and migration capabilities of the model in the power scenario. In addition, the present invention also introduces a visual semantic space alignment-based image-text sample augmentation technology and an information resolution-based model pruning algorithm, which are used to map image features and text features to a unified visual semantic space, generate diverse text samples, enhance the generalization ability of the model, and evaluate the importance of parameters, trim redundant parameters, and reduce the computational complexity of the model. The combination of the two technologies enables the model to further enhance the expression ability of multimodal features while maintaining its lightweight, thereby improving the generalization ability of the model. Based on this, the present invention performs well in power scene description, event location, and knowledge question-answering tasks. Compared with baseline models such as LLaVA and Qwen, it has significant improvements in indicators such as image description and location description, with the highest improvement reaching 12.8%, achieving further improvement in the understanding and transfer capabilities of multimodal features and the efficient operation of power system tasks.
[0101] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0102] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multimodal power sample feature transfer method based on dual cross-modal information decoupling, characterized by: include: A fine-grained image-text alignment strategy is used to hierarchically align visual and linguistic features to capture discriminative multimodal representations. The fine-grained image-text alignment strategy includes an object position alignment strategy and a context alignment strategy, which are respectively used to combine visual features with text features to align power objects and their spatial positions, and to combine visual features with linguistic context features to capture the global relationship between image and text. Enhanced feature extraction is performed on multimodal data based on a dual-stream attention mechanism; wherein the multimodal data includes visual features and language features processed by the fine-grained image-text alignment strategy; the dual-stream attention mechanism includes an image-adaptive attention mechanism and a text-guided attention mechanism; Performing cross-modal information decoupling on multimodal data to separate semantic features representing universal semantics and modality-specific features representing modality characteristics, and then fine-grained alignment and recombination of the semantic features and modality-specific features to obtain cross-modal fusion features. The fine-grained alignment and recombination process includes establishing fine-grained associations between semantic features and modality-specific features through an attention mechanism, and adaptively weighted fusion of features based on association weights. Migrate the deeply fused multimodal power sample features to downstream tasks of the power system.
2. The multimodal power sample feature transfer method based on dual cross-modal information decoupling according to claim 1 is characterized in that: Based on the object position alignment strategy, the power objects and their spatial positions are aligned by combining visual features and text features, specifically: The object position alignment strategy includes two branches: power object and spatial position; First, the power object branch fuses visual features with text features through a cross-attention mechanism to extract features related to power objects; Secondly, the spatial location branch captures the specific location information of the power object in the image through spatial prior guidance; Finally, the outputs of the power object branch and the spatial position branch are fused through matrix multiplication to generate a refined feature representation that takes both the power object and the spatial position into consideration.
3. The multimodal power sample feature transfer method based on dual cross-modal information decoupling according to claim 1 is characterized in that: Based on the context alignment strategy, visual features are combined with language context features to capture the global relationship between image and text. Specifically: Taking visual features as queries and language context features as keys and values, feature fusion is achieved through the pixel attention mechanism, and the fused features are adjusted through Tanh gating to enhance feature expression capabilities.
4. The multimodal power sample feature transfer method based on dual cross-modal information decoupling according to claim 1 is characterized in that: The fine-grained image-text alignment strategy further includes: introducing a channel modulation operation to recalibrate multimodal features through channel dependencies, specifically: Through channel contraction and expansion operations, channel weights are generated and used to recalibrate multimodal features.
5. The multimodal power sample feature transfer method based on dual cross-modal information decoupling according to claim 1 is characterized in that: Based on the dual-stream attention mechanism, enhanced feature extraction is performed on multimodal data, specifically: The image adaptive attention mechanism extracts key area information in the image by self-adjusting visual features; The text-guided attention mechanism guides the attention allocation of visual features through text features.
6. The multimodal power sample feature transfer method based on dual cross-modal information decoupling according to claim 1 is characterized in that: Before migrating the deeply fused multimodal power sample features to downstream tasks of the power system, it also includes: using the image-text sample augmentation technology based on visual semantic space alignment to map image features and text features to a unified visual semantic space to generate diverse text samples, and using the information resolution-based model pruning algorithm to evaluate the importance of parameters and trim redundant parameters.
7. The multimodal power sample feature transfer method based on dual cross-modal information decoupling according to claim 6 is characterized in that: We use image-text sample augmentation technology based on visual semantic space alignment to map image features and text features into a unified visual semantic space, generating diverse text samples. Specifically: Extract image features through visual encoder and map them into semantic space; Leverage pre-trained language models to generate diverse text descriptions that match image semantics.
8. The multimodal power sample feature transfer method based on dual cross-modal information decoupling according to claim 6 is characterized in that: The information resolution-based model pruning algorithm is used to evaluate the importance of parameters and prune redundant parameters. Specifically: The central server uses LoRA to initialize the trainable parameters and sends the initialized parameters to each data source. The data source uses private data to fine-tune the LoRA trainable parameters and calculates the Fisher matrix of each parameter module during the fine-tuning process to evaluate the importance of the parameters. The central server aggregates the Fisher information of each parameter module from multiple data sources, performs weighted summation, sorts the Fisher information of each parameter module, and selects the parameter module with the largest information value for importance marking; The data source cuts out the remaining unimportant parameter modules based on the marking information of the central server and only retains the important parameter modules.
9. A multimodal power sample feature transfer system based on dual cross-modal information decoupling using the multimodal power sample feature transfer method based on dual cross-modal information decoupling according to any one of claims 1 to 8, characterized in that: include: A multimodal representation capture module is configured to hierarchically align visual and linguistic features using a fine-grained image-text alignment strategy to capture discriminative multimodal representations. The fine-grained image-text alignment strategy includes an object position alignment strategy and a context alignment strategy, respectively combining visual and textual features to align power objects and their spatial locations, and combining visual features with linguistic contextual features to capture the global relationship between image and text. Feature representation enhancement module: used to perform enhanced feature extraction on multimodal data based on a dual-stream attention mechanism; wherein the multimodal data includes visual features and language features processed by the fine-grained image-text alignment strategy; the dual-stream attention mechanism includes an image-adaptive attention mechanism and a text-guided attention mechanism; Cross-modal deep fusion module: used to perform cross-modal information decoupling processing on multimodal data, separating semantic features that represent general semantics and modality-specific features that represent modality characteristics, and performing fine-grained alignment and recombination processing on the semantic features and modality-specific features to obtain cross-modal fusion features. The fine-grained alignment and recombination processing includes: establishing fine-grained associations between semantic features and modality-specific features through an attention mechanism, and performing adaptive weighted fusion of features based on the association weights; Multimodal power sample feature migration module: used to migrate the multimodal power sample features after deep fusion to downstream tasks of the power system.
10. An electronic device, characterized in that: include: memory for storing computer programs; A processor is configured to implement the steps of a multimodal power sample feature migration method based on dual cross-modal information decoupling as described in any one of claims 1 to 8 when executing the computer program.
11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a multimodal power sample feature migration method based on dual cross-modal information decoupling according to any one of claims 1 to 8.
Citation Information
Cited By
Electric power multi-modal sample knowledge enhancement question-answering method, system and device and medium
CN120851224A
Electric power multi-modal sample knowledge augmented question and answer method, system, device and medium
CN120851224B