A component segmentation method and system based on Hodge decomposition and multimodal fusion

By using Hodge decomposition and multimodal fusion, multimodal features are decomposed into gradient, curl and harmonic components, which solves the problem of structural semantics being difficult to distinguish in the segmentation of power components in the existing technology, and achieves high-precision and robust component recognition.

CN120707579BActive Publication Date: 2025-10-31HUNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511195689.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-10-31
Estimated Expiration
2045-08-26

AI Technical Summary

Technical Problem

Existing multimodal segmentation methods struggle to effectively distinguish and coordinate various structural semantics in power component segmentation, leading to blurred boundaries, missing structural information, or misjudgments, which fails to meet the demands for high-precision recognition and anomaly detection.

Method used

The method of Hodge decomposition and multimodal fusion is adopted to decompose the multimodal features into gradient components, curl components and harmonic components, which correspond to the structural boundary, local changes and global spatial information of the target, respectively. The precise coordination and dynamic fusion are achieved through component-level cross-attention mechanism.

Benefits of technology

It significantly improves the accuracy, robustness, and structural interpretability of power component segmentation, enabling accurate identification of component boundaries and structures in complex environments and achieving high-precision pixel-level segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707579B_ABST
    Figure CN120707579B_ABST
Patent Text Reader

Abstract

This invention discloses a component segmentation method and system based on Hodge decomposition and multimodal fusion. The method consists of five core steps: multimodal data acquisition and preprocessing, Hodge decomposition-driven fractional feature extraction, component-level structured multimodal fusion, structure-aware component collaborative modeling, and segmentation output. Notably, this invention introduces Hodge decomposition theory into the multimodal segmentation network for the first time. For each modal feature field, it orthogonally decomposes it into gradient components, curl components, and harmonic components. Then, for different component features, it achieves deep collaboration and dynamic fusion of different modal information under the same component within the component-level structural space. The final result is a power component segmentation result with accurate boundaries, clear structure, and strong spatial consistency. It fully leverages the prior advantages of physical structure components in multimodal perception, significantly improving component recognition accuracy, robustness, and interpretability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent segmentation technology for power components, specifically to a component segmentation method and system based on Hodge decomposition and multimodal fusion, particularly applicable to power component segmentation and component defect detection in industrial scenarios. Background Technology

[0002] Current mainstream multimodal segmentation technologies for power components typically rely on deep convolutional neural networks or Transformer architectures, using multi-source data such as visible light, infrared images, and LiDAR point clouds as input. Each modality's data is first initially represented by a backbone feature extraction network. Then, in the intermediate layers or the segmentation head stage of the segmentation network, multimodal features are comprehensively utilized through features concatenating, weighted fusion, channel attention, spatial attention, and cross attention. The fused features are then processed through decoding and upsampling to finally output pixel-level segmentation results. The advantage of this approach lies in its ability to fully leverage the complementary information of each modality's data, improving segmentation accuracy and robustness in complex environments. Therefore, it is widely used in practical scenarios such as power line inspection and intelligent identification.

[0003] In recent years, to further enhance modal collaborative perception capabilities, some methods have introduced multi-level, multi-scale attention mechanisms or Transformer structures into the network architecture to achieve deeper feature interactions. However, regardless of the fusion strategy employed, existing methods are essentially still based on the linear superposition of features in high-dimensional space or self-attention weight adjustment. Their fusion process fails to distinguish different types of structural information within multimodal features, especially lacking targeted structured modeling at the physical level, such as component boundaries, rotational structures, and global consistency. Overall, current fusion methods for multimodal segmentation networks still primarily rely on overall feature mixing, making it difficult to achieve effective physical-level decoupling and component-level collaboration for complex structural semantics, thus becoming a major technical bottleneck for further improving segmentation performance. In other words, these existing multimodal segmentation methods, due to their reliance on traditional methods such as feature concatenation, weighted fusion, or attention allocation, often fail to adequately distinguish and express the physical structural layers within multimodal data.

[0004] In real-world power grid inspection scenarios, the boundaries between components often become blurred due to changes in lighting, noise interference, and complex backgrounds. Local structural details and overall spatial relationships are also easily weakened or even lost during the fusion process. Furthermore, most common multimodal cross-attention mechanisms are limited to information interaction between overall modalities, lacking targeted modeling of specific structural features. This results in segmentation models performing poorly when dealing with complex shapes and multi-source noise, limiting the application scope of multimodal segmentation methods in power components.

[0005] With the increasing demands for automation and intelligent inspection in power systems, higher requirements are being placed on the high-precision intelligent segmentation of power components in complex scenarios. However, in practical applications, existing deep learning-based multimodal segmentation methods still have significant shortcomings in structural semantic decoupling, fine boundary recognition, and spatial consistency due to factors such as complex on-site environments, diverse component structures, and significant noise in multimodal data. Specifically, during multimodal data fusion, the local boundaries and complex structures of components are easily interfered with by intermodal noise and redundant information, leading to blurred boundaries, missing structural information, or misjudgments in the segmentation results, making it difficult to meet the actual needs of high-precision identification and anomaly detection of key components in power inspection. Therefore, how to effectively distinguish and coordinate various structural semantics in a multimodal environment to improve the accuracy, robustness, and structural interpretability of power component segmentation is a core technical problem that urgently needs to be solved in this field. Summary of the Invention

[0006] This invention aims to solve the aforementioned problems in the segmentation accuracy of power components, particularly the lack of structural interpretability. It provides a component segmentation method and system based on Hodge decomposition and multimodal fusion. This method is not limited to power component segmentation but can be extended to other fields. This invention can be applied to any segmentation task involving multimodal image input and sensitive to structural information. For example, component defect detection in industrial settings, organ or lesion segmentation in medical images, or extraction of structures such as roads and buildings from remote sensing images can all utilize component-level modeling and fusion methods. Specifically, this invention proposes introducing Hodge decomposition into a multimodal feature fusion and segmentation network. Through learnable differential operators and structure-aware convolutional kernels, the fused multimodal features are decomposed into different components with clear physical meanings, corresponding to the target's structural boundaries, local changes, and global spatial information, respectively. Simultaneously, a component-level multimodal cross-attention mechanism is combined to achieve precise coordination and dynamic fusion of various structural features, significantly improving the segmentation network's ability to perceive complex power component structures, spatial relationships, and boundaries, effectively overcoming the shortcomings of existing methods in structural representation and noise suppression.

[0007] Therefore, the present invention provides the following technical solution:

[0008] A component segmentation method based on Hodge decomposition and multimodal fusion includes the following steps:

[0009] Multimodal data acquisition refers to acquiring and segmenting multimodal images of the target scene;

[0010] Hodge decomposition-driven fractional feature extraction means that for each modal image corresponding to a modality, three parallel sub-networks for gradient component extraction, curl component extraction, and harmonic component extraction are used to extract a class of Hodge decomposition features for each modal image. The Hodge decomposition features include gradient components, curl components, and harmonic components.

[0011] Component-level structured multimodal fusion and structure awareness;

[0012] The component-level structured multimodal fusion involves fusing the same type of Hodge decomposition features from different modes to obtain fused features corresponding to gradient components, curl components, and harmonic components; and the structure perception involves performing structure perception on the features of gradient components, curl components, and harmonic components.

[0013] Feature fusion: fusing current features to obtain comprehensive structural perception features;

[0014] The segmentation output uses the integrated structural perception features as input to the segmentation head to obtain the component segmentation result of the target scene.

[0015] Further, optionally, the mathematical model of the gradient component extraction subnet is expressed as:

[0016]

[0017] In the formula, For gradient component feature maps, Operators for extracting gradient components, For gradient convolution weight matrix, These are the bias terms in the mathematical model corresponding to the gradient components; This represents the input image, specifically the corresponding modal image. The modal information is represented by the first convolutional kernel of the gradient component extraction subnet, which is initialized with the structural features corresponding to the focused gradient components as the target.

[0018] The mathematical model of the curl component extraction subnet is expressed as follows:

[0019]

[0020] In the formula, This is a characteristic map of the curl component. Operator for extracting curl components, Here is the curl convolution weight matrix. The bias term in the mathematical model corresponding to the curl component is defined as follows: The first convolutional kernel of the curl component extraction subnetwork is initialized with the objective of focusing on the structural features corresponding to the curl component.

[0021] The mathematical model of the harmonic component extraction subnet is expressed as follows:

[0022]

[0023] In the formula, This is a characteristic diagram of harmonic components. Operator for harmonic component extraction, The harmonic convolution weight matrix is... For global average pooling, the first layer convolutional kernel of the harmonic component extraction subnet is initialized with the goal of focusing on the structural features corresponding to the harmonic components.

[0024] Further, optionally, a component-level Selective Attention fusion mechanism is introduced, which uses the Selective Attention operator to fuse the same type of Hodge decomposition features of different modalities to obtain the fused features corresponding to the gradient component, curl component, and harmonic component respectively.

[0025] The mathematical model of the Selective Attention operator is expressed as follows:

[0026]

[0027] In the formula, This represents the Selective Attention operator for a class of Hodge decomposition features j. and Representing modes and modality The corresponding Hodge decomposition feature j; j is the category label of the Hodge decomposition feature, corresponding to the gradient component, curl component, and harmonic component; , Both are learnable linear transformation matrices, corresponding to the gradient components of the two modes respectively; Indicates feature splicing; , All are attention mechanism parameters; σ represents nonlinear activation.

[0028] Further, optionally, if component-level structured multimodal fusion is performed first, followed by structure-aware and feature fusion, the generation process of the comprehensive structure-aware features is as follows:

[0029] After component-level structured multimodal fusion, the fused features of gradient components, curl components, and harmonic components are input into the structure-aware branch networks corresponding to various Hodge decomposition features for structure perception and feature extraction, represented as follows:

[0030]

[0031] In the formula, , , These represent the features extracted after the fusion features of gradient components, curl components, and harmonic components are processed by a structure-aware branch network. , , These represent the fusion characteristics of the gradient component, curl component, and harmonic component, respectively. , , These represent structure-aware branch networks;

[0032] The features extracted by the three types of structure-aware branch networks are then weighted and fused according to learnable weights to obtain the comprehensive structure-aware features, represented as:

[0033]

[0034] in, , , All are adaptive fusion weights obtained through end-to-end training. This is for the comprehensive structural perception features.

[0035] Further optionally, the component segmentation method pre-constructs a sample set of the target scene to be segmented, trains the network model based on the sample set and the designed loss function, and then segments the target scene to be segmented.

[0036] Among them, in addition to measuring the predicted segmentation mask With real labels In addition to the loss metric based on similarity between the components, the loss function also includes structure-aware loss. The structure-aware loss is used to constrain the expressive power of quantized features in spatial structure. The component loss is used to ensure that the three structural components of gradient, curl and harmonics maintain physical consistency and orthogonality in spatial distribution and expression, and to prevent structural aliasing.

[0037] Further, alternatively, structure-aware loss Represented as:

[0038]

[0039]

[0040] In the formula, For structural orthogonal loss, Assuming energy loss, , These are the Hodge decomposition features of class i and class j after component-level structured multimodal fusion, respectively, where i and j take grad, curl, and harm, respectively, representing the gradient component, curl component, and harmonic component. The L2 norm (Euclidean norm) is the square of the number of features used to measure the distance or difference between two features. To represent the original overall features of the input image, i.e., the undecomposed, fused features before Hodge decomposition-driven feature extraction, , , These represent the features extracted by the structure-aware branch network after the fusion features of the gradient component, curl component, and harmonic component.

[0041] Further optionally, the modal image is divided into visible light image and infrared image; the component is an electrical component, a part in an industrial scene, a road or building in a remote sensing image, or an organ or lesion in a medical image.

[0042] Furthermore, the present invention also provides a system based on the component segmentation method, including a multimodal data acquisition module, a fractional feature extraction module, a component-level structured multimodal fusion module, a structure-aware component fusion module, and a segmentation output module that are connected sequentially or interconnected.

[0043] The multimodal data acquisition module is used to acquire multimodal images of the target scene for segmentation.

[0044] The fractional feature extraction module is used to extract a class of Hodge decomposition features corresponding to each modal image through three parallel gradient component extraction subnetworks, curl component extraction subnetworks, and harmonic component extraction subnetworks for each modal image. The Hodge decomposition features include gradient components, curl components, and harmonic components.

[0045] The component-level structured multimodal fusion module is used to fuse the same type of Hodge decomposition features of different modes to obtain the fused features corresponding to the gradient component, curl component, and harmonic component, respectively.

[0046] The structure-aware component fusion module is used to perform structure perception on the features of gradient components, curl components, and harmonic components.

[0047] The structure-aware component fusion module is also used for feature fusion, that is, to fuse the current features to obtain comprehensive structure-aware features;

[0048] The segmentation output module is used to take the integrated structural perception features as input to the segmentation head to obtain the component segmentation result of the segmentation target scene.

[0049] The present invention also provides a computer device, comprising:

[0050] One or more processors and a memory storing one or more computer programs;

[0051] The processor invokes the computer program to implement the steps of a component segmentation method based on Hodge decomposition and multimodal fusion.

[0052] The present invention also provides a computer-readable storage medium storing a computer program, which is called by a processor to implement the steps of a component segmentation method based on Hodge decomposition and multimodal fusion.

[0053] Beneficial effects

[0054] 1. This invention proposes a component segmentation method based on Hodge decomposition and multimodal fusion. Its core principle is to decouple and collaboratively model the structural information in the multimodal fusion feature field at the physical component level. Specifically, the multimodal features are decomposed into gradient components, curl components, and harmonic components, corresponding to the boundary structure, local deformation, and global consistency information of the segmented component, respectively. Subsequently, for different component features, deep collaboration and dynamic fusion of different modal information under the same component are achieved within the component-level structural space, effectively enhancing the expressive power of key structural information. Furthermore, a structure-aware branch network is used to deeply model the three types of structural features, and component-level fusion is achieved in the output stage, ultimately obtaining segmented component results with accurate boundaries, clear structures, and strong spatial consistency. The overall solution overcomes the limitations of traditional multimodal segmentation methods that rely solely on feature splicing or overall attention fusion, fully leveraging the prior advantages of physical structural components in multimodal perception, and significantly improving the accuracy, robustness, and interpretability of component recognition in complex power system scenarios.

[0055] 2. In practical applications, the technical solution of this invention has a wide range of potential and outstanding effects. It is not limited to power component segmentation, but is also applicable to any segmentation task that uses multimodal images as input and is sensitive to structural information. Especially in actual power inspection scenarios, the boundaries between components often become blurred due to changes in lighting, noise interference, and complex backgrounds, and local structural details and overall spatial relationships are easily weakened or even lost during the fusion process. To this end, the technical solution of this invention introduces Hodge decomposition theory into multimodal segmentation networks for the first time. For each modal feature field, it orthogonally decomposes it into gradient components, curl components, and harmonic components. The gradient components correspond to the edge and local texture structure of the image, the curl components correspond to structural features such as rotation, bending, and entanglement in the image, and the harmonic components represent spatial structures such as global consistency and background redundancy. Based on this, the technical solution of this invention uses three components to achieve structured modeling at the physical level such as component boundaries, rotational structures, and global consistency, thereby improving the interpretability of the structure. Finally, the model can accurately identify key structures such as the boundaries of power components, local deformation of insulators, and wire entanglement in substation scenarios under complex backgrounds and multi-source noise interference, achieving high-precision pixel-level segmentation. Attached Figure Description

[0056] Figure 1 This is a schematic diagram of the overall architecture of a bimodal component-level structured segmentation network based on Hodge decomposition.

[0057] Figure 2 This is a diagram showing the connection relationships of the hardware components of the computer device provided by the present invention. Detailed Implementation

[0058] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. The technical features involved in the various embodiments of the invention described below can be combined with each other as long as they do not conflict with each other.

[0059] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0061] To achieve high-precision structured segmentation of power components in complex power scenarios, this invention proposes a multimodal component-level structured segmentation network based on Hodge decomposition. The overall structure of the dual-modal component-level structured segmentation network, using visible light and infrared images as examples, is as follows: Figure 1 As shown, the process consists of five core steps, with the following technical approach: Step S1 is multimodal data acquisition and preprocessing; Step S2 is Hodge decomposition-driven fractional feature extraction; Step S3 is component-level structured multimodal fusion; Step S4 is structure-aware component collaborative modeling; and Step S5 is segmentation output. It should be noted that this structure diagram is a single-layer illustration. In practical applications, steps S2 to S4 can be stacked into a multi-layered progressive structure according to task requirements to further improve the model's multi-scale structural modeling capability and segmentation accuracy. Furthermore, the execution order of steps S3 and S4 can be adaptively adjusted.

[0062] The present invention will be described in detail below with reference to specific embodiments. In the example, an electric component is used as the segmentation object. It should be understood that the technical concept of the present invention is not limited to the application field of the segmentation object, but only to the electric field.

[0063] Example 1

[0064] Step S1: Multimodal Data Acquisition. The purpose of this step is to acquire multimodal image data of the segmented target scene.

[0065] In this embodiment, the multimodal image data acquired by the on-site multi-sensor platform consists of visible light and infrared images. Conventional image preprocessing methods such as scaling, cropping, and normalization are used to uniformly adjust the two types of images to have the same spatial size and pixel distribution. After preprocessing, a visible light image with strictly one-to-one spatial pixel correspondence is obtained. and infrared images H is the height of the image, and W is the width of the image. Given a set of real numbers, the above processing provides a standardized input basis for subsequent feature component extraction and multimodal collaborative fusion, effectively ensuring the consistency and comparability of the data.

[0066] In other feasible embodiments, the multimodal image data is not limited to the image types described above, and the specific preprocessing techniques are not limited to those implemented in this embodiment. In practical applications, the modal image type and preprocessing techniques can be set according to application requirements, and this invention does not impose specific limitations on them.

[0067] Step S2: Segmentation feature extraction based on Hodge decomposition.

[0068] In two-dimensional space, the Hodge decomposition theorem states that for any vector field... All of them can be uniquely decomposed into gradient components (irrotational), curl components (dispersion-free), and harmonic components (irrotational and divergent), which can be expressed mathematically as follows:

[0069]

[0070] in, The gradient component corresponds to the irrotational structure of the field, reflecting local boundary and texture changes; For curl components, It is a curl potential function or a vector potential field. The divergence-free structure of the corresponding field represents deformations such as local rotation and entanglement. In the feature extraction stage, this invention proposes for the first time a Hodge decomposition-driven fractional feature extraction sub-network. Specifically, after inputting the original image, a three-way parallel component sub-network is constructed for each mode, directly used to extract the three types of feature components corresponding to the physical structure. It should be understood that as long as the component sub-network can extract the corresponding gradient component, curl component, and harmonic component, it meets the requirements of this invention.

[0071] To encourage the three sub-networks to focus on the feature components of their corresponding physical structures, this embodiment of the invention preferably guides the sub-networks to focus on the features of different physical structures by setting the initialization method of the first-layer convolutional kernel. For example, for gradient components, curl components, and harmonic components, differential kernels, rotation kernels, and smoothing kernels are used respectively to guide the focus on different structures. This method can effectively achieve functional structural decoupling. Secondly, it is also preferable to add structural loss and orthogonality constraints during training to guide each sub-network to focus on extracting the corresponding structural components. In this case, there are no strict requirements on the structure of the three types of sub-networks; for example, a stacked convolutional structure can be used. A detailed explanation follows.

[0072] S2.1: Gradient Component Extraction Subnet. The purpose of this gradient component extraction subnet is to extract the gradient components from the Hodge decomposition features of each modality image.

[0073] The gradient components correspond to the edges and local texture structure of the image, and are represented as follows:

[0074]

[0075] In the formula, For gradient component feature maps, Operators for extracting gradient components, For gradient convolution weight matrix, This represents the bias term in the mathematical model corresponding to the gradient components. This represents the input image, specifically the corresponding modal image. This represents modal information; in this embodiment, m can take the following values: or , representing visible light or infrared light, further expanded to:

[0076]

[0077] In the formula, , These are the gradient component feature maps corresponding to the visible light and infrared modal images, respectively. , These are gradient component extraction operators for visible light and infrared modal images, respectively. , These are visible light and infrared modal images, respectively. , These are the gradient convolution weight matrices corresponding to the visible light and infrared modal images, respectively; , These are the bias terms corresponding to the gradient components in the visible light and infrared modal images, respectively.

[0078] In this embodiment, a learnable convolution kernel in the form of a first-order difference is used to approximate the gradient operator of a continuous field. This is because the first-order difference form can better simulate gradient changes in an image, especially edge structures. Furthermore... Initialize as a first-order difference convolution template, such as in the horizontal direction:

[0079]

[0080] Or vertically:

[0081]

[0082] In the formula, and These are convolutional kernels in the horizontal and vertical directions, meaning that horizontal and vertical convolutional kernels are set simultaneously in the same gradient subnet to extract horizontal and vertical edge features from the image at the same time.

[0083] It should be understood that the technical solution of the present invention does not limit the initial value of the weights. It must be a first-order difference kernel. In principle, as long as the convolutional kernel can effectively converge to the directional features representing image edges or texture changes during training, it meets the requirements. This is because this component essentially extracts non-rotational (gradient) information from the image, focusing on areas of drastic structural change such as component boundaries and insulation damage. For example, in the technical solution of this invention, the first-layer convolutional kernel is initialized as a gradient direction difference kernel (such as a Sobel form), thus guiding the subnetwork to respond to image edges and detailed structures in the early stages of the network. Secondly, in other embodiments, it is also preferable to enhance its response strength in areas such as boundaries and abrupt structural changes by guiding the loss function.

[0084] In this embodiment, according to and Assigning weights as convolution kernels Subsequent parameter updates can be adaptively optimized through end-to-end training to optimally approximate the physical gradient structure of the modality image. End-to-end training and adaptive optimization means that this gradient subnetwork is not trained independently, but rather as part of the entire backbone network, participating synchronously in forward propagation and backward gradient updates during the overall network training process. In other words, setting the initial convolutional kernel to a first-order difference form only provides prior guidance on the physical structure; during actual training, its weight parameters... and bias All of them are learnable and are jointly optimized and updated through the main loss function.

[0085] S2.2 Curl Component Extraction Subnetwork. The purpose of this curl component extraction subnetwork is to extract the curl component from the Hodge decomposition features of each modality image. The core design of the curl subnetwork is to introduce a loop difference kernel or a convolutional kernel with rotational symmetry during the initialization phase to guide the network to respond to information such as local rotation, bending, and texture distortion in the image. For example, commonly used Laplace loops and Sobel crossover kernels have a certain degree of curl sensitivity.

[0086] The curl component corresponds to structural features in the image such as rotation, bending, and entanglement, and is represented as:

[0087]

[0088] In the formula, This is a characteristic map of the curl component. Operator for extracting curl components, Here is the curl convolution weight matrix. This represents the bias term in the mathematical model corresponding to the curl component. This embodiment focuses on visible light or infrared images, and further elaborates as follows:

[0089]

[0090] In the formula, , These are the curl component feature maps corresponding to the visible light and infrared modal images, respectively. , These are the curl component extraction operators for visible light and infrared modal images, respectively. , These are the curl convolution weight matrices corresponding to the visible light and infrared modal images, respectively; , These are the bias terms corresponding to the curl components in visible light and infrared modal images, respectively.

[0091] In this embodiment, a learnable curl convolution kernel is used to approximate a two-dimensional curl operator in the sub-network. Initialize it with templates such as loop difference, convolution groups, rotationally symmetric filters, etc., like curl approximation operators for two-dimensional fields or higher-order loop difference kernels. Here, it can be initialized with the following 3×3 discrete curl operator to sense the local rotation and loop structure of the image:

[0092]

[0093] In the formula, This is a curl component convolution kernel; subsequent weights can be adaptively trained. This component essentially extracts non-divergent but spin-based structures, focusing on local rotational deformations such as cable bends and component entanglements. It should be understood that the above convolution kernel design is a heuristic prior guidance method, aiming to help the subnet focus more quickly on rotational structures or local bending information in the image during the early stages of network training. This is merely an example from an effectiveness perspective, and the technical solution of this invention is not limited to this type of convolution kernel design.

[0094] S2.3 Harmonic Component Extraction Subnet. The purpose of this harmonic component extraction subnet is to extract the harmonic components from the Hodge decomposition features of each modality image. The harmonic components mainly correspond to regions with consistent background and slow-changing areas in the image. This invention guides the subnet to focus on low-frequency, smooth region features by initializing it with a large-size mean convolution kernel or Gaussian kernel combined with global average pooling. During training, the harmonic subnet is decoupled from the gradient and curl subnets, and it extracts the residual portion after removing edge and rotation information, thus achieving effective modeling of harmonic features.

[0095] Harmonic components represent spatial structures such as global uniformity and background redundancy. Theoretically, they are residual components that are neither rotational nor divergent, and are expressed as:

[0096]

[0097] In the formula, This is a characteristic diagram of harmonic components. Operator for harmonic component extraction, The harmonic convolution weight matrix is... This is a global average pooling method. This embodiment is for visible light or infrared images, and further elaborates as follows:

[0098]

[0099] In the formula, , These are the harmonic component feature maps corresponding to the visible light and infrared modal images, respectively. , These are the harmonic component extraction operators corresponding to visible light and infrared modal images, respectively. , These are the harmonic convolution weight matrices corresponding to the visible light and infrared modal images, respectively.

[0100] In this embodiment, global average pooling and low-pass convolution kernels with large receptive fields are used for extraction. Initialized with a large mean or Gaussian kernel, it can be trained subsequently; global pooling ensures a consistent response to the global space. This component focuses on the global scene structure and low-frequency background. It should be understood that in other feasible embodiments, it is not necessary to set it to a large mean or Gaussian kernel; this embodiment is a design adopted to enhance the response to low-frequency, background-consistent regions. In principle, as long as the convolutional kernel or subnet structure can effectively extract low-frequency, smooth variation features in the image and weaken edge and local rotation information, the modeling requirements of harmonic components can be met. Therefore, other forms of smoothing kernels or global context modeling structures (such as long-range attention, low-pass filters, etc.) can also be used instead.

[0101] The three components mentioned above approximate the physical gradient component (irrotational, reflecting structural boundaries), curl component (divergence-free, characterizing local rotational changes), and harmonic component (irrotational and divergence-free, capturing global consistency) in Hodge decomposition, respectively. By using structured learnable convolutional weights, end-to-end extraction of spatial physical components is achieved, endowing each branch of the network with explicit physical semantics and structural interpretability. It should be noted that, in obtaining... , and After obtaining the three pairs of Hodge decomposition features for infrared and visible light, basic gradient, curl, and harmonic features are already present, allowing for fusion in step S3. However, to obtain deeper features, each pair of features is often further enhanced using the corresponding component extraction subnet. This allows for the acquisition of more layers of Hodge decomposition fusion results in step S3. This embodiment will first illustrate the technical solution of the present invention using only extraction and one layer as examples.

[0102] Step S3: Component-level structured multimodal fusion

[0103] This phase innovatively proposes a component-level Selective Attention fusion mechanism. Specifically, each type of physical structural component (such as gradient, curl, and harmonics) interacts and fuses only between components of the same type in different modes, thereby achieving precise alignment, complementarity, and enhancement of structural information.

[0104] This embodiment takes gradient components as an example, showing the features after fusing gradient components from different modes. Represented as:

[0105]

[0106] in, The Selective Attention operator, designed for gradient components, is implemented as follows:

[0107]

[0108] , Both are learnable linear transformation matrices, used for the gradient components of the two modes respectively. (Current implementation...) , It is merely a standard fully connected layer weight matrix, used for linear mapping of input features, facilitating cross-modal feature alignment and attention fusion. It should be understood that in the technical solution of this invention... , As long as the basic linear transformation capability, learnability, and dimensionality matching requirements are met, the application needs of this method can be satisfied. Indicates feature splicing; , σ represents the attention mechanism parameter; σ represents non-linear activation (such as Sigmoid or Softmax).

[0109] The curl component and harmonic components use the same structure; only the input features need to be replaced. The mathematical model is expressed as follows:

[0110]

[0111] In the formula, and These are the characteristics of the fused curl components of different modes and the characteristics of the fused harmonic components of different modes, respectively. and These are SelectiveAttention operators designed for the curl and harmonic components, respectively. It should be understood that each SelectiveAttention structure can be set to an independent set of parameters, or they can share some weights to enhance structural consistency. The essence of component-level SelectiveAttention is to perform weighted interaction only on information between similar physical components, avoiding the mixing of features from different physical structures and enhancing structural interpretability. The fusion method is not limited to simple weighted concatenation; depthwise separable convolutions and structure-aware gating can also be used to improve the dynamic modeling capability. This mechanism significantly improves the complementary recognition capability of complex structures in multimodal scenarios, achieving efficient and accurate alignment and enhancement of power component boundaries, rotational deformations, and the global background.

[0112] It should be understood that Selective Attention falls under the category of attention mechanisms. The mathematical model described above differs from existing mainstream forms such as multi-head self-attention, channel attention, spatial attention, gated fusion, or graph attention. Instead, it is a customized fusion structure based on feature concatenation and linear mapping, used to achieve fine-grained alignment and dynamic enhancement of modal features at the physical component level. In other feasible embodiments, the fusion mechanism in step S3 can also employ other attention structures, such as multi-head self-attention, spatial / channel attention, gated fusion, graph convolution, etc., or use operators such as feature concatenation and weighted averaging for information fusion.

[0113] Step S4: Component-aware structural modeling and collaborative segmentation.

[0114] The fused three physical components (gradient, curl, and harmonics) are input into dedicated structure-aware branch networks to further extract and enhance their respective structural representations. Each structure-aware branch can include deep structures such as multi-layer local convolutions, pyramid pooling, and Transformer blocks to capture complex spatial relationships from fine-grained to global scales.

[0115]

[0116] in, , , The following structures are used: stacked local convolutional blocks (such as ResBlock, DilatedConv, etc.) for extracting spatial details and boundary information; spatial pyramid pooling (SPP) and / or pyramid attention (PAM) modules to enhance multi-scale semantic awareness; and structured Transformer modules (such as SwinTransformer, ViT, etc.) to strengthen long-range dependencies and global structure. After passing through the structure-aware branch, the three component features respectively complete high-level semantic aggregation and structure enhancement, fully capturing boundary changes, rotational deformations, and background consistency in complex scenes. It should be understood that deep structures such as multi-layer local convolutions, pyramid pooling, and Transformer blocks, and their connections, can be selectively selected, replaced, or adjusted according to the application object and accuracy requirements.

[0117] Finally, the three types of structure-aware components are weighted and fused according to learnable weights to obtain the final comprehensive structure-aware feature representation:

[0118]

[0119] in, , , All weights are adaptive fusion weights obtained through end-to-end training, which automatically balance the contributions of each structural component under different tasks and scenarios.

[0120] The above-mentioned comprehensive perception features This component-coordinated structure-aware modeling and fusion not only enhances the ability to represent and distinguish various types of physical structural information, but also brings stronger interpretability and scene adaptability to the segmentation model. It serves as input to the segmentation head and is used to output the final pixel-level segmentation mask for electrical components.

[0121] It should be noted that some implementation processes (such as component-level fusion and structure awareness) can be parallelized, cascaded, or fused according to actual engineering needs, and the process order can be flexibly adjusted. For example, in some embodiments, structure awareness is performed first, followed by multimodal fusion. It should also be emphasized that component extraction and physical attribution are always based on Hodge decomposition theory. Therefore, in actual engineering implementation, apart from the core component modeling using Hodge decomposition, other alternative solutions can be implemented and structurally extended around Hodge decomposition, and all fall within the protection scope of this invention.

[0122] Step S5: Segmentation and Optimization

[0123] The fused structure-aware features The input is fed into a segmentation head (such as mainstream segmentation decoder structures like U-Net, DeepLab, and Mask R-CNN), which performs multi-scale upsampling and pixel-level classification, ultimately outputting a segmentation mask for the electrical components. .

[0124] During model training, the loss function not only employs the conventional pixel-level cross-entropy loss. To measure the difference between the predicted segmentation mask and the ground truth label The similarity between them is also used to jointly introduce structure-aware loss. This is specifically designed to constrain the expressive power of fractional features in spatial structure. Ultimately, the joint optimization objective can be expressed as:

[0125]

[0126] Among them, structure-aware loss The component loss is used to ensure that the three structural components of gradient, curl, and harmonics maintain physical consistency and orthogonality in spatial distribution and expression, and to prevent structural aliasing. , These are all weighting coefficients, which can be adjusted according to specific tasks and datasets, balancing pixel-level accuracy with structural component representation capabilities.

[0127] In this embodiment, during the structure perception and decoding stages, a dedicated structure perception branch (using multi-layer convolution, pooling, Transformer, etc.) is designed for each component, enabling each type of physical component to obtain multi-scale structure enhancement from local to global. Finally, the three-component features are fused as the input to the segmentation head, specifically introducing a structure orthogonal loss. and energy attribution loss To maintain the orthogonality between components and the uniqueness of component attribution, specifically:

[0128]

[0129] Ultimately, the structural loss is obtained:

[0130]

[0131] In the formula, , These are the Hodge decomposition features of class iii and class jjj after component-level structured multimodal fusion, respectively, where i and j can be taken as "grad" (gradient component), "curl" (curl component), and "harm" (harmonic component), respectively. The L2 norm (Euclidean norm) is the square of the number of features used to measure the distance or difference between two features. To represent the original overall features of the input image, i.e., the undecomposed, fused overall features before Hodge decomposition-driven feature extraction, structural orthogonal loss is used. and energy attribution loss All of these fall under the category of component losses. It should be understood that structural orthogonality loss, energy attribution loss, structural consistency loss, etc., can be achieved using various metrics such as inner product minimization, cosine similarity, and contrast loss. All loss designs that can effectively constrain the attribution of component structures to physical independence are within the scope of protection of this invention.

[0132] Through these physical mechanisms and loss constraints, the model can achieve spatial attribution and redundancy removal of physical components, ensure the uniqueness and interpretability of structural representation, and effectively suppress noise and feature aliasing.

[0133] This invention can be widely applied to intelligent inspection and component identification scenarios in power systems. Similarly, it can be applied to any segmentation task that uses multimodal images as input and is sensitive to structural information. Taking substation on-site inspection as an example, the system simultaneously acquires visible light and infrared images of power equipment through a multi-sensor platform. After preprocessing such as scaling, cropping, and normalization, a multimodal input with one-to-one spatial pixel correspondence is obtained. Subsequently, using the Hodge decomposition-based fractional feature modeling network described in this invention, the physical structure of the gradient component, curl component, and harmonic component of the input image is decoupled, and multimodal selective attention fusion is achieved at the component level. Finally, the model can accurately identify key structures such as power component boundaries, local deformation of insulators, and wire entanglement in substation scenarios under complex backgrounds and multi-source noise interference, achieving high-precision pixel-level segmentation. In practical applications, the method of this invention demonstrates deep collaboration and adaptive modeling capabilities for intermodal structural information, which can effectively improve the recognition accuracy, structural interpretability, and spatial consistency of power components in harsh environments, providing a reliable structured perception foundation for downstream tasks such as power operation and maintenance, equipment testing, and safety monitoring.

[0134] Example 2

[0135] This invention provides a system based on the above-described component segmentation method, comprising a multimodal data acquisition module, a fractional feature extraction module, a component-level structured multimodal fusion module, a structure-aware component fusion module, and a segmentation output module that are sequentially or interconnected.

[0136] The multimodal data acquisition module is used to acquire multimodal images of the segmented target scene.

[0137] The fractional feature extraction module is used to extract a class of Hodge decomposition features corresponding to each modal image through three parallel gradient component extraction subnetworks, curl component extraction subnetworks, and harmonic component extraction subnetworks. The Hodge decomposition features include gradient components, curl components, and harmonic components.

[0138] The component-level structured multimodal fusion module is used to fuse the same type of Hodge decomposition features of different modes to obtain the fused features corresponding to the gradient component, curl component, and harmonic component, respectively.

[0139] The structure-aware component fusion module is used to perform structure perception on the features of gradient components, curl components, and harmonic components; the structure-aware component fusion module is also used for feature fusion, that is, to fuse the current features to obtain comprehensive structure-aware features.

[0140] The segmentation output module is used to take the comprehensive structural perception features as input to the segmentation head to obtain the segmentation result of the target scene.

[0141] It should also be understood that the specific implementation process of each module is described in the above method. This invention will not repeat it here. The above division of functional modules is only for illustrative purposes. In some embodiments, some functional modules can be combined and some functional modules can be separated. Each functional module can be implemented in software, hardware, or a combination of software and hardware. The software and hardware devices include, but are not limited to, general-purpose computer equipment, programmable gate arrays, digital signal processors, microprocessors and their corresponding programming or burning software.

[0142] Example 3

[0143] This invention provides a computer device, including: one or more processors and a memory storing one or more computer programs;

[0144] The processor invokes the computer program to implement the steps of a component segmentation method based on Hodge decomposition and multimodal fusion. For the specific implementation process of each step, please refer to the foregoing description of the component segmentation method embodiment based on Hodge decomposition and multimodal fusion.

[0145] In some embodiments, such as Figure 2 As shown, the electronic components of a computer device include:

[0146] The processor 1600 can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. The processor 1600 is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention.

[0147] The memory 1700 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 1700 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1700, and the processor 1600 calls and executes the algorithm program of the component segmentation method based on Hodge decomposition and multimodal fusion of the embodiments of this invention.

[0148] The input / output interface 1800 is used to implement information input and output.

[0149] The communication interface 1900 is used to enable communication and interaction between this device and other devices. Communication can be achieved via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0150] Bus 2000 transmits information between various components of the device, such as processor 1600, memory 1700, input / output interface 1800, and communication interface 1900.

[0151] The processor 1600, memory 1700, input / output interface 1800 and communication interface 1900 are connected to each other within the device via bus 2000.

[0152] Example 4

[0153] This invention provides a computer-readable storage medium storing a computer program that is invoked by a processor to implement the steps of a component segmentation method based on Hodge decomposition and multimodal fusion.

[0154] For details on the implementation of each step, please refer to the aforementioned embodiment of the component segmentation method based on Hodge decomposition and multimodal fusion.

[0155] The readable storage medium is a computer-readable storage medium, which can be an internal storage unit of the hardware and software device described in any of the foregoing embodiments, such as the hard drive or memory of the controller. The readable storage medium can also be an external storage device of the controller, such as a plug-in hard drive, Smart MediaCard (SMC), Secure Digital (SD) card, or Flash Card equipped on the controller. Further, the readable storage medium can include both internal storage units and external storage devices of the controller. The readable storage medium is used to store the computer program and other programs and data required by the controller. The readable storage medium can also be used to temporarily store data that has been output or will be output.

[0156] Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned readable storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0157] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. This application refers to flowchart illustrations and / or instructions executed by a processor of a method, apparatus (system), and computer program product according to embodiments of this application to create means for implementing the functions specified in one or more flowchart illustrations and / or one or more block diagrams. These computer program instructions may also be stored in a computer-readable storage medium capable of directing a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowchart illustrations and / or one or more block diagrams. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more blocks of a block diagram.

[0158] It should be emphasized that the examples described in this invention are illustrative rather than limiting. Therefore, this invention is not limited to the examples described in the specific embodiments. Any other embodiments derived by those skilled in the art based on the technical solutions of this invention, without departing from the spirit and scope of this invention, whether modifications or substitutions, are also within the protection scope of this invention.

Claims

1. A component segmentation method based on Hodge decomposition and multimodal fusion, characterized in that: Includes the following steps: Multimodal data acquisition refers to acquiring and segmenting multimodal images of the target scene; Hodge decomposition-driven fractional feature extraction means that for each modal image corresponding to a modality, three parallel sub-networks for gradient component extraction, curl component extraction, and harmonic component extraction are used to extract a class of Hodge decomposition features for each modal image. The Hodge decomposition features include gradient components, curl components, and harmonic components. Component-level structured multimodal fusion and structure awareness; The component-level structured multimodal fusion involves fusing the same type of Hodge decomposition features from different modes to obtain fused features corresponding to gradient components, curl components, and harmonic components; and the structure perception involves performing structure perception on the features of gradient components, curl components, and harmonic components. Feature fusion: fusing current features to obtain comprehensive structural perception features; The segmentation output uses the integrated structural perception features as input to the segmentation head to obtain the component segmentation result of the segmentation target scene; The component segmentation method pre-constructs a sample set of the target scene to be segmented, trains a network model based on the sample set and a designed loss function, and then segments the target scene to be segmented; wherein, in addition to measuring the predicted segmentation mask... With real labels In addition to the loss metric based on similarity between the components, the loss function also includes structure-aware loss. The structure-aware loss is used to constrain the expressive power of quantized features in spatial structure. The component loss is used to ensure that the three structural components of gradient, curl, and harmonics maintain physical consistency and orthogonality in spatial distribution and representation. Structure-aware loss Represented as: ; ; In the formula, For structural orthogonal loss, Assuming energy loss, , These are the Hodge decomposition features of class i and class j after component-level structured multimodal fusion, respectively, where i and j take grad, curl, and harm, respectively, representing the gradient component, curl component, and harmonic component. The L2 norm is used to represent the square of the L2 norm, which measures the distance or difference between two features. To represent the original overall features of the input image, i.e., the undecomposed, fused features before Hodge decomposition-driven feature extraction, , , These represent the features extracted by the structure-aware branch network after the fusion features of the gradient component, curl component, and harmonic component.

2. The component segmentation method according to claim 1, characterized in that: The mathematical model of the gradient component extraction subnet is expressed as follows: ; In the formula, For gradient component feature maps, Operators for extracting gradient components, For gradient convolution weight matrix, These are the bias terms in the mathematical model corresponding to the gradient components; This represents the input image, specifically the corresponding modal image. The modal information is represented by the first convolutional kernel of the gradient component extraction subnet, which is initialized with the structural features corresponding to the focused gradient components as the target. The mathematical model of the curl component extraction subnet is expressed as follows: ; In the formula, This is a characteristic map of the curl component. Operator for curl component extraction, Here is the curl convolution weight matrix. The bias term of the mathematical model corresponding to the curl component is used, wherein the first convolutional kernel of the curl component extraction subnet is initialized with the goal of focusing on the structural features corresponding to the curl component. The mathematical model of the harmonic component extraction subnet is expressed as follows: ; In the formula, This is a characteristic diagram of harmonic components. Operator for harmonic component extraction, The harmonic convolution weight matrix is... For global average pooling, the first layer convolutional kernel of the harmonic component extraction subnet is initialized with the goal of focusing on the structural features corresponding to the harmonic components.

3. The component segmentation method according to claim 1, characterized in that: A component-level SelectiveAttention fusion mechanism is introduced, which uses the SelectiveAttention operator to fuse the same type of Hodge decomposition features of different modalities to obtain the fused features corresponding to the gradient component, curl component, and harmonic component respectively. The mathematical model of the Selective Attention operator is expressed as follows: ; In the formula, This represents the Selective Attention operator for a class of Hodge decomposition features j. and Representing modes and modality The corresponding Hodge decomposition feature j; j is the category label of the Hodge decomposition feature, corresponding to the gradient component, curl component, and harmonic component; , Both are learnable linear transformation matrices, corresponding to the gradient components of the two modes respectively; Indicates feature splicing; , All are attention mechanism parameters; σ represents nonlinear activation.

4. The component segmentation method according to claim 1, characterized in that: If component-level structured multimodal fusion is performed first, followed by structure-aware and feature fusion, the generation process of the comprehensive structure-aware features is as follows: After component-level structured multimodal fusion, the fused features of gradient components, curl components, and harmonic components are input into the structure-aware branch networks corresponding to various Hodge decomposition features for structure perception and feature extraction, represented as follows: ; In the formula, , , These represent the features extracted after the fusion features of gradient components, curl components, and harmonic components are processed by a structure-aware branch network. , , These represent the fusion characteristics of the gradient component, curl component, and harmonic component, respectively. , , These represent structure-aware branch networks; The features extracted by the three types of structure-aware branch networks are then weighted and fused according to learnable weights to obtain the comprehensive structure-aware features, represented as: ; in, , , All are adaptive fusion weights obtained through end-to-end training. This is for the comprehensive structural perception features.

5. The component segmentation method according to claim 1, characterized in that: The modal images are divided into visible light images and infrared images; the components are power components, parts in industrial scenes, roads or buildings in remote sensing images, and organs or lesions in medical images.

6. A system based on the component segmentation method according to any one of claims 1-5, characterized in that: It includes a multimodal data acquisition module, a fractional feature extraction module, a component-level structured multimodal fusion module, a structure-aware component fusion module, and a segmentation output module that are connected sequentially or interconnected. The multimodal data acquisition module is used to acquire multimodal images of the target scene for segmentation. The fractional feature extraction module is used to extract a class of Hodge decomposition features corresponding to each modal image through three parallel gradient component extraction subnetworks, curl component extraction subnetworks, and harmonic component extraction subnetworks for each modal image. The Hodge decomposition features include gradient components, curl components, and harmonic components. The component-level structured multimodal fusion module is used to fuse the same type of Hodge decomposition features of different modes to obtain the fused features corresponding to the gradient component, curl component, and harmonic component, respectively. The structure-aware component fusion module is used to perform structure perception on the features of gradient components, curl components, and harmonic components. The structure-aware component fusion module is also used for feature fusion, that is, to fuse the current features to obtain comprehensive structure-aware features; The segmentation output module is used to take the integrated structural perception features as input to the segmentation head to obtain the component segmentation result of the segmentation target scene.

7. A computer device, characterized in that: include: One or more processors; A memory that stores one or more computer programs; The processor calls the computer program to implement: The steps of the component segmentation method according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that: The computer program is stored and is invoked by the processor to implement: The steps of the component segmentation method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Remote sensing image water body segmentation method based on comparative learning and multi-modal fusion

    CN118864865A

  • A method for detecting OLED screen defects based on optical-magnetic multimodal fusion

    CN119780107A