This invention relates to the fields of intelligent inspection of
power equipment and
computer vision technology, specifically to an open-vocabulary substation equipment segmentation and inspection
system based on multimodal cue learning. The
system includes: an anchoring modeling module, which acquires
structured text of the target scene as an initial text prefix input to a text
encoder and acquires sample images as visual anchors; a collaborative
adaptation module, which acquires a real-time image
stream, extracts multi-scale visual features through a visual
encoder, and inputs them into a visual-language orthogonal
coupling projection layer to obtain target cue parameters; a fitting segmentation module, which calculates the inner product of the visual cue feature
tensor and the text cue feature
tensor to generate a cross-
modal attention heatmap and outputs a pixel-level segmentation
mask; and a closed-loop deployment module, which generates
state recognition results based on the pixel-level segmentation
mask and serializes and stores the target cue parameters as a parameter data
package with attribute labels. This invention can achieve open-vocabulary segmentation under
small sample conditions and reduce the risk of catastrophic forgetting.