Multi-view classification method and device for power transmission line inspection images, terminal equipment and storage medium

By combining image segmentation and multi-view classification models, the problem of background interference in the perspective classification of power transmission line inspection images was solved, achieving higher classification accuracy.

CN120894630APending Publication Date: 2025-11-04GUANGDONG POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511008555.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing technologies are prone to inaccurate classification when classifying images from different perspectives during power transmission line inspections due to interference from complex backgrounds.

Method used

By acquiring inspection images and prompts of the transmission lines to be classified, a feature map and prompt embedding matrix are generated using an image segmentation model. An optimal mask for removing the background is then generated, and the segmented image is input into a multi-view classification model for view classification.

Benefits of technology

It improves the accuracy of perspective classification for power transmission line inspection images, reduces interference from background factors, and enhances the accuracy of classification output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120894630A_ABST
    Figure CN120894630A_ABST
Patent Text Reader

Abstract

The invention discloses a power transmission line inspection image multi-view classification method and device, terminal equipment and a storage medium, and belongs to the technical field of image classification. The method comprises the following steps: inputting a to-be-classified power transmission line inspection image and prompt information into an image segmentation model, the image segmentation model generates a feature map according to the to-be-classified power transmission line inspection image, generates a prompt embedding matrix according to the prompt information, and generates an optimal mask for removing a background image of the to-be-classified power transmission line inspection image according to the feature map and the prompt embedding matrix; generating segmented images according to the optimal mask and the to-be-classified power transmission line inspection images; inputting the segmented image of the to-be-classified power transmission line inspection image into a multi-view classification model, so that the multi-view classification model outputs the view category of the power transmission line inspection image; by implementing the method, the problem that the visual angle classification of the power transmission line inspection image is inaccurate due to the influence of background interference factors during the visual angle classification of the power transmission line inspection image in the prior art can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image classification technology, and in particular to a multi-view classification method, apparatus, terminal equipment, and storage medium for power transmission line inspection images. Background Technology

[0002] With the accelerated construction of the global energy internet and the continuous expansion of transmission line scale, traditional manual inspection faces problems such as low efficiency, high risk, and high cost, making it difficult to meet the needs of modern power grid operation and maintenance. Against this backdrop, intelligent inspection technology using drones and robots has become a core direction for industry transformation. Collecting transmission line inspection images through intelligent inspection technology and extracting inspection information based on these images is a crucial method in the inspection field. However, the perspective differences in transmission line inspection images directly affect the recognizability of equipment features. Front and rear views determine the integrity of the insulator skirt structure, side views reflect the spatial geometric relationship between conductors and towers, and top views are used to analyze the line alignment and the layout of the surrounding environment. Therefore, classifying the perspectives of inspection images before acquiring inspection information is a necessary means to improve the accuracy of the information obtained from the inspection images.

[0003] However, transmission line inspection images are affected by complex backgrounds when classified by perspective. In natural scenes, transmission lines coexist with similar line targets (such as roads and rivers) and irregular textures (such as vegetation shadows). Existing methods directly classify perspectives based on the captured transmission line inspection images containing the background, which is easily affected by background interference and leads to inaccurate perspective classification. Summary of the Invention

[0004] This invention provides a multi-view classification method, apparatus, terminal equipment, and storage medium for transmission line inspection images. It can effectively solve the problem of inaccurate view classification of transmission line inspection images caused by background interference factors in the prior art, and improve the view classification accuracy of transmission line inspection images.

[0005] An embodiment of the present invention provides a multi-view classification method for transmission line inspection images, including:

[0006] Acquire inspection images and alerts for transmission lines to be classified;

[0007] The inspection image of the transmission line to be classified and the prompt information are input into the image segmentation model, so that the image segmentation model generates a feature map based on the inspection image of the transmission line to be classified, generates a prompt embedding matrix based on the prompt information, generates an optimal mask to remove the background image of the inspection image of the transmission line to be classified based on the feature map and the prompt embedding matrix, and generates a segmented image of the inspection image of the transmission line to be classified based on the optimal mask and the inspection image of the transmission line to be classified.

[0008] The segmented image of the transmission line inspection image to be classified is input into the multi-view classification model so that the multi-view classification model outputs the view category of the transmission line inspection image; wherein, the view category includes: front and back angle, side view angle and top view angle.

[0009] Furthermore, the construction of the image segmentation model includes:

[0010] Obtain several training samples; wherein each training sample includes a first transmission line inspection image sample, a prompt information sample, and a real segmented image of the first transmission line inspection image sample;

[0011] Construct an initial image segmentation model; wherein, the initial image segmentation model includes: an image encoder, a cue encoder, and a mask decoder;

[0012] The initial image segmentation model is iteratively trained using each training sample until the initial image segmentation model converges, at which point the image segmentation model is generated.

[0013] In each iteration of training, the image encoder scales the first transmission line inspection image sample proportionally to a preset number of pixels to obtain the second transmission line inspection image sample; the second transmission line inspection image sample is then divided into image blocks and linearly projected to obtain linear projection data; the linear projection data is encoded to obtain the initial feature map of the first transmission line inspection image sample; and the initial feature map is then subjected to channel dimensionality reduction processing to obtain the feature map of the first transmission line inspection image sample.

[0014] The prompt encoder determines the prompt type based on the prompt information sample, and obtains all coordinate point data and the point type of each coordinate point within the prompt area of ​​the target transmission line in the first transmission line inspection image sample according to the prompt type. It then generates point set data for the first transmission line inspection image sample based on the coordinate point data and the point type of each coordinate point within the prompt area. For each coordinate point in the point set data, it generates a position vector based on the coordinate point data, a type vector based on the coordinate point type, a fusion vector based on the position vector and type vector, and a prompt embedding matrix for the first transmission line inspection image sample based on the fusion vectors of each coordinate point. The prompt types include point prompts and text prompts, and the point types include foreground coordinate point types and background coordinate point types.

[0015] The mask decoder generates the optimal mask for the first transmission line inspection image sample based on the prompt embedding matrix, the feature map of the first transmission line inspection image sample, and the preset output token. It then generates a predicted segmentation image of the first transmission line inspection image sample based on the optimal mask and the first transmission line inspection image sample.

[0016] The image segmentation loss under the current iteration training is determined based on the predicted segmentation image of the first transmission line inspection image sample and the actual segmentation image of the first transmission line inspection image sample.

[0017] Furthermore, the construction of the multi-view classification model includes:

[0018] Obtain the viewpoint category of the predicted segmented image of each first transmission line inspection image sample;

[0019] Construct an initial multi-view classification model;

[0020] The initial multi-view classification model is trained by using the predicted segmented image of the first transmission line inspection image sample as the input and the view category of the predicted segmented image of the first transmission line inspection image sample as the output. The multi-view classification model is generated when the initial multi-view classification model converges.

[0021] Furthermore, the loss function of the image segmentation model is specifically as follows:

[0022]

[0023] Among them, L_ seg This represents the image segmentation loss of the image segmentation model; This represents the predicted segmented image from the image segmentation model. This represents the true segmented image of the image segmentation model; α is the balance coefficient.

[0024] Furthermore, the loss function of the multi-view classification model is specifically as follows:

[0025]

[0026] Among them, L_ cls p represents the classification loss of a multi-view classification model. c y represents the predicted probability of the multi-view classification model for view category c; c This represents the one-hot encoding of the true view category of view category c.

[0027] Based on the above method embodiments, the present invention provides corresponding apparatus embodiments;

[0028] One embodiment of the present invention provides a multi-view classification device for transmission line inspection images, including: a data acquisition module, an image segmentation module, and a view classification module;

[0029] The data acquisition module is used to acquire inspection images and prompt information of the transmission lines to be classified;

[0030] The image segmentation module is used to input the transmission line inspection image to be classified and the prompt information into the image segmentation model, so that the image segmentation model generates a feature map based on the transmission line inspection image to be classified, generates a prompt embedding matrix based on the prompt information, generates an optimal mask to remove the background image of the transmission line inspection image to be classified based on the feature map and the prompt embedding matrix, and generates a segmented image of the transmission line inspection image to be classified based on the optimal mask and the transmission line inspection image to be classified.

[0031] The view classification module is used to input the segmented image of the transmission line inspection image to be classified into the multi-view classification model, so that the multi-view classification model outputs the view category of the transmission line inspection image; wherein, the view category includes: front and back angle, side view angle and top view angle.

[0032] Furthermore, it also includes: an image segmentation model construction module;

[0033] The image segmentation model construction module is used to acquire several training samples; wherein, each training sample includes a first transmission line inspection image sample, a prompt information sample, and a real segmented image of the first transmission line inspection image sample;

[0034] Construct an initial image segmentation model; wherein, the initial image segmentation model includes: an image encoder, a cue encoder, and a mask decoder;

[0035] The initial image segmentation model is iteratively trained using each training sample until the initial image segmentation model converges, at which point the image segmentation model is generated.

[0036] In each iteration of training, the image encoder scales the first transmission line inspection image sample proportionally to a preset number of pixels to obtain the second transmission line inspection image sample; the second transmission line inspection image sample is then divided into image blocks and linearly projected to obtain linear projection data; the linear projection data is encoded to obtain the initial feature map of the first transmission line inspection image sample; and the initial feature map is then subjected to channel dimensionality reduction processing to obtain the feature map of the first transmission line inspection image sample.

[0037] The prompt encoder determines the prompt type based on the prompt information sample, and obtains all coordinate point data and the point type of each coordinate point within the prompt area of ​​the target transmission line in the first transmission line inspection image sample according to the prompt type. It then generates point set data for the first transmission line inspection image sample based on the coordinate point data and the point type of each coordinate point within the prompt area. For each coordinate point in the point set data, it generates a position vector based on the coordinate point data, a type vector based on the coordinate point type, a fusion vector based on the position vector and type vector, and a prompt embedding matrix for the first transmission line inspection image sample based on the fusion vectors of each coordinate point. The prompt types include point prompts and text prompts, and the point types include foreground coordinate point types and background coordinate point types.

[0038] The mask decoder generates the optimal mask for the first transmission line inspection image sample based on the prompt embedding matrix, the feature map of the first transmission line inspection image sample, and the preset output token. It then generates a predicted segmentation image of the first transmission line inspection image sample based on the optimal mask and the first transmission line inspection image sample.

[0039] The image segmentation loss under the current iteration training is determined based on the predicted segmentation image of the first transmission line inspection image sample and the actual segmentation image of the first transmission line inspection image sample.

[0040] Furthermore, it also includes: a multi-view classification model construction module;

[0041] The multi-view classification model construction module is used to obtain the view category of the predicted segmented image of each first transmission line inspection image sample;

[0042] Construct an initial multi-view classification model;

[0043] The initial multi-view classification model is trained by using the predicted segmented image of the first transmission line inspection image sample as the input and the view category of the predicted segmented image of the first transmission line inspection image sample as the output. The multi-view classification model is generated when the initial multi-view classification model converges.

[0044] Another embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements a multi-view classification method for transmission line inspection images as described in the above-described embodiment of the invention.

[0045] Another embodiment of the present invention provides a storage medium including a stored computer program, wherein, when the computer program is running, it controls the device where the storage medium is located to execute the multi-view classification method for transmission line inspection images described in the above-described embodiment of the invention.

[0046] The following benefits can be obtained by implementing the present invention:

[0047] This invention provides a multi-view classification method, apparatus, terminal equipment, and storage medium for transmission line inspection images. The method acquires the transmission line inspection image to be classified and prompt information, processes the image to be classified using an image segmentation model, generates a segmented image of the transmission line inspection image to be classified after removing the background image, and then inputs it into the multi-view classification model for perspective classification. This makes the classification in the multi-view classification model unaffected by background factors, thereby improving the accuracy of the classification output. Attached Figure Description

[0048] Figure 1 This is a flowchart illustrating a multi-view classification method for transmission line inspection images provided in an embodiment of the present invention.

[0049] Figure 2 This is a schematic diagram of the original image of a power transmission line inspection provided in an embodiment of the present invention.

[0050] Figure 3 This is a schematic diagram of a segmented image of a power transmission line inspection image provided in an embodiment of the present invention.

[0051] Figure 4 This is a schematic diagram of the structure of a multi-view classification device for power transmission line inspection images provided in an embodiment of the present invention. Detailed Implementation

[0052] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0053] like Figure 1 As shown, to address the problem of inaccurate perspective classification of transmission line inspection images due to background interference factors in existing technologies, an embodiment of the present invention provides a multi-view classification method for transmission line inspection images, including:

[0054] Step S1: Obtain inspection images and prompts for the transmission lines to be classified;

[0055] Step S2: Input the inspection image of the transmission line to be classified and the prompt information into the image segmentation model, so that the image segmentation model generates a feature map based on the inspection image of the transmission line to be classified, generates a prompt embedding matrix based on the prompt information, generates an optimal mask to remove the background image of the inspection image of the transmission line to be classified based on the feature map and the prompt embedding matrix, and generates a segmented image of the inspection image of the transmission line to be classified based on the optimal mask and the inspection image of the transmission line to be classified.

[0056] Step S3: Input the segmented image of the transmission line inspection image to be classified into the multi-view classification model so that the multi-view classification model outputs the view category of the transmission line inspection image; wherein, the view category includes: front and back angle, side view angle and top view angle.

[0057] For step S1, obtain the image of the transmission line to be classified and the prompt information. The prompt information can be the coordinate information of a point obtained by the user clicking on a location in the image of the transmission line to be classified in the human-computer interaction interface, the coordinate set of points obtained by the user selecting a location in the image of the transmission line to be classified in the human-computer interaction interface, or the text prompt information entered by the user in the human-computer interaction interface.

[0058] For step S2, the obtained transmission line inspection image to be classified and the prompt information are input into the image segmentation model, so that the image segmentation model generates a feature map based on the transmission line inspection image to be classified, generates a prompt embedding matrix based on the prompt information, generates an optimal mask to remove the background image of the transmission line inspection image to be classified based on the feature map and the prompt embedding matrix, and generates a segmented image of the transmission line inspection image to be classified based on the optimal mask and the transmission line inspection image to be classified.

[0059] In a preferred embodiment, the construction of the image segmentation model includes:

[0060] Obtain several training samples; wherein each training sample includes a first transmission line inspection image sample, a prompt information sample, and a real segmented image of the first transmission line inspection image sample;

[0061] Construct an initial image segmentation model; wherein, the initial image segmentation model includes: an image encoder, a cue encoder, and a mask decoder;

[0062] The initial image segmentation model is iteratively trained using each training sample until the initial image segmentation model converges, at which point the image segmentation model is generated.

[0063] In each iteration of training, the image encoder scales the first transmission line inspection image sample proportionally to a preset number of pixels to obtain the second transmission line inspection image sample; the second transmission line inspection image sample is then divided into image blocks and linearly projected to obtain linear projection data; the linear projection data is encoded to obtain the initial feature map of the first transmission line inspection image sample; and the initial feature map is then subjected to channel dimensionality reduction processing to obtain the feature map of the first transmission line inspection image sample.

[0064] The prompt encoder determines the prompt type based on the prompt information sample, and obtains all coordinate point data and the point type of each coordinate point within the prompt area of ​​the target transmission line in the first transmission line inspection image sample according to the prompt type. It then generates point set data for the first transmission line inspection image sample based on the coordinate point data and the point type of each coordinate point within the prompt area. For each coordinate point in the point set data, it generates a position vector based on the coordinate point data, a type vector based on the coordinate point type, a fusion vector based on the position vector and type vector, and a prompt embedding matrix for the first transmission line inspection image sample based on the fusion vectors of each coordinate point. The prompt types include point prompts and text prompts, and the point types include foreground coordinate point types and background coordinate point types.

[0065] The mask decoder generates the optimal mask for the first transmission line inspection image sample based on the prompt embedding matrix, the feature map of the first transmission line inspection image sample, and the preset output token. It then generates a predicted segmentation image of the first transmission line inspection image sample based on the optimal mask and the first transmission line inspection image sample.

[0066] The image segmentation loss under the current iteration training is determined based on the predicted segmentation image of the first transmission line inspection image sample and the actual segmentation image of the first transmission line inspection image sample.

[0067] Specifically, such as Figure 2 The image shown is a schematic diagram of power transmission line inspection images captured by a drone during a power transmission line inspection, as provided by this invention. Before training the initial image segmentation model, the Labelme annotation tool is used to annotate the foreground of the power transmission line inspection image (i.e., data related to the power transmission line itself in the power transmission line inspection image excluding the background). The annotation results are used as the true segmented images of the first power transmission line inspection image sample, and the power transmission line inspection image carrying the annotation results is used as the first power transmission line inspection image sample. Simultaneously, corresponding prompt information samples are obtained. Training samples are constructed based on the first power transmission line inspection image sample, the prompt information sample, and the true segmented images of the first power transmission line inspection image sample.

[0068] Preferably, to reduce the number of samples required, the first transmission line inspection image sample in the training samples can be randomly rotated by ±15°, horizontally flipped, or randomly cropped by 5% to simulate the viewpoint shift caused by changes in flight attitude during UAV inspection, thus obtaining an augmented image sample, which is then added to the training samples. Preferably, to enhance the robustness of the image segmentation model to changes in illumination, the augmented image sample can also be adjusted by adjusting brightness (±0.2), contrast (±0.15), and adding Gaussian noise (σ=0.01) when applying ±15° rotation, horizontal flipping, or 5% random cropping to the training samples.

[0069] After preparing the training data, an initial image segmentation model is constructed. In this invention, the initial image segmentation model is an encoder-decoder structure. The encoder part includes an image encoder and a cue encoder, and the mask part is a mask decoder. The image encoder is the fundamental component of the initial image segmentation model, responsible for converting the input image into high-dimensional features, i.e., feature map representation, providing global visual information for subsequent cue processing and mask generation. The cue encoder is used to encode different types of cue information into a unified dimensional vector, whose dimension is consistent with the dimension of the feature map output by the image encoder, to facilitate subsequent fusion.

[0070] After the initial image segmentation model was constructed, iterative training was performed on the initial image segmentation model using the prepared training data. During each iteration of training, the following processing operations were performed on each structure in the initial image segmentation model.

[0071] For the image encoder, the input is a first power line inspection image sample. After receiving the first power line inspection image sample, the image encoder first performs image preprocessing on the first power line inspection image sample, scaling the long side of the preprocessed first power line inspection image sample proportionally and scaling the short side to 1024 pixels (the preset pixel size in this invention is 1024 pixels). During the scaling process, if it is necessary to enlarge the first power line inspection image sample, black pixels are filled into the first power line inspection image sample so that the enlarged pixel size is a second power line inspection image sample of 1024x1024 pixels. The Vision Transformer (ViT) structure in the image encoder is used, specifically the ViT-H / 16 (ViT-Huge, patch size 16×16) structure, to divide the second power line inspection image sample into blocks, dividing the second power line inspection image sample into 16×16 small blocks. Then, linear projection is performed on the block-shaped second power line inspection image sample to obtain linear projection data that can be used for Transformer encoding. After Transformer encoding of the line projection data, an initial feature map of the first transmission line inspection image sample is obtained, with a size of 64×64×1024. Channel dimensionality reduction is then performed on the initial feature map. A 1×1 convolution reduces the 1024-pixel channel to 256-pixel channels, resulting in the feature map of the first transmission line inspection image sample. A 3×3 convolution then maintains the feature map at 256 channels, resulting in a final feature map size of 64×64×256 for the first transmission line inspection image sample.

[0072] For the prompt encoder, the input is a prompt information sample corresponding to the first transmission line inspection image sample. The prompt type is determined based on the prompt information sample. Prompt types include point prompts and text prompts. Point prompts obtain the corresponding coordinate point data based on the area clicked or selected by the user. Text prompts obtain the coordinate point data of the area corresponding to the text content described in the first transmission line inspection image sample, parsed from the user's input text. Simultaneously, for each coordinate point, its point type (foreground or background) is determined based on the annotation information of the first transmission line inspection image sample. For each coordinate point, a type vector is generated based on its point type. Position encoding and Fourier transform are performed on the coordinate point data to generate a position vector. Then, vector fusion is performed on the position vector and type vector to obtain the fused vector for the current coordinate point. This vector fusion process ensures that the fused vector simultaneously contains the geometric position and interaction intent information of the coordinate point, while maintaining the spatial correspondence of the original input between coordinate points. An n×256 prompt embedding matrix is ​​generated based on the fused vectors of all coordinate points.

[0073] For the mask decoder, it receives a feature map (64×64×256) output from the image encoder and a cue embedding matrix (n×256) output from the cue encoder. Then, based on the feature map (64×64×256), the cue embedding matrix (n×256), and the three pre-stored learnable output tokens (3×256) in the mask decoder, it processes this data through a first-layer Transformer. The first-layer Transformer concatenates the cue embedding matrix with the output tokens to obtain a (n+3×256) feature sequence, capturing the dependencies between cues and the potential associations of the learnable output tokens. This first-layer normalized output semantically aligned cue embedding matrix-token features are then fed to a second-layer Transformer. The second-layer Transformer fuses the cue embedding matrix-token features with the feature map through cross-attention, establishing a mapping between visual features and cue semantic features. Finally, a second normalization layer generates three learnable tokens (3×256) containing cue guidance. The three learnable tokens (3×256) obtained from the second-layer Transformer are used to generate three sets of convolutional kernel parameters through a dynamic masking mechanism. The image features upsampled to 256×256 are then convolved channel-by-channel to output three candidate mask probability maps (256×256×1), achieving spatial detail restoration of the mask. Finally, the IoU prediction head (Intersection over Union Prediction Head) evaluates the intersection-over-union ratio (IoU) of each candidate mask probability map with the three learnable tokens obtained from the second-layer Transformer, obtaining the mask quality score of each candidate mask map. After sorting by mask quality score, the candidate mask map with the highest wind speed is selected as the optimal mask for the first transmission line inspection image sample.

[0074] Finally, the optimal mask of the first transmission line inspection image sample is fused with the first transmission line inspection image sample at the pixel level to obtain the predicted segmentation image of the first transmission line inspection image sample. Specifically, firstly, the optimal mask probability map (256×256×1) selected from the IoU prediction head needs to be converted into a binary mask through thresholding. A fixed threshold method is adopted, with 0.7 set as the default threshold. Pixels with a probability value ≥0.7 are marked as foreground (1), and the rest are marked as background (0). Secondly, the binary mask is restored to the same size as the first transmission line inspection image sample through upsampling technology. At the same time, bilinear interpolation is used to improve the image resolution when the binary mask is restored to the same size as the first transmission line inspection image sample by linearly weighting the neighbor pixel values. Finally, the processed binary mask is multiplied pixel by pixel (masking operation) with the first transmission line inspection image sample, where the foreground pixels retain the original color and the background pixels are set to black. The predicted segmentation image obtained after processing is as follows. Figure 3 As shown.

[0075] Preferably, to improve the visual quality of the predicted segmentation image, the following post-processing steps can be introduced: Morphological operations: Optimize mask edges through dilation (enlarging the foreground region), erosion (removing noise), opening operations (erosion followed by dilation to eliminate small regions), or closing operations (dilation followed by erosion to fill holes). Connected component analysis: Identify and remove isolated regions with excessively small areas (such as noise points), or merge adjacent connected components to repair broken target boundaries. Edge enhancement: Sharpen mask edges using convolution kernels (3×3 Laplacian operators) to improve the distinction between the target and the background. After the above optimization processing, the predicted segmentation image can be presented in the following ways: Direct overlay: Overlay the binary mask on the original image in a semi-transparent form to visually display the segmented region. Boundary drawing: Draw the mask outline on the original image to highlight the target edges.

[0076] For step S3, the segmented image of the transmission line inspection image to be classified is input into the multi-view classification model so that the multi-view classification model outputs the view category of the transmission line inspection image; wherein, the view category includes: front and back angle, side view angle and top view angle.

[0077] In a preferred embodiment, the construction of the multi-view classification model includes:

[0078] Obtain the viewpoint category of the predicted segmented image of each first transmission line inspection image sample;

[0079] Construct an initial multi-view classification model;

[0080] The initial multi-view classification model is trained by using the predicted segmented image of the first transmission line inspection image sample as the input and the view category of the predicted segmented image of the first transmission line inspection image sample as the output. The multi-view classification model is generated when the initial multi-view classification model converges.

[0081] Specifically, in this invention, the network backbone of the initial multi-view classification model adopts EfficientNet-B4. The initial multi-view classification model includes a Stem layer (backbone initiation layer) and an improved MBConv (MobileInverted Bottleneck Convolution) module. The improved MBConv module includes a single MBConv module with a kernel of 1 and six MBConv modules with a kernel of 6 connected sequentially.

[0082] It should be noted that the core components of the MBConv module include depthwise separable convolution, inverted residual structure, and SE attention mechanism. Among them, depthwise separable convolution is used to decompose standard convolution into depthwise convolution and pointwise convolution, which greatly reduces the number of parameters. Inverted residual structure is used to enhance feature representation ability while maintaining computational efficiency by first increasing the dimensionality and then decreasing the dimensionality. SE attention mechanism is used to adaptively adjust the channel weights through global average pooling and fully connected layers to highlight features related to viewpoint classification (such as the direction of insulator skirts and the spatial layout of conductors).

[0083] In each iteration of the initial multi-view classification model training process, the predicted segmentation image of the first transmission line inspection image sample is input into the Stem layer. The Stem layer uses a 3×3 convolution kernel to extract the basic classification features of the view category in the predicted segmentation image of the first transmission line inspection image sample, and outputs a 48-channel feature map containing the basic classification features to the improved MBConv module. In the improved MBConv module, each MBConv module extracts classification features sequentially. Specifically, the features extracted by the fourth MBConv module with a kernel of 6 (mid-level classification features, 32×32×80) and the features extracted by the seventh MBConv module with a kernel of 6 (high-level classification features, 2×2×320) are fused through skip connections to combine local details and global semantics, resulting in a 32×32×80 feature map. During fusion, bilinear interpolation is used to upsample the features extracted by the seventh MBConv module with a kernel of 6 from 2×2 to 32×32, and a 1×1 convolution is used to reduce the number of channels from 320 to 80, ensuring that the size and number of channels of the mid-level and high-level classification features are perfectly aligned.

[0084] The fused 32×32×80 feature map is output to the classification head, which then outputs the predicted viewpoint category. The classification head consists of a global pooling layer, a Dropout layer, and a fully connected layer connected sequentially. Specifically, the global average pooling layer compresses the fused 32×32×80 feature map into a 1×1×80 global feature vector, preserving global statistical information of multi-scale features; the Dropout layer randomly discards neurons with a probability of 0.4 to prevent overfitting; and the fully connected layer maps the 80-dimensional features to three viewpoint categories (front / back angle, side view angle, and top view angle), outputting a probability distribution through the Softmax function to achieve refined viewpoint category classification.

[0085] In this invention, the image segmentation model uses a weighted sum of Dice loss and BCE (Binary Cross-Entropy) loss to ensure high-precision segmentation of the foreground mask.

[0086] In a preferred embodiment, the loss function of the image segmentation model is specifically:

[0087]

[0088] Among them, L_ seg This represents the image segmentation loss of the image segmentation model; This represents the predicted segmented image from the image segmentation model. This represents the true segmented image of the image segmentation model; α is the balance coefficient.

[0089] In this invention, the loss function of the multi-view classification model adopts the standard cross-entropy loss, which directly optimizes the category prediction probability of multi-class problems.

[0090] In a preferred embodiment, the loss function of the multi-view classification model is specifically:

[0091]

[0092] Among them, L_ cls p represents the classification loss of a multi-view classification model. c y represents the predicted probability of the multi-view classification model for view category c; c This represents the one-hot encoding of the true view category of view category c.

[0093] The joint loss function of the image segmentation model and the multi-view classification model can be expressed as:

[0094] L total =L_ seg +βL_ cls

[0095] Among them, L total β represents the joint loss of the image segmentation model and the multi-view classification model; β represents the classification loss weight of the multi-view classification model, which is determined through hyperparameter tuning to ensure the coordinated optimization of segmentation and classification tasks.

[0096] Based on the above method embodiments, the present invention provides corresponding apparatus embodiments.

[0097] like Figure 4 As shown, an embodiment of the present invention provides a multi-view classification device for transmission line inspection images, including: a data acquisition module, an image segmentation module, and a view classification module;

[0098] The data acquisition module is used to acquire inspection images and prompt information of the transmission lines to be classified;

[0099] The image segmentation module is used to input the transmission line inspection image to be classified and the prompt information into the image segmentation model, so that the image segmentation model generates a feature map based on the transmission line inspection image to be classified, generates a prompt embedding matrix based on the prompt information, generates an optimal mask to remove the background image of the transmission line inspection image to be classified based on the feature map and the prompt embedding matrix, and generates a segmented image of the transmission line inspection image to be classified based on the optimal mask and the transmission line inspection image to be classified.

[0100] The view classification module is used to input the segmented image of the transmission line inspection image to be classified into the multi-view classification model, so that the multi-view classification model outputs the view category of the transmission line inspection image; wherein, the view category includes: front and back angle, side view angle and top view angle.

[0101] In a preferred embodiment, it further includes: an image segmentation model construction module;

[0102] The image segmentation model construction module is used to acquire several training samples; wherein, each training sample includes a first transmission line inspection image sample, a prompt information sample, and a real segmented image of the first transmission line inspection image sample;

[0103] Construct an initial image segmentation model; wherein, the initial image segmentation model includes: an image encoder, a cue encoder, and a mask decoder;

[0104] The initial image segmentation model is iteratively trained using each training sample until the initial image segmentation model converges, at which point the image segmentation model is generated.

[0105] In each iteration of training, the image encoder scales the first transmission line inspection image sample proportionally to a preset number of pixels to obtain the second transmission line inspection image sample; the second transmission line inspection image sample is then divided into image blocks and linearly projected to obtain linear projection data; the linear projection data is encoded to obtain the initial feature map of the first transmission line inspection image sample; and the initial feature map is then subjected to channel dimensionality reduction processing to obtain the feature map of the first transmission line inspection image sample.

[0106] The prompt encoder determines the prompt type based on the prompt information sample, and obtains all coordinate point data and the point type of each coordinate point within the prompt area of ​​the target transmission line in the first transmission line inspection image sample according to the prompt type. It then generates point set data for the first transmission line inspection image sample based on the coordinate point data and the point type of each coordinate point within the prompt area. For each coordinate point in the point set data, it generates a position vector based on the coordinate point data, a type vector based on the coordinate point type, a fusion vector based on the position vector and type vector, and a prompt embedding matrix for the first transmission line inspection image sample based on the fusion vectors of each coordinate point. The prompt types include point prompts and text prompts, and the point types include foreground coordinate point types and background coordinate point types.

[0107] The mask decoder generates the optimal mask for the first transmission line inspection image sample based on the prompt embedding matrix, the feature map of the first transmission line inspection image sample, and the preset output token. It then generates a predicted segmentation image of the first transmission line inspection image sample based on the optimal mask and the first transmission line inspection image sample.

[0108] The image segmentation loss under the current iteration training is determined based on the predicted segmentation image of the first transmission line inspection image sample and the actual segmentation image of the first transmission line inspection image sample.

[0109] In a preferred embodiment, it further includes: a multi-view classification model construction module;

[0110] The multi-view classification model construction module is used to obtain the view category of the predicted segmented image of each first transmission line inspection image sample;

[0111] Construct an initial multi-view classification model;

[0112] The initial multi-view classification model is trained by using the predicted segmented image of the first transmission line inspection image sample as the input and the view category of the predicted segmented image of the first transmission line inspection image sample as the output. The multi-view classification model is generated when the initial multi-view classification model converges.

[0113] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.

[0114] Those skilled in the art will clearly understand that, for convenience and brevity, the specific working process of the device described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0115] Based on the above method embodiments, the present invention provides corresponding terminal device embodiments.

[0116] One embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements a multi-view classification method for transmission line inspection images as described in any one of the present invention.

[0117] The terminal device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.

[0118] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the terminal device, connecting all parts of the terminal device via various interfaces and lines.

[0119] The memory can be used to store the computer program. The processor implements various functions of the terminal device by running or executing the computer program stored in the memory and calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function, etc.; the data storage area may store data created based on the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, RAM, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0120] Based on the above method embodiments, the present invention provides corresponding storage medium embodiments.

[0121] One embodiment of the present invention provides a storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the storage medium is located to execute a multi-view classification method for transmission line inspection images as described in any one of the present invention.

[0122] The storage medium is a computer-readable storage medium, and the computer program is stored in the computer-readable storage medium. When executed by a processor, the computer program can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0123] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A multi-view classification method for transmission line inspection images, characterized in that, include: Acquire inspection images and alerts for transmission lines to be classified; The inspection image of the transmission line to be classified and the prompt information are input into the image segmentation model, so that the image segmentation model generates a feature map based on the inspection image of the transmission line to be classified, generates a prompt embedding matrix based on the prompt information, generates an optimal mask to remove the background image of the inspection image of the transmission line to be classified based on the feature map and the prompt embedding matrix, and generates a segmented image of the inspection image of the transmission line to be classified based on the optimal mask and the inspection image of the transmission line to be classified. The segmented image of the transmission line inspection image to be classified is input into the multi-view classification model so that the multi-view classification model outputs the view category of the transmission line inspection image; wherein, the view category includes: front and back angle, side view angle and top view angle.

2. The multi-view classification method for transmission line inspection images as described in claim 1, characterized in that, The construction of the image segmentation model includes: Obtain several training samples; wherein each training sample includes a first transmission line inspection image sample, a prompt information sample, and a real segmented image of the first transmission line inspection image sample; Construct an initial image segmentation model; wherein, the initial image segmentation model includes: an image encoder, a cue encoder, and a mask decoder; The initial image segmentation model is iteratively trained using each training sample until the initial image segmentation model converges, at which point the image segmentation model is generated. In each iteration of training, the image encoder scales the first transmission line inspection image sample proportionally to a preset number of pixels to obtain the second transmission line inspection image sample; the second transmission line inspection image sample is then divided into image blocks and linearly projected to obtain linear projection data; the linear projection data is encoded to obtain the initial feature map of the first transmission line inspection image sample; and the initial feature map is then subjected to channel dimensionality reduction processing to obtain the feature map of the first transmission line inspection image sample. The prompt encoder determines the prompt type based on the prompt information sample, and obtains all coordinate point data and the point type of each coordinate point within the prompt area of ​​the target transmission line in the first transmission line inspection image sample according to the prompt type. It then generates point set data for the first transmission line inspection image sample based on the coordinate point data and the point type of each coordinate point within the prompt area. For each coordinate point in the point set data, it generates a position vector based on the coordinate point data, a type vector based on the coordinate point type, a fusion vector based on the position vector and type vector, and a prompt embedding matrix for the first transmission line inspection image sample based on the fusion vectors of each coordinate point. The prompt types include point prompts and text prompts, and the point types include foreground coordinate point types and background coordinate point types. The mask decoder generates the optimal mask for the first transmission line inspection image sample based on the prompt embedding matrix, the feature map of the first transmission line inspection image sample, and the preset output token. It then generates a predicted segmentation image of the first transmission line inspection image sample based on the optimal mask and the first transmission line inspection image sample. The image segmentation loss under the current iteration training is determined based on the predicted segmentation image of the first transmission line inspection image sample and the actual segmentation image of the first transmission line inspection image sample.

3. The multi-view classification method for transmission line inspection images as described in claim 2, characterized in that, The construction of the multi-view classification model includes: Obtain the viewpoint category of the predicted segmented image of each first transmission line inspection image sample; Construct an initial multi-view classification model; The initial multi-view classification model is trained by using the predicted segmented image of the first transmission line inspection image sample as the input and the view category of the predicted segmented image of the first transmission line inspection image sample as the output. The multi-view classification model is generated when the initial multi-view classification model converges.

4. The multi-view classification method for transmission line inspection images as described in claim 3, characterized in that, The loss function of the image segmentation model is as follows: Among them, L_ seg This represents the image segmentation loss of the image segmentation model; This represents the predicted segmented image from the image segmentation model. This represents the true segmented image of the image segmentation model; α is the balance coefficient.

5. The multi-view classification method for transmission line inspection images as described in claim 4, characterized in that, The loss function of the multi-view classification model is as follows: Among them, L_ cls p represents the classification loss of a multi-view classification model. c y represents the predicted probability of the multi-view classification model for view category c; c This represents the one-hot encoding of the true view category of view category c.

6. A multi-view classification device for transmission line inspection images, characterized in that, include: Data acquisition module, image segmentation module, and viewpoint classification module; The data acquisition module is used to acquire inspection images and prompt information of the transmission lines to be classified; The image segmentation module is used to input the transmission line inspection image to be classified and the prompt information into the image segmentation model, so that the image segmentation model generates a feature map based on the transmission line inspection image to be classified, generates a prompt embedding matrix based on the prompt information, generates an optimal mask to remove the background image of the transmission line inspection image to be classified based on the feature map and the prompt embedding matrix, and generates a segmented image of the transmission line inspection image to be classified based on the optimal mask and the transmission line inspection image to be classified. The view classification module is used to input the segmented image of the transmission line inspection image to be classified into the multi-view classification model, so that the multi-view classification model outputs the view category of the transmission line inspection image; wherein, the view category includes: front and back angle, side view angle and top view angle.

7. The multi-view classification device for transmission line inspection images as described in claim 6, characterized in that, Also includes: Image segmentation model construction module; The image segmentation model construction module is used to acquire several training samples; wherein, each training sample includes a first transmission line inspection image sample, a prompt information sample, and a real segmented image of the first transmission line inspection image sample; Construct an initial image segmentation model; wherein, the initial image segmentation model includes: an image encoder, a cue encoder, and a mask decoder; The initial image segmentation model is iteratively trained using each training sample until the initial image segmentation model converges, at which point the image segmentation model is generated. In each iteration of training, the image encoder scales the first transmission line inspection image sample proportionally to a preset number of pixels to obtain the second transmission line inspection image sample; the second transmission line inspection image sample is then divided into image blocks and linearly projected to obtain linear projection data; the linear projection data is encoded to obtain the initial feature map of the first transmission line inspection image sample; and the initial feature map is then subjected to channel dimensionality reduction processing to obtain the feature map of the first transmission line inspection image sample. The prompt encoder determines the prompt type based on the prompt information sample, and obtains all coordinate point data and the point type of each coordinate point within the prompt area of ​​the target transmission line in the first transmission line inspection image sample according to the prompt type. It then generates point set data for the first transmission line inspection image sample based on the coordinate point data and the point type of each coordinate point within the prompt area. For each coordinate point in the point set data, it generates a position vector based on the coordinate point data, a type vector based on the coordinate point type, a fusion vector based on the position vector and type vector, and a prompt embedding matrix for the first transmission line inspection image sample based on the fusion vectors of each coordinate point. The prompt types include point prompts and text prompts, and the point types include foreground coordinate point types and background coordinate point types. The mask decoder generates the optimal mask for the first transmission line inspection image sample based on the prompt embedding matrix, the feature map of the first transmission line inspection image sample, and the preset output token. It then generates a predicted segmentation image of the first transmission line inspection image sample based on the optimal mask and the first transmission line inspection image sample. The image segmentation loss under the current iteration training is determined based on the predicted segmentation image of the first transmission line inspection image sample and the actual segmentation image of the first transmission line inspection image sample.

8. The multi-view classification device for transmission line inspection images as described in claim 7, characterized in that, Also includes: Multi-view classification model building module; The multi-view classification model construction module is used to obtain the view category of the predicted segmented image of each first transmission line inspection image sample; Construct an initial multi-view classification model; The initial multi-view classification model is trained by using the predicted segmented image of the first transmission line inspection image sample as the input and the view category of the predicted segmented image of the first transmission line inspection image sample as the output. The multi-view classification model is generated when the initial multi-view classification model converges.

9. A terminal device, characterized in that, The system includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements a multi-view classification method for transmission line inspection images as described in any one of claims 1 to 5.

10. A storage medium, characterized in that, The storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device where the storage medium is located to perform a multi-view classification method for transmission line inspection images as described in any one of claims 1 to 5.