A method and system for detecting screw status of transformer cabinet
By using high-resolution binocular cameras and deep learning networks, the substation cabinet screw status is automatically detected, which solves the disadvantages of manual detection in the prior art and achieves efficient, accurate and stable detection results.
Patent Information
- Application Number
- CN202411201832.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-29
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-08-29
AI Technical Summary
The prior art has disadvantages of manual detection when detecting the screw state of the substation cabinet, including high error rate, low efficiency and inconsistent detection standards, making it difficult to ensure the stability of the detection quality.
A high-resolution binocular camera is used to collect visible light and infrared images of the internal screws of the substation cabinet, combined with a visual encoder, feature decoder and screw state detection head, and detect the screw state through a deep learning network to achieve automated detection.
It improves the accuracy and robustness of detection, reduces the error rate and working intensity of manual detection, and ensures the consistency of detection standards and the stability of quality.
Smart Images

Figure CN119130962B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of transformer cabinet screw detection, and in particular to a transformer cabinet screw status detection method and system. Background Art
[0002] As a key component of the power system, substations are responsible for converting electrical energy between different voltage levels, ensuring stable transmission from power plants to end users. Their stable operation is crucial to maintaining the continuity and reliability of power supply. However, because substations operate continuously, they are prone to problems such as protruding or loosening screws. These can cause circuit interruptions or short circuits, resulting in malfunctions and even power system failures, leading to significant economic losses and safety risks.
[0003] Existing manual inspection methods have many drawbacks. For example, manual inspection is easily affected by worker fatigue and distraction, which not only increases the error rate but also affects the overall quality of inspection. Inspection personnel work in a limited space, with high workload and low efficiency. The lack of automated tools and standardized processes makes it difficult to ensure the consistency of inspection standards and the stability of inspection quality. Summary of the Invention
[0004] The present invention provides a method and system for detecting the screw status of a transformer cabinet to solve the problems existing in the above-mentioned prior art. The technical solution is as follows:
[0005] On the one hand, a method for detecting the screw status of a transformer cabinet is provided, comprising:
[0006] S1. Use a high-resolution binocular camera to capture images of the screws to be inspected inside the substation cabinet, obtaining a pair of visible light image and infrared image pairs;
[0007] S2. Input the visible light image-infrared image pair into a trained screw state detection backbone network to detect and output a screw state detection result. The screw state detection backbone network is composed of a visual encoder, a feature decoder, and a screw state detection head;
[0008] The visual encoder extracts local and global features from the input visible light image and infrared image respectively through a two-stage state space feature extraction module, and then enhances the features through a state space feature enhancement module to obtain feature maps at three different scales: a scale feature map around the screw, a scale feature map of the entire screw, and a scale feature map of the screw center.
[0009] Then, the three feature maps of different scales are input into the feature decoder, which uses a pyramid network based on a visual state space module to further fuse the features to obtain three multi-scale feature maps P1, P2, and P3;
[0010] Finally, the three multi-scale feature maps P1, P2, and P3 are input into the screw state detection head, and the screw state detection result is output.
[0011] Optionally, the visual encoder extracts the visible light feature map R1 and the infrared feature map T1 respectively through two two-stage state space feature extraction modules, and then inputs the two feature maps R1 and T1 into the cascaded four state space feature enhancement modules to obtain the enhanced feature map R of visible light. 1’ and infrared enhanced feature map T 1’ ,The enhanced feature map obtained in this way takes into account the ,advantages of both the clear texture of visible light ,image and the strong contour expression of infrared image;
[0012] Then the two enhanced feature maps R 1’ and T 1’ The two dual-stage state space feature extraction modules of the next stage are respectively input to obtain the visible light feature map R2 and the infrared feature map T2. Then the two feature maps R2 and T2 are input into the cascaded four state space feature enhancement modules to obtain the enhanced feature map R of visible light. 2’ and infrared enhanced feature map T 2’ , R 2’ and T 2’ After adding, a scale feature map around the screw is obtained;
[0013] Then the two enhanced feature maps R 2’ and T 2’ The two dual-stage state space feature extraction modules of the next stage are respectively input to obtain the visible light feature map R3 and the infrared feature map T3. Then the two feature maps R3 and T3 are input into the cascaded four state space feature fusion modules to obtain the enhanced feature map R of visible light. 3’ and infrared enhanced feature map T 3’ , R 3’ and T 3’ After adding, the overall scale characteristic map of the screw is obtained;
[0014] Then the two enhanced feature maps R 3’ and T 3’ The two dual-stage state space feature extraction modules of the next stage are respectively input to obtain the visible light feature map R4 and the infrared feature map T4. Then the two feature maps R4 and T4 are input into the cascaded four state space feature fusion modules to obtain the enhanced feature map R of visible light. 4’and infrared enhanced feature map T 4’ , R 4’ and T 4’ After adding, the screw center scale characteristic map is obtained.
[0015] Optionally, the feature decoder adopts a pyramid network based on a visual state space module. The input feature map first passes through a bottom-up feature fusion path to further extract features and fuse them upward layer by layer to obtain richer semantic information. Then, it passes through a top-down feature enhancement path and fuses with the bottom-up features to enhance the feature details. In this way, feature maps of different scales are fused together to generate a multi-scale feature map.
[0016] The bottom-up feature fusion path refers to first extracting features from the screw center scale feature map through the visual state space module, then upsampling and adding it to the screw overall scale feature map, and then using the visual state space module to fuse the two scale feature maps, then upsampling and adding it to the screw surrounding scale feature map, and then using the visual state space module to fuse the two scale feature maps to obtain the multi-scale feature map P1;
[0017] The top-down feature enhancement path refers to downsampling the multi-scale feature map P1 through depth-separable convolution and adding it to the corresponding feature map in the bottom-up path to obtain a multi-scale feature map P2; the multi-scale feature map P2 uses a visual state space module for feature extraction, and then downsampling it through depth-separable convolution and adding it to the corresponding feature map in the bottom-up path to obtain a multi-scale feature map P3.
[0018] Optionally, the screw state detection head concatenates the three multi-scale feature maps output by the feature decoder through upsampling and downsampling, followed by a 1×1 convolution to adjust the feature map dimension;
[0019] Next, two branches are used to perform localization and classification tasks respectively, which helps the model better learn the characteristics of these two different types of tasks and improve detection performance. The localization branch uses four 3×3 convolutions, and passes through batch normalization and activation function, followed by one 1×1 convolution, and outputs a localization result of size H×W×4; the classification branch uses two 3×3 convolutions, and passes through batch normalization and activation function, followed by one 1×1 convolution, and outputs a classification result of size H×W×C;
[0020] Finally, the detection result (c, x, y, w, h) is output, where c represents the state category and (x, y, w, h) represents the position and size information of the detection box, representing the x and y coordinates of the center of the detection box, and the width w and height h of the detection box respectively.
[0021] Optionally, the dual-stage state-space feature extraction module comprises two branches, combining the CNN and visual state-space modules to simultaneously capture local and global feature information;
[0022] One branch first performs maximum pooling and global pooling on the input feature map along the channel direction, then connects it to two convolutional layers to adjust the size and number of channels. The resulting map is then multiplied channel by channel through an activation function and sent to the main branch to capture rough global semantic information.
[0023] The main branch, after layer normalization, uses two convolutional layers and activation functions to capture local detail information;
[0024] After the two branches are multiplied, they are added to the input feature map after layer normalization and convolution layer, and then the visual state space module is used to further extract long-distance fine global information, and finally output a feature map with both local detail information and global semantic information.
[0025] Optionally, the state space feature enhancement module first extracts two feature maps R respectively extracted by the two dual-stage state space feature extraction modules. i and T i , channel exchange is performed through the shallow state space feature fusion module to obtain the shallow fusion feature map FR i and FT i ;
[0026] Then the FR i and FT i Send it to the deep state space feature fusion module for deep feature fusion to obtain the deep fusion feature map F i ;
[0027] Finally, the two input feature maps R i and T i , and the deep fusion feature map F i Add them separately to get the enhanced feature map R i’ and T i’ .
[0028] Optionally, the shallow state space feature fusion module converts the input feature map R i and T i Divide into 4 equal parts according to the channel dimension and perform channel swapping: i Select the first and third parts from T i Select the second and fourth parts from T, and splice these four parts together in sequence to generate a new fusion feature map. i Select the first and third parts from R iSelect the second and fourth parts and generate new fusion feature maps accordingly; input these two new fusion feature maps into the visual state space module respectively to obtain the shallow feature fusion map FR i and FT i ;
[0029] The deep state space feature fusion module combines the FR i and FT i After layer normalization, they are divided into two branches, FR i A branch of FT i One branch of FR performs feature extraction through the visual state space module. i Another branch and FT i After adding the other branch of , a gating mechanism is formed through a fully connected layer and an activation function, which are multiplied with the feature maps extracted by the visual state space module respectively, encouraging the learning of complementary features while suppressing redundant features;
[0030] The two branches after multiplication are added together, and a skip connection is added after passing through the fully connected layer. Finally, after passing through the fully connected layer and activation function, the deep fusion feature map F is obtained. i .
[0031] Optionally, the visual state space module is designed based on a selective state space model. The selective state space model is a new type of state space model that introduces a selective mechanism to allow the model to selectively propagate or forget information along the sequence length dimension according to the currently processed input, thereby improving the performance of the model on important modalities. The state space model converts a one-dimensional input sequence into , after potential mapping Mapped to the hidden space, N represents the dimension of the hidden layer, to produce the output The state space model realizes the modeling of the mapping relationship between the input sequence and the output sequence. In the process of learning the input-output mapping, the useful features in the input sequence are automatically learned to better represent the input sequence. Ultimately, the model understands the patterns, structures, and regularities in the input data and converts this information into correct output, thus realizing feature extraction, feature fusion, or feature enhancement.
[0032] The visual state space module is first layer-normalized and then mapped to a high-dimensional space through a fully connected layer. It is then connected to a depth-wise separable convolution for preliminary feature extraction. A SiLU activation function is then added to introduce nonlinear factors to increase the learning ability of the model.
[0033] Since the selective state-space model can only accept sequential input, an eight-way scanning state-space module is added to convert the two-dimensional image blocks into one-dimensional sequential vectors. The selective state-space model is then used to model the input sequence to achieve efficient global information exchange.
[0034] Then add a layer of normalization to further improve the generalization ability of the model;
[0035] A gated connection branch is also added, consisting of a fully connected layer and a SiLU activation function for controlling the weights of the skip connections, to enhance the network's expressiveness and learning capabilities.
[0036] Finally, a fully connected layer is used to map the output to the same dimension as the input, and then a skip connection is added to speed up the network convergence.
[0037] Optionally, the eight-directional scanning state space module first uses a convolution operation to crop the feature map into non-overlapping image blocks;
[0038] Then the image block is expanded along the rows, columns and diagonal scanning to form a sequence of 8 image blocks;
[0039] Then the 8 image blocks are fed into the selective state space model to extract semantic information and obtain the feature sequence of the 8 image blocks;
[0040] Finally, the feature sequences of the eight image blocks are added and averaged to obtain a new feature sequence, the new feature sequence is reshaped into a single image, and the output image is restored to the input size.
[0041] On the other hand, a system for detecting the screw status of a transformer cabinet is provided, the system comprising:
[0042] The acquisition module is used to use a high-resolution binocular camera to capture images of the screws to be inspected inside the substation cabinet, obtaining a pair of visible light image and infrared image pairs;
[0043] A detection module is configured to input the visible light image-infrared image pair into a trained screw state detection backbone network, detect and output a screw state detection result, wherein the screw state detection backbone network is composed of a visual encoder, a feature decoder, and a screw state detection head;
[0044] The visual encoder extracts local and global features from the input visible light image and infrared image respectively through a two-stage state space feature extraction module, and then enhances the features through a state space feature enhancement module to obtain feature maps at three different scales: a scale feature map around the screw, a scale feature map of the entire screw, and a scale feature map of the screw center.
[0045] Then, the three feature maps of different scales are input into the feature decoder, which uses a pyramid network based on a visual state space module to further fuse the features to obtain three multi-scale feature maps P1, P2, and P3;
[0046] Finally, the three multi-scale feature maps P1, P2, and P3 are input into the screw state detection head, and the screw state detection result is output.
[0047] On the other hand, an electronic device is provided, which includes a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the above-mentioned transformer cabinet screw status detection method.
[0048] On the other hand, a computer-readable storage medium is provided, wherein the storage medium stores at least one instruction, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned method for detecting the screw status of a substation cabinet.
[0049] The beneficial effects brought about by the technical solution provided by the present invention include at least:
[0050] 1) This paper adopts a method based on a selective state space model, which takes into account the advantages of both CNN and Transformer models, alleviates the disadvantage of high computational complexity of previous models for long sequence input information, and integrates local and global information while achieving linear computational complexity.
[0051] 2) Infrared images have clear contour information and are complementary to visible light image information. The present invention uses infrared images as a second modality in addition to visible light images, effectively breaking away from the limitation of existing target detection technology that only uses one modality, significantly improving the accuracy and robustness of detection, as well as the adaptability of the model under various environmental conditions.
[0052] 3) To address the problem of a single cross-modal feature fusion method in existing target detection methods, the state-space feature fusion module proposed in this paper maps visible light image features and infrared image features to a latent space for alignment, thereby reducing the differences between cross-modal features and achieving deep cross-modal feature fusion.
[0053] 4) The image serialization method based on eight-directional scanning in this invention can integrate information from different directions and realize the global receptive field compared with the existing image serialization method based on single-directional scanning. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0055] Figure 1 This is a flow chart of a method for detecting the status of screws in a transformer cabinet provided by an embodiment of the present invention;
[0056] Figure 2 Schematic diagram of the backbone network structure for screw status detection provided by an embodiment of the present invention;
[0057] Figure 3 This is a general flow chart of the method for detecting the status of screws in a transformer cabinet provided by an embodiment of the present invention;
[0058] Figure 4 Schematic diagram of the structure of a visual encoder provided by an embodiment of the present invention;
[0059] Figure 5 1 is a schematic diagram of the structure of a feature decoder provided by an embodiment of the present invention;
[0060] Figure 6 This is a structural diagram of a screw status detection head provided by an embodiment of the present invention;
[0061] Figure 7 Schematic diagram of the structure of a two-stage state space feature extraction module provided by an embodiment of the present invention;
[0062] Figure 8 Schematic diagram of the state space feature enhancement module structure provided by an embodiment of the present invention;
[0063] Figure 9 Schematic diagram of the structure of the shallow state space feature fusion module provided by an embodiment of the present invention;
[0064] Figure 10 Schematic diagram of the deep state space feature fusion module provided by an embodiment of the present invention;
[0065] Figure 11 Schematic diagram of the structure of the visual state space module provided by an embodiment of the present invention;
[0066] Figure 12 Schematic diagram of the structure of the eight-direction scanning state space module provided by an embodiment of the present invention;
[0067] Figure 13 This is a block diagram of a transformer cabinet screw status detection system provided by an embodiment of the present invention;
[0068] Figure 14It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0069] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0070] The embodiment of the present invention provides a method for detecting the status of screws in a transformer cabinet. The method can be implemented by an electronic device, which can be a terminal or a server. Figure 1 The flowchart of a method for detecting the screw status of a transformer cabinet is shown. The processing flow of the method may include the following steps:
[0071] S1. Use a high-resolution binocular camera to capture images of the screws to be inspected inside the substation cabinet, obtaining a pair of visible light image and infrared image pairs;
[0072] S2. Input the visible light image-infrared image pair into a trained screw state detection backbone network to detect and output a screw state detection result. The screw state detection backbone network is composed of a visual encoder, a feature decoder, and a screw state detection head;
[0073] Among them, such as Figure 2 As shown, the visual encoder performs local and global feature extraction on the input visible light image and infrared image respectively through a two-stage state space feature extraction module, and then performs feature enhancement through a state space feature enhancement module to obtain feature maps of three different scales: a scale feature map around the screw, a scale feature map of the entire screw, and a scale feature map of the screw center;
[0074] Then, the three feature maps of different scales are input into the feature decoder, which uses a pyramid network based on a visual state space module to further fuse the features to obtain three multi-scale feature maps P1, P2, and P3;
[0075] Finally, the three multi-scale feature maps P1, P2, and P3 are input into the screw state detection head, and the screw state detection result (ie, the category and coordinates of the screw state) is output.
[0076] The training of the screw state detection backbone network in the embodiment of the present invention is performed using image data in the training set, such as Figure 3 Shown, including:
[0077] 1) Data collection and annotation.
[0078] In this embodiment of the present invention, a high-resolution binocular camera is used to capture multiple images of the interior of a substation cabinet, obtaining multiple pairs of visible light and infrared images. The screw conditions in the images are then classified into three types according to different situations: tight screws, protruding and loose screws, and loose left and right screws. The images are annotated using the labelimg tool, using the YOLO target detection annotation format (c0, x0, y0, w0, h0). Here, c0 represents the state category, and (x0, y0, w0, h0) represent the position and size information of the detection frame, representing the horizontal coordinate x0 and vertical coordinate y0 of the center of the detection frame, as well as the width w0 and height h0 of the detection frame, respectively. The annotation file is stored in txt format.
[0079] 2) Dataset division.
[0080] In this embodiment of the present invention, infrared-visible light image pairs and labels are divided into training and validation sets in a 4:1 ratio to form a complete data set. The model is trained on the training set and the model performance is evaluated on the validation set. The model with the best detection results on the validation set is selected as the model for final inference and application.
[0081] 3) Data preprocessing.
[0082] In the embodiment of the present invention, all images are enhanced by random horizontal mirror flipping, random scale scaling, random size cropping, contrast enhancement, etc. to expand the training data set.
[0083] Alternatively, as Figure 4 As shown in the figure, the visual encoder extracts the visible light feature map R1 and the infrared feature map T1 respectively through two two-stage state space feature extraction modules, and then inputs the two feature maps R1 and T1 into the cascaded four state space feature enhancement modules to obtain the enhanced feature map R of visible light. 1’ and infrared enhanced feature map T 1’ ,The enhanced feature map obtained in this way takes into account the ,advantages of both the clear texture of visible light ,image and the strong contour expression of infrared image;
[0084] Then the two enhanced feature maps R 1’ and T 1’ The two dual-stage state space feature extraction modules of the next stage are respectively input to obtain the visible light feature map R2 and the infrared feature map T2. Then the two feature maps R2 and T2 are input into the cascaded four state space feature enhancement modules to obtain the enhanced feature map R of visible light. 2’ and infrared enhanced feature map T 2’ , R 2’ and T 2’ After adding, a scale feature map around the screw is obtained;
[0085] Then the two enhanced feature maps R 2’ and T 2’ The two dual-stage state space feature extraction modules of the next stage are respectively input to obtain the visible light feature map R3 and the infrared feature map T3. Then the two feature maps R3 and T3 are input into the cascaded four state space feature fusion modules to obtain the enhanced feature map R of visible light. 3’ and infrared enhanced feature map T 3’ , R 3’ and T 3’ After adding, the overall scale characteristic map of the screw is obtained;
[0086] Then the two enhanced feature maps R 3’ and T 3’ The two dual-stage state space feature extraction modules of the next stage are respectively input to obtain the visible light feature map R4 and the infrared feature map T4. Then the two feature maps R4 and T4 are input into the cascaded four state space feature fusion modules to obtain the enhanced feature map R of visible light. 4’ and infrared enhanced feature map T 4’ , R 4’ and T 4’ After adding, the screw center scale characteristic map is obtained.
[0087] Alternatively, as Figure 5 As shown in the figure, the feature decoder adopts a pyramid network based on the visual state space module. The input feature map first passes through a bottom-up feature fusion path to further extract features and fuse them layer by layer to obtain richer semantic information. Then, it passes through a top-down feature enhancement path and fuses with the bottom-up features to enhance the feature details. In this way, feature maps of different scales are fused together to generate a multi-scale feature map.
[0088] The bottom-up feature fusion path refers to first extracting features from the screw center scale feature map through the visual state space module, then upsampling and adding it to the screw overall scale feature map, and then using the visual state space module to fuse the two scale feature maps, then upsampling and adding it to the screw surrounding scale feature map, and then using the visual state space module to fuse the two scale feature maps to obtain the multi-scale feature map P1;
[0089] The top-down feature enhancement path refers to downsampling the multi-scale feature map P1 through depth-separable convolution and adding it to the corresponding feature map in the bottom-up path to obtain a multi-scale feature map P2; the multi-scale feature map P2 uses a visual state space module for feature extraction, and then downsampling it through depth-separable convolution and adding it to the corresponding feature map in the bottom-up path to obtain a multi-scale feature map P3.
[0090] Alternatively, as Figure 6 As shown, the screw state detection head concatenates the three multi-scale feature maps output by the feature decoder through upsampling and downsampling, followed by a 1×1 convolution to adjust the feature map dimension;
[0091] Next, two branches are used to perform localization and classification tasks respectively, which helps the model better learn the characteristics of these two different types of tasks and improve detection performance. The localization branch uses four 3×3 convolutions, and passes through batch normalization and activation function, followed by one 1×1 convolution, and outputs a localization result of size H×W×4; the classification branch uses two 3×3 convolutions, and passes through batch normalization and activation function, followed by one 1×1 convolution, and outputs a classification result of size H×W×C;
[0092] Finally, the detection result (c, x, y, w, h) is output, where c represents the state category and (x, y, w, h) represents the position and size information of the detection box, representing the x and y coordinates of the center of the detection box, and the width w and height h of the detection box respectively.
[0093] Alternatively, as Figure 7 As shown, the dual-stage state space feature extraction module includes two branches, which combines the CNN and visual state space modules to capture local and global feature information simultaneously;
[0094] One branch first performs maximum pooling and global pooling on the input feature map along the channel direction, then connects it to two convolutional layers to adjust the size and number of channels. The resulting map is then multiplied channel by channel through an activation function and sent to the main branch to capture rough global semantic information.
[0095] The main branch, after layer normalization, uses two convolutional layers and activation functions to capture local detail information;
[0096] After the two branches are multiplied, they are added to the input feature map after layer normalization and convolution layer, and then the visual state space module is used to further extract long-distance fine global information, and finally output a feature map with both local detail information and global semantic information.
[0097] Alternatively, as Figure 8 As shown, the state space feature enhancement module first extracts the two feature maps R i and T i , channel exchange is performed through the shallow state space feature fusion module to obtain the shallow fusion feature map FR i and FT i ;
[0098] Then the FRi and FT i Send it to the deep state space feature fusion module for deep feature fusion to obtain the deep fusion feature map F i ;
[0099] Finally, the two input feature maps R i and T i , and the deep fusion feature map F i Add them separately to get the enhanced feature map R i’ and T i’ .
[0100] Alternatively, as Figure 9 As shown, the shallow state space feature fusion module converts the input feature map R i and T i Divide into 4 equal parts according to the channel dimension and perform channel swapping: i Select the first and third parts from T i Select the second and fourth parts from T, and splice these four parts together in sequence to generate a new fusion feature map. i Select the first and third parts from R i Select the second and fourth parts and generate new fusion feature maps accordingly (for example, input feature map R i and T i The size of the fused feature map is 640×640×32, which is divided into 4 parts according to the channel dimension, that is, the size of each part is 640×640×8. After fusion as described above, the size of the fused feature map is also 640×640×32); these two new fused feature maps are respectively input into the visual state space module to obtain the shallow feature fusion map FR i and FT i ;
[0101] like Figure 10 As shown, the deep state space feature fusion module combines the FR i and FT i After layer normalization, they are divided into two branches, FR i A branch of FT i One branch of FR performs feature extraction through the visual state space module. i Another branch and FT i After adding the other branch of , a gating mechanism is formed through a fully connected layer and an activation function, which are multiplied with the feature maps extracted by the visual state space module respectively, encouraging the learning of complementary features while suppressing redundant features;
[0102] The two branches after multiplication are added together, and a skip connection is added after passing through the fully connected layer. Finally, after passing through the fully connected layer and activation function, the deep fusion feature map F is obtained. i .
[0103] Optionally, the visual state space module is designed based on a selective state space model. The selective state space model is a new type of state space model that introduces a selective mechanism to allow the model to selectively propagate or forget information along the sequence length dimension according to the currently processed input, thereby improving the performance of the model on important modalities. The state space model converts a one-dimensional input sequence into , after potential mapping Mapped to the hidden space, N represents the dimension of the hidden layer, to produce the output The state-space model models the mapping relationship between input sequences and output sequences. In the process of learning the input-output mapping, it automatically learns useful features in the input sequence, better represents the input sequence, and ultimately understands the patterns, structures, and regularities in the input data, and converts this information into correct outputs, thereby achieving feature extraction, feature fusion, or feature enhancement. (Since digital computers can generally only process discrete-time signals, embodiments of the present invention also require discretization of the continuous-time state-space model to obtain a discrete state-space model.)
[0104] This selective mechanism enables the selective state-space model to maintain efficient computational performance when processing long sequences while improving the ability to focus on key information.
[0105] like Figure 11 As shown in the figure, the visual state space module is first normalized and then mapped to a high-dimensional space through a fully connected layer. It is then connected to a depth-wise separable convolution to initially extract features. A SiLU activation function is then added to introduce nonlinear factors to increase the learning ability of the model.
[0106] Since the selective state-space model can only accept sequential input, an eight-way scanning state-space module is added to convert the two-dimensional image blocks into one-dimensional sequential vectors. The selective state-space model is then used to model the input sequence to achieve efficient global information exchange.
[0107] Then add a layer of normalization to further improve the generalization ability of the model;
[0108] A gated connection branch is also added, consisting of a fully connected layer and a SiLU activation function for controlling the weights of the skip connections, to enhance the network's expressiveness and learning capabilities.
[0109] Finally, a fully connected layer is used to map the output to the same dimension as the input, and then a skip connection is added to speed up the network convergence.
[0110] Since the selective state space model can only process serialized information, the input feature map needs to be serialized.
[0111] Alternatively, as Figure 12 As shown, the eight-way scanning state space module first uses a convolution operation to crop the feature map into non-overlapping image blocks (e.g., 16x16 size);
[0112] Then the image block is expanded along the rows, columns and diagonal scanning to form a sequence of 8 image blocks;
[0113] Then the 8 image blocks are fed into the selective state space model to extract semantic information and obtain the feature sequence of the 8 image blocks;
[0114] Finally, the feature sequences of the eight image blocks are added and averaged to obtain a new feature sequence, the new feature sequence is reshaped into a single image, and the output image is restored to the input size.
[0115] like Figure 13 As shown, an embodiment of the present invention further provides a system for detecting the screw status of a transformer cabinet, the system comprising:
[0116] The acquisition module 1310 is used to use a high-resolution binocular camera to acquire an image of the screw to be inspected inside the transformer cabinet to obtain a pair of visible light image and infrared image pairs;
[0117] A detection module 1320 is configured to input the visible light image-infrared image pair into a trained screw state detection backbone network, and output a screw state detection result. The screw state detection backbone network is composed of a visual encoder 13201, a feature decoder 13202, and a screw state detection head 13203.
[0118] The visual encoder 13201 performs local and global feature extraction on the input visible light image and infrared image respectively through a two-stage state space feature extraction module, and then performs feature enhancement through a state space feature enhancement module to obtain feature maps at three different scales: a screw perimeter scale feature map, a screw overall scale feature map, and a screw center scale feature map;
[0119] Then, the three feature maps of different scales are input into the feature decoder 13202. The feature decoder 13202 uses a pyramid network based on a visual state space module to further fuse the features to obtain three multi-scale feature maps P1, P2, and P3.
[0120] Finally, the three multi-scale feature maps P1, P2, and P3 are input into the screw state detection head 13203, and the screw state detection result is output.
[0121] An embodiment of the present invention provides a system for detecting the screw status of a transformer cabinet, and its functional structure corresponds to an embodiment of the present invention provides a method for detecting the screw status of a transformer cabinet, which will not be described in detail here.
[0122] Figure 14 This is a structural diagram of an electronic device 1400 provided in an embodiment of the present invention. The electronic device 1400 may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 1401 and one or more memories 1402, wherein the memory 1402 stores at least one instruction, and the at least one instruction is loaded and executed by the processor 1401 to implement the steps of the above-mentioned substation screw status detection method.
[0123] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory device containing instructions. These instructions are executable by a processor in a terminal to implement the aforementioned method for detecting the status of a transformer screw. For example, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device.
[0124] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.
[0125] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for detecting the screw status of a transformer cabinet, characterized in that: The method comprises: S1, using a high-resolution binocular camera to collect images of screws to be detected inside the transformer cabinet, and obtaining a pair of visible light image-infrared image pairs; S2, inputting the visible light image-infrared image pair into a trained screw state detection backbone network, detecting and outputting a screw state detection result, wherein the screw state detection backbone network is composed of a visual encoder, a feature decoder and a screw state detection head; The visual encoder extracts local and global features of the input visible light image and infrared image through a two-stage state space feature extraction module, and then enhances the features through a state space feature enhancement module to obtain feature maps of three different scales: a screw perimeter scale feature map, a screw overall scale feature map, and a screw center scale feature map; Then, the three feature maps of different scales are input into the feature decoder, and the feature decoder uses a pyramid network based on a visual state space module to further fuse the features to obtain three multi-scale feature maps P1, P2, and P3; The feature decoder adopts a pyramid network based on a visual state space module. The input feature map first passes through a bottom-up feature fusion path to further extract features and fuse them upward layer by layer to obtain richer semantic information. Then, it passes through a top-down feature enhancement path and fuses with the bottom-up features to enhance the feature details. In this way, feature maps of different scales are fused together to generate a multi-scale feature map. Among them, the bottom-up feature fusion path refers to that the screw center scale feature map is first extracted by the visual state space module, then up-sampled and added to the screw overall scale feature map, and then the visual state space module is used to fuse the two scale feature maps, then up-sampled and added to the screw surrounding scale feature map, and then the visual state space module is used to fuse the two scale feature maps to obtain a multi-scale feature map P1; The top-down feature enhancement path refers to downsampling the multi-scale feature map P1 through a depth-separable convolution and adding it to the corresponding feature map in the bottom-up path to obtain a multi-scale feature map P2; the multi-scale feature map P2 is extracted by a visual state space module, and then down-sampled through a depth-separable convolution and added to the corresponding feature map in the bottom-up path to obtain a multi-scale feature map P3; Finally, the three multi-scale feature maps P1, P2, and P3 are input into the screw state detection head, and the screw state detection result is output; The screw state detection head concatenates the three multi-scale feature maps output by the feature decoder through upsampling and downsampling, followed by a 1×1 convolution to adjust the feature map dimension; Next, two branches are used to perform positioning tasks and classification tasks respectively, which helps the model better learn the characteristics of these two different types of tasks and improve the detection performance. The positioning branch uses 4 3×3 convolutions, and passes through batch normalization and activation function, followed by 1 1×1 convolution, and outputs a positioning result of size H×W×4; the classification branch uses 2 3×3 convolutions, and passes through batch normalization and activation function, followed by 1 1×1 convolution, and outputs a classification result of size H×W×C; Finally, the detection result (c, x, y, w, h) is output, where c represents the category of the state, and (x, y, w, h) represents the position and size information of the detection box, which respectively represent the x and y coordinates of the center of the detection box, and the width w and height h of the detection box.
2. The method according to claim 1, characterized in that The visual encoder extracts the visible light feature map R1 and the infrared feature map T1 respectively through two two-stage state space feature extraction modules, and then inputs the two feature maps R1 and T1 into the cascaded four state space feature enhancement modules to obtain the enhanced feature map R1 of the visible light. 1’ and infrared enhanced feature map T 1’ ,The enhanced feature map obtained in this way takes into account the advantages of clear texture of visible light image and strong contour expression of infrared image; Then the two enhanced feature maps R 1’ and T 1’ The two dual-stage state space feature extraction modules of the next stage are respectively input to obtain the visible light feature map R2 and the infrared feature map T2. Then, the two feature maps R2 and T2 are input to the cascaded four state space feature enhancement modules to obtain the enhanced feature map R of visible light. 2’ and infrared enhanced feature map T 2’ , R 2’ and T 2’ After adding, a scale characteristic map around the screw is obtained; Then the two enhanced feature maps R 2’ and T 2’ The two dual-stage state space feature extraction modules of the next stage are respectively input to obtain the visible light feature map R3 and the infrared feature map T3. Then, the two feature maps R3 and T3 are input into the cascaded four-state space feature fusion module to obtain the enhanced feature map R of visible light. 3’ and infrared enhanced feature map T 3’ , R 3’ and T 3’ After adding, the overall scale characteristic diagram of the screw is obtained; Then the two enhanced feature maps R 3’ and T 3’ The two dual-stage state space feature extraction modules of the next stage are respectively input to obtain the visible light feature map R4 and the infrared feature map T4. Then, the two feature maps R4 and T4 are input into the cascaded four state space feature fusion modules to obtain the enhanced feature map R of visible light. 4’ and infrared enhanced feature map T 4’ , R 4’ and T 4’ After adding, the screw center scale characteristic map is obtained.
3. The method according to claim 1, characterized in that The dual-stage state-space feature extraction module includes two branches, combining the CNN and visual state-space modules to capture local and global feature information simultaneously; One branch first performs maximum pooling and global pooling on the input feature map along the channel direction, connects it to two convolutional layers to adjust the size and number of channels, and then multiplies it channel by channel to the main branch after passing the activation function to capture rough global semantic information; The main branch, after layer normalization, uses two convolutional layers and activation functions to capture local detail information; After the two branches are multiplied, they are added to the input feature map after layer normalization and convolution, and then the long-distance fine global information is further extracted through the visual state space module, and finally a feature map with both local detail information and global semantic information is output.
4. The method according to claim 1, characterized in that: The state space feature enhancement module first extracts the two feature maps R respectively extracted by the two dual-stage state space feature extraction modules. i and T i , channel exchange is performed through the shallow state space feature fusion module to obtain the shallow fusion feature map FR i and FT i ; Then the FR i and FT i It is sent to the deep state space feature fusion module for deep feature fusion to obtain the deep fusion feature map F i ; Finally, the two input feature maps R i and T i , and the deep fusion feature map F i Add them separately to get the enhanced feature map R i’ and T i’ .
5. The method according to claim 4, characterized in that The shallow state space feature fusion module transforms the input feature map R i and T i Divide into 4 equal parts according to the channel dimension and perform channel swapping: i Select the first and third parts from T i Select the second and fourth parts from T, and splice these four parts together in sequence to generate a new fusion feature map. i Select the first and third parts from R i Select the second and fourth parts and generate new fusion feature maps accordingly; The two new fusion feature maps are input into the visual state space module respectively to obtain the shallow feature fusion map FR i and FT i ; The deep state space feature fusion module combines the FR i and FT i After layer normalization, they are divided into two branches: FR i A branch of FT i One branch of FR extracts features through the visual state space module. i Another branch and FT i After adding the other branch of , a gating mechanism is formed through a fully connected layer and an activation function, which is multiplied with the feature map extracted by the visual state space module, encouraging complementary feature learning while suppressing redundant features; The two branches after multiplication are added together, and a skip connection is added after passing through the fully connected layer. Finally, after passing through the fully connected layer and the activation function, the deep fusion feature map F is obtained. i .
6. The method according to claim 1, characterized in that The visual state space module is designed based on the selective state space model. The selective state space model is a new type of state space model. By introducing a selective mechanism, the model can selectively propagate or forget information along the sequence length dimension according to the currently processed input, thereby improving the performance of the model in important modes. The state space model converts a one-dimensional input sequence into Through potential mapping Mapped to the hidden space, N represents the dimension of the hidden layer, to generate the output y(t). The state space model realizes the modeling of the mapping relationship between the input sequence and the output sequence. In the process of learning the input-output mapping, the useful features in the input sequence are automatically learned to better represent the input sequence, and finally the patterns, structures and regularities in the input data are understood, and this information is converted into the correct output to realize feature extraction, feature fusion or feature enhancement. The visual state space module is first normalized and then mapped to a high-dimensional space through a fully connected layer, and then connected to a depthwise separable convolution to preliminarily extract features, and then a SiLU activation function is added to introduce nonlinear factors to increase the learning ability of the model; Since the selective state space model can only accept serialized input, an eight-way scanning state space module is added to transform the two-dimensional image blocks into one-dimensional serialized vectors, and the selective state space model is used to model the input sequence to achieve efficient global information interaction; Then add a layer of normalization to further improve the generalization ability of the model; A gated connection branch is also added, which consists of a fully connected layer and a SiLU activation function for controlling the weights of the skip connection to enhance the network's expressiveness and learning capabilities. Finally, a fully connected layer is used to map the output to the same dimension as the input, and then a skip connection is added to speed up the network convergence.
7. The method according to claim 6, characterized in that The eight-direction scanning state space module first uses a convolution operation to crop the feature map into non-overlapping image blocks; Then the image block is expanded along the row, column and diagonal scanning to form a sequence of 8 image blocks; Then the 8 image blocks are sent to the selective state space model to extract semantic information and obtain the feature sequence of the 8 image blocks; Finally, the feature sequences of the eight image blocks are added and averaged to obtain a new feature sequence, the new feature sequence is reshaped into a single image, and the output image is restored to the input size.
8. A transformer cabinet screw status detection system, characterized in that: The system comprises: The acquisition module is used to acquire the image of the screw to be detected inside the transformer cabinet using a high-resolution binocular camera to obtain a pair of visible light image-infrared image pairs; A detection module, used for inputting the visible light image-infrared image pair into a trained screw state detection backbone network, detecting and outputting a screw state detection result, wherein the screw state detection backbone network is composed of a visual encoder, a feature decoder and a screw state detection head; The visual encoder extracts local and global features of the input visible light image and infrared image through a two-stage state space feature extraction module, and then enhances the features through a state space feature enhancement module to obtain feature maps of three different scales: a screw perimeter scale feature map, a screw overall scale feature map, and a screw center scale feature map; Then, the three feature maps of different scales are input into the feature decoder, and the feature decoder uses a pyramid network based on a visual state space module to further fuse the features to obtain three multi-scale feature maps P1, P2, and P3; The feature decoder adopts a pyramid network based on a visual state space module. The input feature map first passes through a bottom-up feature fusion path to further extract features and fuse them upward layer by layer to obtain richer semantic information. Then, it passes through a top-down feature enhancement path and fuses with the bottom-up features to enhance the feature details. In this way, feature maps of different scales are fused together to generate a multi-scale feature map. Among them, the bottom-up feature fusion path refers to that the screw center scale feature map is first extracted by the visual state space module, then up-sampled and added to the screw overall scale feature map, and then the visual state space module is used to fuse the two scale feature maps, then up-sampled and added to the screw surrounding scale feature map, and then the visual state space module is used to fuse the two scale feature maps to obtain a multi-scale feature map P1; The top-down feature enhancement path refers to downsampling the multi-scale feature map P1 through a depth-separable convolution and adding it to the corresponding feature map in the bottom-up path to obtain a multi-scale feature map P2; the multi-scale feature map P2 is extracted by a visual state space module, and then down-sampled through a depth-separable convolution and added to the corresponding feature map in the bottom-up path to obtain a multi-scale feature map P3; Finally, the three multi-scale feature maps P1, P2, and P3 are input into the screw state detection head, and the screw state detection result is output; The screw state detection head concatenates the three multi-scale feature maps output by the feature decoder through upsampling and downsampling, followed by a 1×1 convolution to adjust the feature map dimension; Next, two branches are used to perform positioning tasks and classification tasks respectively, which helps the model better learn the characteristics of these two different types of tasks and improve the detection performance. The positioning branch uses 4 3×3 convolutions, and passes through batch normalization and activation function, followed by 1 1×1 convolution, and outputs a positioning result of size H×W×4; the classification branch uses 2 3×3 convolutions, and passes through batch normalization and activation function, followed by 1 1×1 convolution, and outputs a classification result of size H×W×C; Finally, the detection result (c, x, y, w, h) is output, where c represents the category of the state, and (x, y, w, h) represents the position and size information of the detection box, which respectively represent the x and y coordinates of the center of the detection box, and the width w and height h of the detection box.
Citation Information
Patent Citations
Multi-modal pedestrian detection method based on improved YOLO model
CN111767882A
Infrared and visible light image fusion method based on multi-scale hybrid converter
CN117274760A