A method for identifying the on / off state of a switch based on the self-attention mechanism
By introducing a Transformer structure with a self-attention mechanism into the YOLOv5 model, the SwF-YOLOv5 network is built, and the problems of low switching state monitoring efficiency and insufficient reliability in the substation are solved, intelligent identification is achieved, and the intelligent operation and maintenance level of the substation is improved.
Patent Information
- Application Number
- CN202210176075.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-02-25
AI Technical Summary
The monitoring of switch split-combination status in the prior art in substations relies on manual inspection or hard contact signal, which has problems of low efficiency and insufficient reliability.
The Transformer structure based on the self-attention mechanism is adopted, and it is integrated into the YOLOv5 object detection model to build a SwF-YOLOv5 network to identify the switch split state. This method enhances feature extraction and fusion capabilities through self-attention mechanism and improves detection accuracy.
It realizes intelligent identification of switch split and joint states, improves the intelligent operation and maintenance level and operation safety and reliability of the substation, and reduces the drawbacks of manual inspection.
Smart Images

Figure CN114596487B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of substation operation and maintenance, and particularly to a method for identifying the opening and closing states of switches based on a self-attention mechanism. Background Art
[0002] Monitoring the operating state of primary equipment is of great significance for the safe, stable, and efficient operation of substations. Currently, checking the opening and closing states of switches is an important part of the daily inspection of substations. At present, it is mainly recorded by means of manual inspection or the method of sending hard contact signals on the background monitor. Under the opportunity of the booming development of artificial intelligence technology, carrying out research on the intelligent identification of the opening and closing states of switches and applying the object detection technology based on deep learning to the discrimination of switch states will help to reduce the disadvantages brought by manual inspection, improve the intelligent operation and maintenance level of substations, and enhance the safety and reliability of the overall operation, which has important practical significance.
[0003] Currently, the mainstream deep learning object detection algorithms (such as YOLOv5) are all constructed based on convolutional neural networks. However, the latest research has found that the Transformer structure based on the self-attention mechanism has shown revolutionary performance improvement in the field of computer vision. Convolution can be regarded as a kind of template matching, and the same template is used for filtering at different positions in the image. The attention unit in the Transformer is an adaptive filter, and the template weight is determined by the combinability of two pixels. This adaptive calculation module has stronger modeling ability. Therefore, integrating the self-attention mechanism module of the Transformer structure into an excellent object detection algorithm based on a convolutional network can enhance the feature extraction and performance expression of the algorithm and improve the overall detection accuracy. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for identifying the opening and closing states of switches based on a self-attention mechanism. By integrating the network structure of the self-attention mechanism into the framework of the benchmark YOLOv5 object detection model, the ability to extract key features in the image and the overall modeling ability can be improved, and beneficial detection effects are obtained.
[0005] The technical solution adopted by the present invention is as follows.
[0006] A method for identifying the opening and closing states of switches based on a self-attention mechanism, comprising:
[0007] S1: Based on the image samples of the substation switch opening and closing indicators and the corresponding manual annotation information, construct a training sample library. The manual annotation information includes the coordinate information and classification information of the rectangular area of the key opening and closing identification. The classification information includes whether the rectangular area of the key opening and closing identification belongs to the categories of open, closed or pointer. The coordinate information includes the upper left corner coordinates, width and height information of the rectangular area of the key opening and closing identification;
[0008] S2: Select the YOLOv5 network as the benchmark network. Replace the backbone feature extraction sub-network of this benchmark network with the Swin Transformer self-attention mechanism network. Replace the neck feature fusion sub-network of this benchmark network with the FPT feature pyramid network. Connect the outputs of the last three multi-scale feature maps in the Swin Transformer network to the three input nodes of the FPT network respectively to form the SwF-YOLOv5 object detection network structure;
[0009] S3: Use the training sample library to train the SwF-YOLOv5 object detection network in S2 to obtain a detection model file;
[0010] S4: Input the image to be detected into the detection model file to obtain the coordinate information and classification information of the rectangular area of the key opening and closing identification of the image to be detected; furthermore, adopt a discrimination algorithm based on the overlap degree to discriminate the opening and closing states of the switches identified in the image to be detected.
[0011] Furthermore, in step S1, the manual annotation information of each image sample is annotated based on expert experience.
[0012] Furthermore, use an open-source image annotation tool to manually annotate the rectangular area of the key opening and closing identification in each image sample based on expert experience to obtain an annotation file in json format;
[0013] Use a python script to convert the annotation file in json format into a normalized txt format file supported by the YOLOv5 algorithm, and randomly divide the training sample library into a training set and a test set at a ratio of 4:1.
[0014] Furthermore, use a python script to count the number of rectangular areas of the key opening and closing identification in each category, and adopt the method of re-collecting image samples or transforming the original image samples using data augmentation techniques for categories with a quantity less than the set threshold to expand the training sample library.
[0015] Furthermore, step S3 includes:
[0016] Step S31: Set up a model training environment on the GPU server, and use the training set to train the SwF-YOLOv5 network to obtain an intermediate model file for identifying the key identification rectangular areas of switch opening and closing in images.
[0017] Step S32: Use the test set to verify the intermediate model file, including: if the verification result does not meet the set requirements, optimize the annotation of each sample image in the training sample library, re-establish the training set and the test set, and return to Step S31 until the verification result meets the set requirements, then use the intermediate model file as the final detection model file.
[0018] Further, in Step S4, a discrimination algorithm based on the overlap degree is adopted, and the results of discriminating the switch opening and closing states identified in the image to be detected include three categories: open, closed, and unknown.
[0019] Further, in Step S4, a discrimination algorithm based on the overlap degree is adopted to discriminate the switch opening and closing states identified in the image to be detected, including the following two cases:
[0020] (1) The case where there is a key identification rectangular area of "pointer" category in the image to be detected
[0021] a) If there are no key identification rectangular areas of "open" and "closed" categories in the image to be detected, then discriminate the switch opening and closing state identified in the image to be detected as unknown;
[0022] b) If there is no key identification rectangular area of "open" category in the image to be detected and only a key identification rectangular area of "closed" category exists, then calculate the area overlap rate between the key identification rectangular area of "pointer" category and the key identification rectangular area of "closed" category. If any of the calculated results is greater than zero, then discriminate the switch opening and closing state identified in the image to be detected as closed; if all the calculated results are less than or equal to zero, then discriminate the switch opening and closing state identified in the image to be detected as unknown;
[0023] c) If there is no key identification rectangular area of "closed" category in the image to be detected and only a key identification rectangular area of "open" category exists, then calculate the area overlap rate between the key identification rectangular area of "pointer" category and the key identification rectangular area of "open" category. If any of the calculated results is greater than zero, then discriminate the switch opening and closing state identified in the image to be detected as open; if all the calculated results are less than or equal to zero, then discriminate the switch opening and closing state identified in the image to be detected as unknown;
[0024] d) If both the "open" and "closed" key identification rectangular regions of the category to be detected in the image exist, calculate the area overlap rate between the key identification rectangular region of the category "pointer" and the key identification rectangular region of the category "open", and record the maximum value of the calculated results as radio 分 ; Calculate the area overlap rate between the key identification rectangular region of the category "pointer" and the key identification rectangular region of the category "closed", and record the maximum value of the calculated results as radio 合 :
[0025] d1) If radio 分 = 0 and radio 合 ≠ 0, then determine that the switch opening and closing state indicated by the image to be detected is closed;
[0026] d2) If radio 合 = 0 and radio 分 ≠ 0, then determine that the switch opening and closing state indicated by the image to be detected is open;
[0027] d3) If radio 合 = 0 and radio 分 = 0, then determine that the switch opening and closing state indicated by the image to be detected is unknown;
[0028] d4) If radio 分 ≠ 0 or radio 合 ≠ 0, then perform the following processing on the key identification rectangular region of the category "pointer": First, perform grayscale conversion and erosion processing, then extract the contour of the "pointer" object, calculate the minimum bounding rectangle, and then determine the position of the pointer based on the angle of the minimum bounding rectangle and combine the coordinate information of the key identification rectangular regions of the categories "open" and "closed" to determine the switch opening and closing state indicated by the image to be detected;
[0029] (2) Case where there is no key identification region of the category "pointer" in the image to be detected
[0030] Respectively count the number of key identification rectangular regions of the category "open" and the number of key identification rectangular regions of the category "closed" in the image to be detected:
[0031] a) If the number of key identification rectangular regions of the category "open" exceeds the number of key identification rectangular regions of the category "closed", then determine that the switch opening and closing state indicated by the image to be detected is open;
[0032] b) If the number of key rectangular regions for closing and opening belonging to the category "closed" exceeds the number of key rectangular regions for closing and opening belonging to the category "open", then it is determined that the switch closing and opening state indicated by the image to be detected is closed;
[0033] c) If the number of key rectangular regions for closing and opening belonging to the category "open" is equal to the number of key rectangular regions for closing and opening belonging to the category "closed", then it is determined that the switch closing and opening state indicated by the image to be detected is unknown.
[0034] Furthermore, the calculation formula for the area overlap ratio of two key rectangular regions S01 and S02 for closing and opening is: ratio = (S01 ∩ S02) / (S01 ∪ S02), where S01 ∩ S02 represents the area of the overlapping part of S01 and S02, and S01 ∪ S02 represents the area formed after the overlapping of S01 and S02.
[0035] Through the above solutions, the present invention can implement an improved YOLOv5 network model based on the self-attention mechanism. When using it to identify the closing and opening indication states of substation switches, on the one hand, it can ensure the reliability of the recognition results and meet the needs of intelligent verification. On the other hand, the introduced self-attention mechanism module can enhance the feature expression of the network and improve the detection effect.
[0036] Beneficial Effects
[0037] Compared with the prior art, the advantages of the present invention are: (1) The Swin Transformer network based on the self-attention mechanism is introduced into the backbone feature extraction network of the benchmark YOLOv5 algorithm. This network adopts a hierarchical Transformer self-attention structure and a local attention enhancement structure, and has a more powerful modeling and representation ability compared with the bottleneck convolutional network in the benchmark YOLOv5 algorithm; (2) The FPT pyramid feature network based on the self-attention mechanism is introduced into the neck feature fusion network of the benchmark YOLOv5 algorithm. This network can achieve cross-space and scale feature interaction, and can fuse and generate richer context feature information compared with the feature fusion network of the FPN + PANet structure in the benchmark YOLOv5 algorithm; (3) Due to the adoption of the self-attention mechanism structure, the extraction and fusion of features in the image are optimized, and it has higher detection accuracy.
[0038] At the same time, the detection model file trained by the present invention has a high recognition accuracy, can meet the application requirements for intelligent recognition of the switch closing and opening states in images, eliminate the risk defects brought by manual verification, and improve the intelligent level of substation operation and maintenance. Description of the Drawings
[0039] Figure 1It is a schematic flow diagram of a method for identifying the opening and closing states of switches based on the self-attention mechanism of the present invention;
[0040] Figure 2 It is a schematic structural diagram of the YOLOv5 network;
[0041] Figure 3 It is a schematic structural diagram of the SwF-YOLOv5 network of the present invention. Specific embodiments
[0042] The following is further described in conjunction with the accompanying drawings and specific embodiments.
[0043] This embodiment introduces a method for identifying the opening and closing states of switches based on the self-attention mechanism, as Figure 1 shown, including:
[0044] 1. Image sample acquisition and annotation
[0045] Collect image samples of the opening and closing state indicators beside the switch equipment in the substation to obtain a sample library with a sufficient number of samples and comprehensive coverage of the opening and closing features; the image samples contain key rectangular regions for judging the opening or closing state of the switch, such as the opening and closing regions marked with text, or the opening and closing regions marked with red and green, or the current state region marked with a pointer, etc. Use data augmentation methods such as optical transformation, geometric transformation, adding noise, and data source expansion to randomly process the image sample data to obtain a training sample library.
[0046] Furthermore, use a python script to count the number of key rectangular regions for opening and closing of each category (hereinafter referred to as key regions), and for categories with a small number, re-take pictures or use data augmentation techniques such as pixel content transformation and spatial geometric transformation on the original images to obtain an expanded training sample library. In this embodiment, there are a total of 5085 training samples, including three categories of key regions: opening, closing, and pointer.
[0047] Furthermore, use an open-source image annotation tool to manually annotate the key regions of the image samples in the training sample library. The annotation information includes the coordinate information and classification information of the key regions. After annotation, a json-format annotation file is obtained. The classification information includes whether the key region belongs to the category of opening, closing, or pointer, and the coordinate information includes the upper left corner coordinates, width, and height information of the key region.
[0048] Furthermore, use a python script to convert the json-format annotation file into a normalized txt-format file supported by the YOLOv5 algorithm, and randomly divide the sample library into a training set and a test set in a ratio of 4:1 as the data source for subsequent training of the network model.
[0049] 2. Construction of the SwF-YOLOv5 Network
[0050] The YOLOv5 network is selected as the baseline network. As Figure 2 shown, this baseline network exhibits extremely excellent detection performance and popularization value in actual use. It integrates a large number of cutting-edge computer vision technologies, significantly improves the performance of object detection, and enhances the speed of model training and the convenience of model application. This baseline network mainly consists of a backbone feature extraction network (Backbone network), a neck feature fusion network (Neck network), and a detection head prediction network (Prediction network). It uses CSPDarknet53 based on the bottleneck structure as the Backbone, the FPN+PANet structure based on multi-feature map fusion as the Neck, and the YOLO detection head Head for regression and classification tasks based on the position and category of the detection target.
[0051] The Transformer network is a classic network proposed by Google in 2017 based on the self-attention mechanism structure. It has completely revolutionized the field of natural language processing (NLP), has absolute technical advantages, and has become the standard network in this field. The Transformer network has many advantages that convolutional neural networks and recurrent neural networks do not have, such as general and powerful modeling capabilities, large throughput and large-scale parallel processing capabilities, etc., and has been widely used in the field of NLP.
[0052] The Swin Transformer network is a Transformer network proposed in 2021 that adopts a local self-attention enhancement mechanism. It uses the Transformer self-attention mechanism in the field of computer vision to extract the characteristic representation of images. This network has stronger dynamic computing capabilities and stronger modeling capabilities compared to convolutional neural networks, and can adaptively calculate the local and global pixel relationships, making it very valuable for popularization; in addition, the hierarchical structure in its network can obtain feature map representations of different scales, which is very suitable for replacing the CSPDarknet53 structure in the backbone network of the baseline YOLOv5. Therefore, the present invention combines this network with the YOLOv5 network structure to obtain better feature extraction capabilities.
[0053] The FPT network is a multi-directional fusion feature pyramid network proposed in 2020. Its core is also the Transformer network using the self-attention mechanism, which can deeply capture the non-local context information of objects at different scales. By using three specially designed Transformer structures, it transforms any feature pyramid into another feature pyramid of the same size but with richer context in a top-down and bottom-up interaction manner. Since the output dimension of FPT is the same as the input, it can be freely embedded into various detection algorithms containing feature pyramids. Therefore, in the present invention, this network is combined with the YOLOv5 network structure to achieve better feature fusion ability.
[0054] The approach of the present invention is to replace the backbone network based on the bottleneck convolutional neural network in the YOLOv5 baseline network with the Swin Transformer network for image feature extraction, and replace the neck feature fusion network based on FPN and PANet with the FPT feature pyramid fusion network. The output is drawn from the last three multi-scale feature map nodes of the Swin Transformer network and connected to the three input feature map nodes of the FPT feature pyramid fusion network to obtain the SwF-YOLOv5 network model structure as Figure 3 shown.
[0055] 3. Training of the detection model
[0056] In this embodiment, a containerized model training environment is built on a GPU server, and an intermediate model file is obtained by training the improved YOLOv5 network for 300 rounds. Then, the intermediate model file is evaluated using the test set samples. The evaluation metrics can comprehensively consider mainstream metrics in the evaluation of deep learning models such as mAP (multi-class average precision), Precision (accuracy), Recall (recall rate), and Flops (computing power required by the model). Determine whether the evaluation metrics meet the requirements of technical specifications or popularization and application. If not, rebuild the training environment by optimizing sample annotation, image data augmentation processing, adjusting the parameters of the Swin Transformer network, etc. Then, through repeated iterative training and evaluation, the final detection model file and the corresponding training sample library are obtained.
[0057] 4. Model inference and recognition of the image splitting and combining state
[0058] Taking the image to be detected as the input, through the inference operation of the final detection model, the classification information and coordinate information of the key areas in the image to be detected are obtained.
[0059] According to the classification information and coordinate information of the key regions in the image to be detected, a discrimination algorithm based on the overlap degree is adopted to discriminate the opening and closing states of the switch identified by the current image. The recognition results include three categories: open, closed, and unknown.
[0060] The specific implementation process of this algorithm is as follows:
[0061] (1) In the case where there is a rectangular key identification region for opening and closing whose category is "pointer" in the image to be detected
[0062] a) If there are no rectangular key identification regions for opening and closing whose categories are "open" and "closed" in the image to be detected, then the opening and closing state of the switch identified by the image to be detected is discriminated as unknown;
[0063] b) If there is no rectangular key identification region for opening and closing whose category is "open" in the image to be detected and there is only a rectangular key identification region for opening and closing whose category is "closed", then calculate the area overlap rate between the rectangular key identification region for opening and closing whose category is "pointer" and the rectangular key identification region for opening and closing whose category is "closed". If any of the calculated results is greater than zero, then the opening and closing state of the switch identified by the image to be detected is discriminated as closed. If all the calculated results are less than or equal to zero, then the opening and closing state of the switch identified by the image to be detected is discriminated as unknown;
[0064] c) If there is no rectangular key identification region for opening and closing whose category is "closed" in the image to be detected and there is only a rectangular key identification region for opening and closing whose category is "open", then calculate the area overlap rate between the rectangular key identification region for opening and closing whose category is "pointer" and the rectangular key identification region for opening and closing whose category is "open". If any of the calculated results is greater than zero, then the opening and closing state of the switch identified by the image to be detected is discriminated as open. If all the calculated results are less than or equal to zero, then the opening and closing state of the switch identified by the image to be detected is discriminated as unknown;
[0065] d) If there are rectangular key identification regions for opening and closing whose categories are "open" and "closed" in the image to be detected, then calculate the area overlap rate between the rectangular key identification region for opening and closing whose category is "pointer" and the rectangular key identification region for opening and closing whose category is "open", and take the maximum value of the calculated results and denote it as radio 分 ; Calculate the area overlap rate between the rectangular key identification region for opening and closing whose category is "pointer" and the rectangular key identification region for opening and closing whose category is "closed", and take the maximum value of the calculated results and denote it as radio 合 :
[0066] d1) If radio 分 = 0 and radio 合If it is not equal to 0, then it is determined that the on / off state of the switch identified by the image to be detected is closed;
[0067] d2) If radio 合 = 0 and radio 分 ≠ 0, then it is determined that the on / off state of the switch identified by the image to be detected is open;
[0068] d3) If radio 合 = 0 and radio 分 = 0, then it is determined that the on / off state of the switch identified by the image to be detected is unknown;
[0069] d4) If radio 分 ≠ 0 and radio 合 ≠ 0, then the following processing is performed on the rectangular area of the on / off key identifier whose category is "pointer": First, perform grayscale processing, erosion, etc., then extract the contour of the "pointer" object, calculate the minimum bounding rectangle, and then judge the position of the pointer according to the angle of the minimum bounding rectangle, and combine the coordinate information of the rectangular area of the on / off key identifier whose category is "open" or "closed" to judge the on / off state of the switch identified by the image to be detected;
[0070] (2) The case where there is no rectangular area of the on / off key identifier whose category is "pointer" in the image to be detected
[0071] Respectively count the number of rectangular areas of the on / off key identifier whose category is "open" and the number of rectangular areas of the on / off key identifier whose category is "closed" in the image to be detected:
[0072] a) If the number of rectangular areas of the on / off key identifier whose category is "open" exceeds the number of rectangular areas of the on / off key identifier whose category is "closed", then it is determined that the on / off state of the switch identified by the image to be detected is open;
[0073] b) If the number of rectangular areas of the on / off key identifier whose category is "closed" exceeds the number of rectangular areas of the on / off key identifier whose category is "open", then it is determined that the on / off state of the switch identified by the image to be detected is closed;
[0074] If the number of rectangular areas of the on / off key identifier whose category is "open" is equal to the number of rectangular areas of the on / off key identifier whose category is "closed", then it is determined that the on / off state of the switch identified by the image to be detected is unknown.
[0075] Furthermore, the calculation method of the area overlap rate radio of the two key regions S01 and S02 is as follows:
[0076] If there is an intersection between S01 and S02 itself, or an intersection exists when the long side of the S01 area is extended to have a maximum overlapping area with the S02 area, it is considered that there is an overlapping area between these two key areas. The formula for calculating the area overlapping ratio is: ratio = (S01 ∩ S02) / (S01 ∪ S02), that is, the ratio of the intersection area to the union area of the two key areas.
[0077] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0078] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of processes and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0079] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0080] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0081] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims. All of these fall within the protection scope of the present invention.
Claims
1. A method for identifying the opening and closing states of switches based on the self-attention mechanism, characterized in that, it includes: S1: Based on the image samples of the substation switch opening and closing indicators and the corresponding manual annotation information, construct a training sample library. The manual annotation information includes the coordinate information and classification information of the rectangular area of the key opening and closing identification. The classification information includes whether the rectangular area of the key opening and closing identification belongs to the categories of open, closed or pointer. The coordinate information includes the upper left corner coordinates, width and height information of the rectangular area of the key opening and closing identification; S2: Select the YOLOv5 network as the benchmark network, replace the backbone feature extraction sub-network of the benchmark network with the SwinTransformer self-attention mechanism network, replace the neck feature fusion sub-network of the benchmark network with the FPT feature pyramid network, and connect the outputs of the last three multi-scale feature maps in the Swin Transformer network to the three input nodes of the FPT network respectively to form the SwF-YOLOv5 object detection network structure; S3: Use the training sample library to train the SwF-YOLOv5 object detection network in S2 to obtain a detection model file; S4: Input the image to be detected into the detection model file to obtain the coordinate information and classification information of the rectangular area of the key opening and closing identification of the image to be detected; and then adopt a discrimination algorithm based on the overlap degree to discriminate the opening and closing states of the switches identified in the image to be detected.
2. A method for identifying the opening and closing states of switches based on the self-attention mechanism according to claim 1, characterized in that, in step S1, the manual annotation information of each image sample is annotated based on expert experience.
3. A method for identifying the opening and closing states of switches based on the self-attention mechanism according to claim 2, characterized in that, Use an open-source image annotation tool to manually annotate the rectangular area of the key opening and closing identification in each image sample based on expert experience to obtain an annotation file in json format; Use a python script to convert the annotation file in json format into a normalized txt format file supported by the YOLOv5 algorithm, and randomly divide the training sample library into a training set and a test set in a ratio of 4:
1.
4. A method for identifying the opening and closing states of switches based on the self-attention mechanism according to claim 3, characterized in that, Use a python script to count the number of rectangular areas of the key opening and closing identifications of each category, and adopt the method of re-collecting image samples or transforming the original image samples using data augmentation technology to expand the training sample library for the categories with the number less than the set threshold.
5. A method for identifying the opening and closing states of switches based on the self-attention mechanism according to claim 3, characterized in that, step S3 in it includes: Step S31: Build a model training environment on the GPU server, and use the training set to train the SwF-YOLOv5 network to obtain an intermediate model file for identifying the rectangular area of the key opening and closing identification in the image; Step S32: Verify the intermediate model file using the test set, including: if the verification result does not meet the set requirements, optimize the annotation of each sample image in the training sample library, re - establish the training set and the test set, and return to Step S31 until the verification result meets the set requirements, then use the intermediate model file as the final detection model file.
6. A method for identifying the on - off state of a switch based on the self - attention mechanism according to claim 1, characterized in that, in step S4, a discrimination algorithm based on the degree of overlap is adopted, and the results of discriminating the on - off state of the switch identified by the image to be detected include: three categories: open, closed, and unknown.
7. A method for identifying the on - off state of a switch based on the self - attention mechanism according to claim 6, characterized in that, in step S4, a discrimination algorithm based on the degree of overlap is adopted to discriminate the on - off state of the switch identified by the image to be detected, including the following two cases: (1) The case where there is a key on - off identification rectangular area belonging to the category of "pointer" in the image to be detected a) If there is no key on - off identification rectangular area belonging to the categories of "open" and "closed" in the image to be detected, then the on - off state of the switch identified by the image to be detected is determined to be unknown; b) If there is no key on - off identification rectangular area belonging to the category of "open" in the image to be detected and only a key on - off identification rectangular area belonging to the category of "closed" exists, then calculate the area overlap rate between the key on - off identification rectangular area belonging to the category of "pointer" and the key on - off identification rectangular area belonging to the category of "closed". If any of the calculated results is greater than zero, then the on - off state of the switch identified by the image to be detected is determined to be closed; if all of the calculated results are less than or equal to zero, then the on - off state of the switch identified by the image to be detected is determined to be unknown; c) If there is no key on - off identification rectangular area belonging to the category of "closed" in the image to be detected and only a key on - off identification rectangular area belonging to the category of "open" exists, then calculate the area overlap rate between the key on - off identification rectangular area belonging to the category of "pointer" and the key on - off identification rectangular area belonging to the category of "open". If any of the calculated results is greater than zero, then the on - off state of the switch identified by the image to be detected is determined to be open; if all of the calculated results are less than or equal to zero, then the on - off state of the switch identified by the image to be detected is determined to be unknown; d) If both the separation and closing key identification rectangular regions of the category to be detected in the image to be detected exist, calculate the area overlap rate between the separation and closing key identification rectangular regions of the category "pointer" and the separation key identification rectangular region of the category "separation" respectively, and record the maximum value of the calculated results as radio 分 ; Calculate the area overlap rate between the separation and closing key identification rectangular regions of the category "pointer" and the closing key identification rectangular region of the category "closing" respectively, and record the maximum value of the calculated results as radio 合 : d1) If radio 分 = 0 and radio 合 ≠ 0, then determine that the switch opening / closing state identified by the image to be detected is closed; d2) If radio 合 = 0 and radio 分 ≠ 0, then determine that the switching on / off state indicated by the image to be detected is off; d3) If radio 合 = 0 and radio 分 = 0, then it is determined that the switch opening / closing state identified by the image to be detected is unknown; d4) If radio 分 ≠ 0, turn on radio 合 ≠ 0, then perform the following processing on the key identification rectangular area for separating and combining whose category is "pointer": first perform grayscale processing and erosion processing, then extract the contour of the "pointer" object, calculate the minimum bounding rectangle, and then judge the position of the pointer according to the angle of the minimum bounding rectangle, and combine the coordinate information of the key identification rectangular area for separating and combining whose category is "separated" or "closed" to judge the switch separating and combining state identified by the image to be detected; (2) The case where there is no key on - off identification area belonging to the category of "pointer" in the image to be detected Count the number of key on - off identification rectangular areas belonging to the category of "open" and the number of key on - off identification rectangular areas belonging to the category of "closed" in the image to be detected respectively: a) If the number of key on - off identification rectangular areas belonging to the category of "open" exceeds the number of key on - off identification rectangular areas belonging to the category of "closed", then the on - off state of the switch identified by the image to be detected is determined to be open; b) If the number of key on - off identification rectangular areas belonging to the category of "closed" exceeds the number of key on - off identification rectangular areas belonging to the category of "open", then the on - off state of the switch identified by the image to be detected is determined to be closed; c) If the number of split - merge key identification rectangular regions whose category is "separated" is equal to the number of split - merge key identification rectangular regions whose category is "closed", then it is determined that the switch split - merge state identified by the image to be detected is unknown.
8. A method for identifying the switch split - merge state based on the self - attention mechanism according to claim 6, characterized in that, The calculation formula for the area overlap ratio of two split - merge key identification rectangular regions S01 and S02 is: ratio=(S01∩S02) / (S01∪S02), where S01∩S02 represents the area of the overlapping part of S01 and S02, and S01∪S02 represents the area formed after the overlap of S01 and S02.