Electric power production place indication type equipment state detection and identification method and system

Through deep learning object detection technology and data enhancement methods, an exclusive image data set for the power industry is built, which solves the problems of insufficient equipment status detection accuracy and high resource consumption in power production sites, and achieves efficient and accurate equipment status recognition.

CN120088543APending Publication Date: 2025-06-03SHENZHEN TIEYUE ELECTRIC CO LTD

Patent Information

Application Number
CN202510136970.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art has insufficient equipment status detection and recognition accuracy in power production sites, poor ability to adapt to complex scenarios, and high resource consumption for training and reasoning.

Method used

Deep learning object detection technology is adopted to achieve efficient detection and recognition of equipment status by building exclusive image data sets in the power industry and combining mainstream object detection models. Specific steps include image data acquisition and labeling, data augmentation, object detection model training and inference, and post-processing to generate the final device state value.

Benefits of technology

High-precision extraction of equipment status is achieved, and the problem of insufficient model feature extraction effect caused by large differences between the open source model pre-trained data set and the power industry data is solved, which significantly reduces resource consumption and time costs, and improves the accuracy and recall of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088543A_ABST
    Figure CN120088543A_ABST
Patent Text Reader

Abstract

The invention relates to the field of intelligent monitoring, and discloses an electric power production place indication type equipment state detection and identification method and system, and the method comprises the following steps: (1) collecting equipment pictures for marking, and carrying out the classification and state division; (2) performing data enhancement, and adjusting category labels of directional equipment; (3) training a target detection model by using enhanced data, and outputting a detection frame and a category label; (4) inputting a to-be-detected picture, predicting an equipment area and a category, and removing repeated targets; (5) post-processing a detection result, combining similar equipment, and only outputting a state value; and (6) generating a final state value of the directional equipment according to the mapping rule and outputting the final state value. Through deep learning of target detection and a power industry exclusive image data set, in combination with a mainstream target detection model, efficient and accurate detection and identification of a power scene complex equipment state are realized, the model feature extraction capability is improved, and the resource consumption and time cost of training and reasoning are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent monitoring, and specifically to a method and system for detecting and identifying the state of an indicating device in a power production site. Background Art

[0002] With the continuous improvement of the automation and intelligence levels in power production sites, various indicating devices (such as indicator lights, pressure plates, air switches, rotary switches, and handle switches, etc.) are widely used in the state monitoring of power equipment. These devices play an important role in monitoring the operating state of power equipment and ensuring the safety of the power system. In the prior art, the methods for monitoring and identifying the device state mainly include manual inspection, semi-automated methods based on traditional image processing techniques, as well as multi-modal techniques and early object detection methods that have emerged in recent years. Manual inspection relies on on-site operators to observe and record the state of the device, usually with low efficiency and is easily affected by human factors. Traditional image processing techniques identify the device state by extracting simple features such as edges and colors, which have a certain degree of automation ability, but have limited adaptability to complex environments and the diversity of device states. Multi-modal techniques achieve comprehensive identification of the device state by combining visual models and language models, which have strong theoretical advantages, but their application depends on pre-trained models and fine-tuning of specific image-text pair datasets, and has high resource requirements. Early object detection methods usually use simple network structures for detection and classification, but their performance is insufficient compared to current mainstream deep learning methods, and the process is relatively complex.

[0003] However, the prior art still has some problems that need to be urgently solved when facing the complex scenarios in power production sites. On the one hand, although manual inspection and traditional image processing techniques can meet the basic state detection requirements to a certain extent, for the recognition problems caused by characteristics such as variable lighting conditions, complex background interference, and diverse device state categories in the power equipment scenario, the accuracy and robustness of these methods are insufficient, and it is difficult to achieve intelligent real-time monitoring. On the other hand, although multi-modal techniques in recent years can theoretically improve the ability to identify the device state, since the power industry belongs to a long-tail field, its characteristics are quite different from the pre-trained datasets of open-source models, resulting in poor model adaptability, high training and inference resource consumption, and it is difficult to operate efficiently in resource-constrained scenarios. In addition, although early object detection techniques can combine deep learning for detection to a certain extent, their network performance is relatively backward, and the detection and classification tasks need to be completed step by step, with a complex overall process and low utilization efficiency of computing resources, making it difficult to meet the requirements of high precision, real-time performance, and efficient resource utilization in power production sites. Summary of the Invention

[0004] In view of the deficiencies of the prior art, the present invention provides a method and system for detecting and identifying the status of an indicator-type device in a power production site, which solves the problems of insufficient accuracy in detecting and identifying the status of devices in the prior art in a power production site, poor adaptability to complex scenarios, and high consumption of training and inference resources.

[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for detecting and identifying the status of an indicator-type device in a power production site, comprising the following steps: (1) Collect picture data of indicator lights, pressure plates, air switches, knob switches, and handle switches. The pictures need to cover conditions of different types, styles, scenarios, and angles. Manually label the collected picture data, and divide the data categories according to the device category and status. Among them, the indicator lights are combined and labeled according to the color and on / off status, the pressure plates are combined and labeled according to the style and open / closed status, the air switches are labeled according to the tripping direction, and the knob switches and handle switches are labeled according to the pointing direction; (2) Perform data augmentation operations on the labeled picture data, including mirroring, flipping, brightness and saturation transformation, multi-scale scaling, mosaic splicing, and grayscale processing. Dynamically adjust the category labels of the enhanced images of directional devices. Among them, the mirror operation adjusts the category label from left to right, and the up and down flipping operation adjusts the category label from top to bottom; (3) Use the enhanced picture data to train a target detection model. The target detection model is constructed based on a deep learning algorithm. By inputting the picture data, it outputs the device detection frame and category label. A custom loss function is used to optimize the parameters during model training; (4) Input the picture to be detected into the trained target detection model, predict the possible device areas and category information in the picture, and remove duplicate detection targets through the non-maximum suppression algorithm, only retaining the detection result with the highest confidence in each category; (5) Post-process the detection results, merge the detection results of the same type of device, and only output the device status value. Among them, the indicator light only outputs the on or off status, the pressure plate only outputs the open, closed, or vacant status, and the air switch, knob switch, and handle switch only output the direction status; (6) According to the detection results of the directional devices and the preset direction-status mapping rules, perform status mapping on the detection results to generate and output the final status value of the directional devices.

[0006] Preferably, the custom loss function used during the training process of the target detection model includes foreground and background classification loss, category classification loss, and bounding box regression loss, where: (1) The foreground and background classification loss is calculated through the cross-entropy of the probability value of predicting whether the target is a device area and the true label value, and is used to measure the ability of the model to distinguish between device and non-device areas; (2) The class classification loss is calculated by the cross-entropy between the predicted probability values of the device classes and the true class labels, and is used to measure the accuracy of the model in identifying device classes; (3) The bounding box regression loss is obtained by calculating the similarity between the predicted bounding box and the true bounding box. The similarity metric can use the intersection over union (IOU) or improved IOU formulas such as CIOU, DIOU, and SIOU.

[0007] Preferably, the optimization strategy adopted during the training process of the object detection model is based on the stochastic gradient descent method. During the optimization process, the model parameters are updated through the following formula: w t+1 =w t -η·m t where, w t represents the model parameters at the t-th iteration, η represents the learning rate, and m t represents the exponentially weighted average of the current gradient.

[0008] Preferably, the non-maximum suppression algorithm is used to remove objects with low confidence or high bounding box overlap, and only the detection box with the highest confidence in each class of devices is retained. The condition for duplicate removal is that the overlap ratio between the bounding boxes is greater than the set threshold, and the threshold can be 0.5 or set according to actual needs.

[0009] Preferably, when the data augmentation operations include mirroring and flipping, the class labels of the directional devices will be dynamically adjusted according to the augmentation operations, where: (1) The mirror operation adjusts the left direction to the right direction and the right direction to the left direction; (2) The up-down flipping operation adjusts the up direction to the down direction and the down direction to the up direction; (3) The mirror and flipping operations can be used alone or in combination to generate more transformation effects. For example: If the device direction in the original image is upper right, it can be obtained as lower right through up-down flipping; It can be obtained as upper left through the mirror operation; By simultaneously using the mirror and up-down flipping operations, it can be obtained as lower left.

[0010] Preferably, the detection results of the directional devices are converted into status values according to the preset direction-status mapping rules. The specific rules include: (1) The up direction of the air switch corresponds to the closing state, and the down direction corresponds to the tripping state; (2) The direction of the knob switch can be configured in advance according to actual needs. For example, the left direction corresponds to the standby state, the right direction corresponds to the running state, or different gears and start-stop states corresponding to the device functions; (3) The direction of the handle switch can be pre-configured according to actual needs. For example, the upward direction corresponds to the input state, and the downward direction corresponds to the exit state. The specific mapping relationship is set by the user according to the actual application scenario.

[0011] Preferably, in the post-processing of the detection results, the output content includes the status value of the device, detection box information, and category label information, where: (1) The indicator light only outputs the on or off state; (2) The pressure plate only outputs the open / closed or vacant state; (3) The air switch, knob switch, and handle switch only output the status value corresponding to the direction.

[0012] Preferably, the target detection model is a YOLOv5 or YOLOv8 network structure. The network model includes a feature extraction layer, a convolutional layer, and a classification layer, and completes device target detection and status recognition through an end-to-end training process.

[0013] Preferably, in the training data of the target detection model, the sample quantity of directional devices is extended to more than three times the original sample quantity through data augmentation to improve the detection recall rate of the model for low-frequency direction states.

[0014] A system for detecting and identifying the status of indicator devices in a power production site includes: An image acquisition module for acquiring device pictures of indicator lights, pressure plates, air switches, knob switches, and handle switches; a data annotation module for manually annotating the acquired pictures and classifying them; A data augmentation module for performing augmentation operations such as mirroring, flipping, and multi-scale scaling on the annotated pictures and adjusting the directional category labels; A target detection model training module for training a target detection model based on deep learning; A target detection module for detecting the device category and status information in the pictures; A status post-processing module for merging the detection results of the same type of devices and outputting the final status value according to the direction-status mapping rule.

[0015] The present invention provides a method and system for detecting and identifying the status of indicator devices in a power production site. It has the following beneficial effects: 1. Through deep learning object detection technology, by constructing an image dataset exclusive to the power industry and combining mainstream object detection models, the present invention realizes the efficient detection and recognition of equipment status, achieving the technical effect of accurately extracting the complex equipment status in power scenarios. Compared with the multi-modal technology that relies on open-source vision models and language models in the prior art, the present invention solves the problem that the feature extraction effect of the model is insufficient due to the large difference between the pre-trained dataset of the open-source model and the power industry data, and at the same time reduces the high requirements for the construction and annotation capabilities of specific image-text datasets, significantly reducing the resource consumption and time cost of training and inference.

[0016] 2. The present invention adopts an end-to-end object detection model structure, directly completing the detection and status recognition of equipment based on deep learning algorithms, avoiding the complexity of multi-step operations in the prior art that detect through traditional image processing and simple convolutional networks and then use SVM classification. It achieves the effects of higher overall model network consistency and stronger image feature extraction ability. Compared with the problems of separate detection and classification and low inference efficiency in the prior art, the present invention solves the problem that detection and status recognition cannot be completed integrally, and significantly improves the accuracy and recall rate of the model.

[0017] 3. By adopting the mainstream YOLO series deep learning algorithms, the present invention designs an end-to-end object detection model, which can complete the recognition of equipment status with only one training and inference, achieving the technical effects of higher resource utilization rate and shorter inference time. Compared with the method in the prior art that uses early object detection models to locate equipment and relies on classification models to complete status recognition, the present invention solves the problems of large resource consumption and high training complexity caused by the backward performance of the algorithm network, significantly optimizing the overall performance.

[0018] 4. By introducing a variety of data augmentation methods, including strategies such as random mirroring, flipping, and dynamic adjustment of directional class labels, the present invention specifically optimizes the detection performance of directional equipment, achieving the technical effect of a significant increase in the detection accuracy and recall rate of switch-like samples by the model. Compared with the problem in the prior art that the model performance decreases due to the scarcity of some rare status sample data, the present invention effectively solves the problem of insufficient detection of long-tail sample categories and enhances the generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is the flowchart of the method steps of the present invention; Figure 2 is the system module diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0021] Embodiment 1: Please refer to the attached Figure 1 , the embodiment of the present invention provides a method for detecting and identifying the state of an indicator-type device in a power production site, including the following steps: (1) Collect picture data of indicator lights, pressure plates, air switches, knob switches, and handle switches. The pictures need to cover different types, styles, scenarios, and angles. Manually label the collected picture data, and divide the data categories according to the device category and status. Among them, the indicator lights are combined and labeled according to the color and on / off status, the pressure plates are combined and labeled according to the style and open / closed status, the air switches are labeled according to the tripping direction, and the knob switches and handle switches are labeled according to the pointing direction; (2) Perform data augmentation operations on the labeled picture data, including mirroring, flipping, brightness and saturation transformation, multi-scale scaling, mosaic splicing, and grayscale processing. Dynamically adjust the category labels of the enhanced images of directional devices. Among them, the mirror operation adjusts the category label from left to right, and the up / down flipping operation adjusts the category label from up to down; (3) Use the enhanced picture data to train a target detection model. The target detection model is constructed based on a deep learning algorithm. By inputting the picture data, it outputs the device detection box and category label. A custom loss function is used to optimize the parameters during model training; (4) Input the picture to be detected into the trained target detection model, predict the possible device areas and category information in the picture, and remove duplicate detection targets through the non-maximum suppression algorithm, only retaining the detection result with the highest confidence in each category; (5) Post-process the detection results, merge the detection results of the same type of device, and only output the device status value. Among them, the indicator light only outputs the on or off status, the pressure plate only outputs the open, closed, or vacant status, and the air switch, knob switch, and handle switch only output the direction status; (6) According to the detection results of the directional device and the preset direction-status mapping rule, perform status mapping on the detection results, generate the final status value of the directional device and output it; The custom loss function used during the training process of the target detection model includes foreground and background classification loss, category classification loss, and bounding box regression loss, where: (1) The foreground and background classification loss is calculated by the cross-entropy between the predicted probability value of whether the target is the device area and the true label value, and is used to measure the ability of the model to distinguish between the device and non-device areas; (2) The class classification loss is calculated by the cross-entropy between the predicted probability value of the device class and the true class label, and is used to measure the accuracy of the model in identifying the device class; (3) The bounding box regression loss is calculated by computing the similarity between the predicted bounding box and the true bounding box. The similarity metric can use the intersection over union (IOU) or improved IOU formulas such as CIOU, DIOU, and SIOU; The optimization strategy adopted during the training process of the object detection model is based on the stochastic gradient descent method. During the optimization process, the model parameters are updated through the following formula: w t+1 =w t -η·m t where, w t represents the model parameters at the t-th iteration, η represents the learning rate, and m t represents the exponentially weighted average of the current gradient; The non-maximum suppression algorithm is used to remove targets with low confidence or high bounding box overlap, and only the detection box with the highest confidence in each device class is retained. The condition for duplicate removal is that the overlap ratio between the bounding boxes is greater than the set threshold, and the threshold can be 0.5 or set according to actual requirements; When data augmentation operations include mirroring and flipping, the class labels of directional devices will be dynamically adjusted according to the augmentation operations, where: (1) The mirror operation adjusts the left direction to the right direction and the right direction to the left direction; (2) The up-down flipping operation adjusts the up direction to the down direction and the down direction to the up direction; (3) The mirror and flipping operations can be used alone or in combination to generate more transformation effects. For example: If the device direction in the original image is upper right, the lower right can be obtained through up-down flipping; The upper left can be obtained through the mirror operation; By simultaneously using the mirror and up-down flipping operations, the lower left can be obtained; The detection results of directional devices are converted to status values according to the preset direction-status mapping rules. The specific rules include: (1) The up direction of the air switch corresponds to the closed state, and the down direction corresponds to the tripped state; (2) The direction of the knob switch can be configured in advance according to actual needs. For example, the left direction corresponds to the standby state, the right direction corresponds to the running state, or different gears and start-stop states corresponding to the device functions; (3) The direction of the handle switch can be pre-configured according to actual needs. For example, the upward direction corresponds to the input state, and the downward direction corresponds to the exit state. The specific mapping relationship is set by the user according to the actual application scenario; During the post-processing of the detection results, the output content includes the status value of the device, detection box information, and class label information, where: (1) The indicator light only outputs the on or off state; (2) The pressure plate only outputs the open / closed or vacant state; (3) The air switch, knob switch, and handle switch only output the status value corresponding to the direction; The target detection model is a YOLOv5 or YOLOv8 network structure. The network model includes a feature extraction layer, a convolutional layer, and a classification layer, and completes device target detection and status recognition through an end-to-end training process; In the training data of the target detection model, the number of samples of directional devices is extended to more than three times the original number through data augmentation to improve the detection recall rate of the model for low-frequency direction states.

[0022] Specifically, step (1): Image data collection and annotation is the basic step in the training of the target detection model, the starting point of the entire method and technical process, and directly affects the implementation effects of subsequent data augmentation, model training, and detection accuracy. Specifically, this step aims to collect diverse device picture data and annotate the data in combination with the device category and status to ensure that the model can accurately identify the device status in various complex scenarios. Generally, to meet the detection requirements of the power production site, the picture collection should cover as comprehensively as possible the characteristics of different types of devices, including different device styles, working states, scene environments, and device angles. As an option, when collecting pictures, the performance of the device in special environments should also be considered, such as changes in lighting conditions or interference of complex backgrounds on the device edges.

[0023] In this embodiment, the collection scope of the picture data includes, but is not limited to, power equipment such as indicator lights, pressure plates, air switches, knob switches, and handle switches. Specifically, the collection of the indicator light should include the color differences of the device in the on and off states; the collection of the pressure plate should cover the open state, closed state, and possible vacant state; the collection of the air switch should record various possibilities of the device tripping direction, including four directions: up, down, left, and right; the knob switch and handle switch should cover eight states of the device pointing direction, including up, down, left, right, upper left, upper right, lower left, and lower right.

[0024] In a possible implementation, the captured images should preferably reflect the real environment of the power production site. For example, considering the possible reflection phenomenon on the surface of the equipment, the angle and intensity of the light source should be controlled during image capture, and multiple shots should be taken to cover different lighting conditions. In addition, to simulate the detection of equipment in complex scenarios, there may be other non-equipment interference targets (such as wires, other switches, etc.) in the image background. These complexities should be actively retained during image capture to enhance the model's adaptability to complex scenarios.

[0025] As the basis for data annotation, the captured images need to be annotated manually. Generally, the annotation categories are subdivided according to the type and status of the equipment. Specifically: The annotation of the indicator lights needs to be combined according to the color (such as red, green, yellow) and the on / off status (on or off). As an option, different category labels can be assigned to the on and off states under the same color so that the subsequent model can distinguish the status changes.

[0026] The annotation of the pressure plate needs to be combined with the style and opening / closing status of the equipment. For example, in some embodiments, the pressure plate may have different shape features (such as rectangular, circular, etc.), and these features can be used as additional information for annotation.

[0027] The annotation of the air switch needs to be clearly classified according to the tripping direction of the equipment, usually including four directions: up, down, left, and right. To enhance the model's understanding of the direction change, additional descriptions of the direction category can be added during annotation.

[0028] The annotation of the rotary switch and the handle switch is classified according to the direction pointed by the equipment. In a possible implementation, to achieve higher direction recognition accuracy, position information descriptions such as the relative position between the knob and the switch body can be added during annotation based on the direction.

[0029] In some embodiments, the annotation categories can be stored in the form of a label file. The label file should contain the following basic information: The coordinates of the detection box of the target equipment in the image, usually represented by the upper left corner coordinates (x 1 , y 1) and the lower right corner coordinates (x 2 , y 2) .

[0030] The category label corresponding to each detection box, and the label value can be a unique identifier composed of the type and status of the equipment.

[0031] For directional equipment, the label file should clearly record the label value of the equipment direction. For example, if the direction of an air switch is "up", its label value can be defined as "direction - up".

[0032] After the annotation is completed, to ensure the accuracy and consistency of the data, the annotation results can be manually reviewed. In some embodiments, the annotation files can be checked to see if they correspond to the original images by random sampling to avoid omission and errors of class labels or coordinate information.

[0033] In a possible extended manner, the device images collected and annotated can be sourced from multiple channels. For example, images can be obtained through on-site camera devices or historical images can be extracted from existing power equipment operation monitoring systems. For the use of historical images, it is also necessary to ensure the accuracy of their scene and status information through screening.

[0034] Step (2) is a key link in the preprocessing of image data, directly affecting the training effect and generalization ability of the target detection model. Generally, in the power equipment scenario, the changes in equipment status are diverse and complex. Some equipment statuses or directions (such as the "lower left" or "upper right" directions of a knob switch) may have a low frequency of occurrence in the original data, resulting in poor learning effects of the model on these statuses. Therefore, when constructing the training dataset, it is necessary to expand the number of samples through various data augmentation methods and dynamically adjust the class labels for directional equipment to ensure that the model can fully adapt to the equipment statuses in different scenarios. As an option, different enhancement strategies can be combined according to the characteristics of the equipment such as illumination, direction, and angle changes to optimize the sample distribution.

[0035] In this embodiment, for the collected and annotated device image data, various enhancement methods are first used to expand the pictures. Specifically, the enhancement methods include but are not limited to mirroring, flipping, brightness and saturation transformation, multi-scale scaling, mosaic splicing, and grayscale processing, etc. As an option, for directional equipment (such as air switches, knob switches, handle switches), when performing enhancement operations related to direction, the class labels need to be dynamically adjusted to ensure that the class labels are consistent with the features of the enhanced pictures.

[0036] In some embodiments, for the mirror operation, the class labels of directional equipment need to be updated according to the horizontal flipping rule. For example, if the original device class is "direction - left", after horizontal mirroring, its class label needs to be adjusted to "direction - right"; if the original class is "direction - right", it is adjusted to "direction - left". Similarly, the vertical flipping operation also needs to update the directional class labels. For example, for a device with the original class of "direction - up", after vertical flipping, it needs to be adjusted to "direction - down"; if the original class is "direction - down", it needs to be adjusted to "direction - up".

[0037] For a knob switch or a handle switch, the "upper left" direction is adjusted to "lower right", the "upper right" direction is adjusted to "lower left", the "lower left" direction is adjusted to "upper right", and the "lower right" direction is adjusted to "upper left".

[0038] For the air switch, only the categories in the horizontal and vertical directions need to be updated, and the diagonal categories are not involved.

[0039] Generally, the enhanced image needs to correspond one-to-one with the original annotation information. To ensure the effectiveness and consistency of the enhanced data, the annotation information of each enhanced image needs to be updated to the new label file. The coordinates of the detection box recorded in the label file need to be consistent with the size of the enhanced image. Specifically, if multi-scale scaling is performed on the original image, the update rule for the detection box coordinates is: x′ = s x ·x, y′ = s y ·y, w′ = s x ·w, h′ = s y ·h where x, y are the coordinates of the center point of the original detection box, w, h are the width and height of the detection box; s x , s y is the scaling ratio of the image in the horizontal and vertical directions; x′, y′, w′, h′ are the coordinates of the center point and the size of the detection box after scaling.

[0040] As a possible implementation, grayscale processing can be used to simulate the performance of the device in low-light or monochromatic light environments. The enhanced grayscale image needs to maintain the integrity of the original annotation file, but there is no need to adjust the category labels. In some embodiments, the grayscale image is combined with the original color image to increase the diversity of the dataset.

[0041] For mosaic splicing processing, the enhancement method can combine four images into one large image, and at the same time update the coordinates of the detection box of the target device in each image. For example, if four images are spliced into a large image with a size of 2W × 2H, and the size of each image is W × H, the update formula for the original detection box is: For the image in the upper left area, the detection box coordinates do not need to be adjusted.

[0042] For the image in the upper right area, the horizontal coordinate needs to be added with the image width W: x′ = x + W For the image in the lower left area, the vertical coordinate needs to be added with the image height H; y′ = y + H For the image in the lower right area, the horizontal coordinate needs to be added with W, and the vertical coordinate needs to be added with H After the data augmentation operation is completed, the finally extended dataset should cover all possible states of the device. As an implementation strategy, the augmented dataset can ensure that there are at least a certain number of samples for each device state. For example, for the "lower left" direction state of the knob switch, the amount of augmented data should reach more than three times the number of original samples, thus effectively solving the problem of insufficient model detection ability in rare states.

[0043] In some embodiments, the augmented dataset needs to be randomized before model training to avoid overfitting problems caused by sample bias introduced by specific augmentation strategies. Specifically, the augmented dataset can be mixed using random sampling, and different types of augmented samples can be evenly distributed into the training set and the validation set to improve the generalization ability of the model. Step (3) is to learn the state features and class information of the device based on the image data after the aforementioned data augmentation processing, so as to achieve accurate detection and state recognition of the device. Generally, the object detection model needs to have efficient feature extraction capabilities and diverse scene adaptation capabilities to meet the detection requirements of device state changes in complex environments in power production sites. As an option, using a lightweight and high-precision deep learning network, such as YOLOv5 or YOLOv8, is an effective way to achieve the above goals. Specifically, the training process of the model not only needs to consider the diversity of the input data, but also needs to design appropriate loss functions for device categories and states, and improve the performance of the model through optimization algorithms.

[0044] In this embodiment, the YOLOv8 network is used for training the object detection model. The YOLOv8 model has strong real-time detection capabilities and an end-to-end training structure, and is composed of an input module, a feature extraction module, and a detection head module. The input module is used to receive the picture data after data augmentation and perform normalization processing on it. Generally, the picture data will be normalized for pixel values, and the pixel values will be adjusted to the range of 0, 1.

[0045] The feature extraction module is based on the CSPNet (CrossStagePartialNetwork) structure and is used to extract multi-scale feature information. CSPNet processes the input features by dividing them into two parts and then fuses the feature outputs of the two parts, thereby reducing the computational complexity and improving the detection performance. In a possible implementation, the backbone network of the feature extraction module consists of multiple convolutional layers and residual modules, and multi-scale features are output at different feature layers for detecting devices of different sizes.

[0046] The detection head module is based on the Feature Pyramid Network (FPN) and the Path Aggregation Network (PANet), and fuses the extracted features to output the category and detection box information of the device. As an option, different detection box output formats can be adopted, such as directly outputting the center point coordinates and size information, or outputting the four vertex coordinates of the bounding box.

[0047] During the training process of the model, the design of the loss function is crucial. Generally, the loss function needs to comprehensively consider the positioning accuracy of the device detection box, the classification accuracy of the device category, and the classification ability of the foreground and background. In this embodiment, the loss function: Model input prediction value, True value of image annotation, Model weight parameters, Input image, Number of true targets in a single batch.

[0048] l obj = -ω obj (y obj logσp + 1 - y obj log1 - σp)) ω obj : Foreground-background loss weight, y obj : True foreground or background category, p: Predicted foreground score; l cls = -ω cls (y cls logσp + 1 - y cls log1 - σp)) ω cls : Class loss weight, y cls : True single-class binary classification value, p: Predicted single-class score; l box = ω box (1 - IOUy box ,gt) ω box : Class loss weight, y box : Predicted target box coordinates, gt: True (annotated) target box coordinates; Weight parameter gradient and descent optimization strategy (using SGDM): η: Optimization learning rate, Gradient, β: Momentum coefficient (default 0.9), m t : Exponential moving average gradient of the t-th iteration.

[0049] In some embodiments, to improve the training effect, a multi-scale training strategy can be adopted. Specifically, the input resolution of the picture is randomly adjusted during each training iteration, and the common resolution range is from 320×320 to 640×640. This method can enhance the model's detection ability for targets of different scales.

[0050] Step (4) is the core step in the model inference stage, and its function is to identify and detect the status of the devices in the image to be detected based on the trained object detection model. Generally, device detection needs to consider complex conditions in the actual scenario, including complex device backgrounds and overlapping targets. Therefore, relying solely on the detection results initially output by the model may have problems such as duplicate targets, low-confidence results, or overlapping detection frames. To solve the above problems, the present invention uses the non-maximum suppression (NMS) algorithm to optimize the initial detection results, so as to remove duplicate targets and ensure the uniqueness and accuracy of the final detection results.

[0051] In this embodiment, the image to be detected is first analyzed by the object detection model, and the results output by the model include information such as the position of the detection frame, the confidence of the detection frame, and the device category label. Specifically, the position of the detection frame is represented by the upper-left corner coordinates (x 1 ,y 1 ) and the lower-right corner coordinates (x 2 ,y 2 ); the confidence of the detection frame is the probability value that the model predicts this detection frame as the target device area, denoted as P conf ; the category label represents the classification result of the device status in the detection frame, denoted as l cls .

[0052] Generally, the object detection model may generate multiple overlapping detection frames for the same device, especially when there is a complex background near the target device or the device edge is unclear. To remove these duplicate targets, the present invention processes the detection results output by the model through the non-maximum suppression (NMS) algorithm. As an option, the NMS algorithm uses the confidence of the detection frame and the overlap ratio as criteria, preferentially retains the detection frame with the highest confidence, and removes other detection frames whose overlap ratio exceeds the set threshold.

[0053] Specifically, the processing steps of the NMS algorithm include: First, all detection frames are sorted in descending order according to the confidence value P conf of the detection frame.

[0054] Then, starting from the detection frame with the highest confidence, calculate the overlap ratio between this detection frame and the remaining detection frames, and the overlap ratio is based on the intersection over union.

[0055] Finally, when the overlap ratio of a detected bounding box with the selected high-confidence detected bounding boxes exceeds a threshold (usually set to 0.5), then remove this detected bounding box. This process is repeated until all detected bounding boxes are processed.

[0056] In a possible implementation, to further improve the accuracy of the detection results, the confidence threshold and overlap ratio threshold of the NMS algorithm can be dynamically adjusted. For example, in a scenario where target devices are densely distributed, the set threshold can be set to a lower value (such as 0.4) to avoid losing targets due to overlapping detected bounding boxes; while in a scenario where target devices are relatively sparse, the set threshold can be set to a higher value (such as 0.6) to reduce redundant detected bounding boxes.

[0057] In some embodiments, to further optimize the processing efficiency of NMS, a soft NMS algorithm can be used to replace the traditional NMS algorithm. The soft NMS algorithm reduces the confidence value of detected bounding boxes with a relatively large overlap ratio instead of directly removing them, thereby reducing the risk of misdeleting targets.

[0058] In the present invention, the output results of the NMS algorithm include: The unique detected bounding box for each device category.

[0059] The final confidence value of each detected bounding box.

[0060] The device category label corresponding to each detected bounding box.

[0061] In some embodiments, the processing results of the NMS algorithm can be further filtered in combination with business requirements. For example, for the indicator light detection task, only the detection results with a confidence higher than 0.8 are retained; while for directional devices (such as rotary switches and air switches), the confidence threshold can be lowered to ensure the integrity of the direction state.

[0062] Step (5) is the last link in the device state recognition process and also an important step for the business transformation of the target detection results. Generally, the results output by the target detection model include detected bounding boxes, category labels, and confidence values, and this information can characterize the physical location and state of the device. However, to meet the state monitoring requirements in the actual power production site, it is necessary to post-process the target detection results, merge the target results of the same type of devices, and only retain the core information related to the device state to avoid the output of redundant information.

[0063] Specifically, for indicator lights, the target detection model may output multiple category results, such as "green light on", "red light on", "other types of lights on", etc. In the post-processing stage, these detailed category results will be merged and simplified, and only two states of "on" or "off" are output; For the pressure plate, the object detection model may output the open and closed state categories of different types of pressure plates, such as "Type A pressure plate open", "Type B pressure plate open", etc. In the post-processing stage, these category results will be merged and simplified, and only the two states of "open" or "closed" will be output.

[0064] In this embodiment, first, the detection results optimized by non-maximum suppression (NMS) are classified and merged. Generally, for the detection boxes of the same device category, they can be merged according to the category label and the spatial distribution of the detection boxes, and only one state value is retained for subsequent output. For example, if there are multiple detection boxes in the indicator light category, the detection boxes of the same category need to be merged, and the detection result with the highest confidence is selected as the final output.

[0065] As an option, for directional devices (such as rotary switches, air switches, and handle switches), the mapping of the state value needs to be combined with the device direction and the category label. Specifically, in the category label of the directional device, the direction information of the device is usually included, such as "direction - up", "direction - left", etc. According to the pre-set direction-state mapping rules, the direction category can be converted into a specific device state value. The specific rules are as follows: The "up" direction of the air switch is mapped to the "closed state", and the "down" direction is mapped to the "tripped state"; The "left" direction of the rotary switch is mapped to the "standby state", and the "right" direction is mapped to the "operating state"; The "up" direction of the handle switch is mapped to the "connected state", and the "down" direction is mapped to the "disconnected state".

[0066] In a possible implementation, for devices that require special state differentiation, such as a rotary switch with multiple functional states, the state mapping rules can be further refined. For example: If the direction of the rotary switch is "upper right", it can be mapped to the "semi-automatic operating state"; If the direction is "lower left", it can be mapped to the "fault standby state".

[0067] During the process of extracting the state value, redundant information irrelevant to the actual application also needs to be removed. For example, in the detection of the indicator light, only the state values of "on" or "off" need to be output, and the details of the detection box position or brightness change do not need to be recorded. In some embodiments, if the necessary state information is missing in the detection result of a certain device, it can be marked as "unknown state" when outputting the state value, so that the user can perform manual verification in the subsequent process.

[0068] Specifically, the process of extracting the state value can be divided into the following key steps: Read the list of detection results optimized by NMS, and extract the category label and confidence value of each detection box.

[0069] For the same type of detection boxes, they are merged according to the spatial distribution of the center points of the detection boxes. The merging rules can be based on the overlapping area of the detection boxes or the Euclidean distance of the center points. For example, when the distance between the center points of two detection boxes is less than a certain threshold (such as half of the physical width of the device), it is considered that they belong to the same target device.

[0070] According to the class label and orientation information of the detection box, combined with the preset mapping rules, the final status value of the device is generated.

[0071] In some embodiments, to improve the accuracy of the status value output, the detection results with confidence values lower than the threshold can also be excluded. For example, if the confidence value of a detection box is lower than 0.6, the status information of this detection box will not be output. This way of confidence filtering can effectively reduce the false alarm ratio in the detection results.

[0072] After the post-processing is completed, the final status value will be output in a structured format for docking with the power equipment monitoring system. For example, the JSON format can be used to record the status information of each device, which includes fields such as device type, status value, detection timestamp, etc.

[0073] Step (6) is the last link of the device status recognition method and also an important part of realizing the business logic conversion based on the detection results. Generally, the results output by the target detection model and the post-processing module contain information such as the class, orientation, or status of the device. These information need to be further transformed into clear and definite status values through the mapping rules to meet the actual requirements of monitoring the operation status of devices in the power production site. As an option, the design of the mapping rules can be customized according to the device type and application scenario to ensure that the output of the final status value is both accurate and meaningful.

[0074] In this embodiment, first, the post-processing result in the foregoing step (5) is used as the input data, which already contains the optimized device class label and orientation information. Specifically, for each detection result, it includes field information such as device class, orientation label, confidence value, etc. In the status mapping stage, according to the device class and the preset orientation-status mapping rules, the orientation label is transformed into a specific device status value.

[0075] Generally, the mapping rules for different types of devices are quite different. For example, for an air switch, its orientation label usually directly corresponds to the switch status. The "up" orientation is mapped to the "closed state", and the "down" orientation is mapped to the "tripped state". For a knob switch, the orientation label needs to be mapped more complexly in combination with the function of the device. For example, the "left" orientation represents the "standby state", and the "right" orientation represents the "operating state". The specific mapping rules are as follows: Air switch: The direction "up" corresponds to the state "switch on".

[0076] The direction "down" corresponds to the state "switch off".

[0077] The direction "left" or "right" can respectively correspond to custom states, such as "maintenance" or "standby".

[0078] Rotary switch: The direction "left" indicates the "standby state".

[0079] The direction "right" indicates the "operating state".

[0080] The direction "upper right" can indicate the "semi-automatic operating state".

[0081] The direction "lower left" can indicate the "fault standby state".

[0082] Handle switch: The direction "up" indicates the "connected state".

[0083] The direction "down" indicates the "disconnected state".

[0084] In a possible implementation, to improve the generality of the mapping rules, the mapping relationship between device categories and states can be stored through a configuration file. The storage format of the mapping rules can adopt structured data formats such as JSON or XML.

[0085] In this embodiment, to ensure the accuracy of the state mapping, the matching process of the mapping rules requires multiple rounds of verification. Specifically: First, filter out the corresponding mapping rules according to the device category. For example, for an air switch, only the direction mapping rules of the air switch need to be loaded, and the rules of other devices are ignored.

[0086] Second, match the direction label in the detection result with the mapping rules. For example, if the detection result is "direction - up", directly search for the mapping rules and output the "switch on" state.

[0087] If the direction label in the detection result does not find a matching value in the mapping rules, this state needs to be marked as an "unknown state" for manual inspection.

[0088] In some embodiments, to improve the output efficiency of the state values, a batch processing method can be used to map multiple detection results. For example, in the same power cabinet, there may be multiple devices, and their detection results can load all the corresponding mapping rules at once and perform parallel processing on all the detection results. This method can significantly reduce the time overhead of repeatedly loading the rules.

[0089] Embodiment 2: Please refer to the appendix Figure 2, an indicator device status detection and recognition system for power production sites, characterized by including: An image acquisition module, which is used to acquire device pictures of indicator lights, pressure plates, air switches, knob switches, and handle switches; a data annotation module, which is used to manually annotate the acquired pictures and classify them. A data enhancement module, which is used to perform enhancement operations such as mirroring, flipping, multi-scale scaling, etc. on the annotated pictures and adjust the directional category labels. A target detection model training module, which is used to train a target detection model based on deep learning. A target detection module, which is used to detect the device category and status information in the pictures. A status post-processing module, which is used to merge the detection results of the same type of devices and output the final status value according to the direction-status mapping rule.

[0090] Specifically, the system of the present invention is composed of an image acquisition module, a data annotation module, a data enhancement module, a target detection model training module, a target detection module, and a status post-processing module. Each module closely cooperates through standardized data flow to jointly complete the tasks of power equipment status detection and recognition. The overall logic of the system takes device image data as the core, forming a complete closed loop from data acquisition to status value output. The cooperation between modules ensures the accuracy, robustness of the detection, and the efficient execution of the business logic.

[0091] As the starting point of the system, the image acquisition module is responsible for collecting device image data. The acquired images need to cover various states and different angles of the devices, considering the complexities such as light changes, occlusion, and background interference in the actual scenario to ensure the diversity of the data. The acquired images are directly transmitted to the data annotation module, which manually annotates the device images. The output annotation information includes device category, direction label, detection box coordinates, etc., forming a preliminary standardized data set to provide a basis for subsequent processing.

[0092] The standardized data set output by the data annotation module is transmitted to the data enhancement module. The data enhancement module expands the annotated data, using various enhancement methods such as mirroring, flipping, brightness and saturation transformation, multi-scale scaling, and grayscale processing to generate diverse training samples. For directional devices, the data enhancement module also dynamically adjusts the category labels to ensure the consistency between the labels and the image content. The output of the data enhancement module provides high-quality and high-coverage training data for model training, which can significantly improve the adaptability of the model in the actual scenario.

[0093] The enhanced dataset enters the object detection model training module. The training module uses a deep learning object detection model (such as YOLOv8) for multiple rounds of iterative optimization to build the final detection model. This module calculates the classification error, localization error, and device status prediction error based on a customized loss function, and adjusts the model parameters through the gradient descent algorithm to ensure that the model can accurately identify the status characteristics of different types of devices. After training, the model parameter file output by the object detection model training module provides core capability support for the object detection module.

[0094] The object detection module loads the model parameter file output by the object detection model training module and processes the actually input image to be detected. The module outputs preliminary results such as device detection boxes, class labels, and confidence values. Due to the possible existence of multiple detection boxes or results with low confidence, the object detection module closely cooperates with the status post-processing module, and optimizes the detection results through the non-maximum suppression (NMS) algorithm to remove duplicate targets and only retain the detection box with the highest confidence.

[0095] Based on the optimized results of the object detection module, the status post-processing module calls the preset direction-status mapping rules according to the device category and direction label to convert the detection results into specific device status values. For example, for an air switch, "direction - up" is mapped to "closed state", and "direction - down" is mapped to "tripped state"; for a knob switch, "direction - left" is mapped to "standby state", and "direction - right" is mapped to "operating state". The status post-processing module further filters the detection results, eliminates the status information with a confidence lower than the threshold, and outputs the final status value in a structured format for easy external system call and business integration.

[0096] The modules of the system achieve efficient collaboration through standardized interfaces, which not only ensures the seamless connection of data flow but also makes the overall architecture highly scalable. The functions of each module are independent and interdependent. For example, the annotation information output by the data annotation module directly determines the effects of the data enhancement module and the model training module, while the performance of the object detection module directly affects the accuracy of the status post-processing module. Finally, through the collaborative work of the modules, the system completes the full-process task from device images to status values, can meet the needs of diverse scenarios in power production sites, and provides an efficient and intelligent solution for device status monitoring.

[0097] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method and system for detecting and identifying the status of indicating equipment in a power production site, characterized in that: The following steps are involved: (1) Collect image data of indicator lights, pressure plates, air switches, knob switches, and handle switches. The images must cover conditions of different types, styles, scenes, and angles. Manually annotate the collected image data and divide the data categories according to the equipment category and status. Indicator lights are annotated by color and on / off status, pressure plates are annotated by style and open / close status, air switches are annotated by tripping direction, and knob switches and handle switches are annotated by pointing direction. (2) Performing data enhancement operations on the labeled image data, including mirroring, flipping, brightness and saturation transformation, multi-scale scaling, mosaic stitching and grayscale processing, and dynamically adjusting the category label of the enhanced image of the directional device. The mirroring operation adjusts the category label from left to right, and the upside-down flipping operation adjusts the category label from top to bottom; (3) Use the enhanced image data to train the target detection model. The target detection model is built based on a deep learning algorithm. By inputting image data, the device detection box and category label are output. A custom loss function is used to optimize the parameters during model training. (4) Input the image to be detected into the trained object detection model, predict the possible device area and category information in the image, remove duplicate detection targets through the non-maximum suppression algorithm, and only retain the detection results with the highest confidence in each category; (5) Post-process the test results, merge the test results of similar equipment, and only output the equipment status value, among which the indicator light only outputs the on or off state, the pressure plate only outputs the open, closed or empty state, and the air switch, knob switch and handle switch only output the direction state; (6) According to the detection result of the directional device and the preset direction-state mapping rule, the detection result is mapped to the state, and the final state value of the directional device is generated and output.

2. A method and system for detecting and identifying the status of indicative equipment in a power production site according to claim 1, characterized in that: The custom loss functions used in the object detection model training process include foreground and background classification loss, category classification loss, and bounding box regression loss, where: (1) The foreground and background classification loss is calculated by the cross entropy between the probability value of predicting whether the target is a device area and the true label value, which is used to measure the model's ability to distinguish between device and non-device areas; (2) The category classification loss is calculated by the cross entropy between the predicted probability value of the device category and the true category label, which is used to measure the accuracy of the model in identifying the device category; (3) The bounding box regression loss is obtained by calculating the similarity between the predicted bounding box and the true bounding box. The similarity metric can be the intersection over union (IOU) or improved IOU formulas such as CIOU, DIOU, and SIOU.

3. A method and system for detecting and identifying the status of indicative equipment in a power production site according to claim 1, characterized in that: The optimization strategy used in the target detection model training process is based on the stochastic gradient descent method. The model parameters are updated by the following formula during the optimization process: w t+1 =w t -η·m t Among them, w t represents the model parameters of the tth iteration, η represents the learning rate, m t Represents the exponentially weighted average of the current gradient.

4. A method and system for detecting and identifying the status of an indicating device in a power production site according to claim 1, characterized in that: The non-maximum suppression algorithm is used to remove targets with low confidence or high bounding box overlap, and only retains the detection frame with the highest confidence in each type of device. The condition for deduplication is that the overlap ratio between bounding boxes is greater than a set threshold, and the threshold can be 0.5 or set according to actual needs.

5. A method and system for detecting and identifying the status of indicative equipment in a power production site according to claim 1, characterized in that: When the data augmentation operation includes mirroring and flipping, the category label of the directional device is dynamically adjusted according to the augmentation operation, wherein: (1) The mirror operation adjusts the left direction to the right direction, and the right direction to the left direction; (2) The upside-down flip operation adjusts the upper direction to the lower direction, and the lower direction to the upper direction; (3) Mirror and flip operations can be used alone or in combination to generate more transformation effects, such as: The device orientation in the original image is upper right, and by flipping it upside down you get lower right; The upper left can be obtained by mirroring; By using the mirroring and up-down flipping operations at the same time, we can get the lower left.

6. A method and system for detecting and identifying the status of indicative equipment in a power production site according to claim 1, characterized in that: The detection result of the directional device is converted into a state value according to a preset direction-state mapping rule, and the specific rules include: (1) The upward direction of the air switch corresponds to the closed state, and the downward direction corresponds to the tripped state; (2) The direction of the knob switch can be configured in advance according to actual needs. For example, the left direction corresponds to the standby state, the right direction corresponds to the running state, or it can correspond to different gears and start / stop states according to the function of the equipment; (3) The direction of the handle switch can be configured in advance according to actual needs. For example, the upper direction corresponds to the input state, and the lower direction corresponds to the exit state. The specific mapping relationship is set by the user according to the actual application scenario.

7. A method and system for detecting and identifying the status of indicative equipment in a power production site according to claim 1, characterized in that: During the post-processing of the detection results, the output content includes the device status value, detection box information and category label information, where: (1) The indicator light only outputs the on or off state; (2) The pressure plate only outputs the open or closed or empty state; (3) Air switches, knob switches and handle switches only output status values ​​corresponding to the direction.

8. A method and system for detecting and identifying the status of indicative equipment in a power production site according to claim 1, characterized in that: The target detection model is a YOLOv5 or YOLOv8 network structure. The network model includes a feature extraction layer, a convolution layer and a classification layer. Device target detection and state recognition are completed through an end-to-end training process.

9. A method and system for detecting and identifying the status of indicative equipment in a power production site according to claim 1, characterized in that: In the training data of the target detection model, the number of samples of directional devices is expanded to more than three times the original number of samples through data enhancement to improve the model's detection recall rate for low-frequency directional states.

10. A system for detecting and identifying the status of an indicating device in a power production site, according to a method for detecting and identifying the status of an indicating device in a power production site according to any one of claims 1 to 9, characterized in that: include: An image acquisition module is used to acquire device images of indicator lights, pressure plates, air switches, knob switches, and handle switches; The data annotation module is used to manually annotate the collected images and classify them into categories; The data enhancement module is used to perform enhancement operations such as mirroring, flipping, and multi-scale scaling on the annotated images and adjust the directional category labels; The target detection model training module is used to train the target detection model based on deep learning; The object detection module is used to detect the device category and status information in the image; The state post-processing module is used to merge the detection results of similar devices and output the final state value according to the direction-state mapping rule.

Citation Information

Patent Citations

  • Model training method, image processing method and related device

    CN114155365A

  • Multi-state knob switch identification method

    CN114202731A

  • Sample data generation method and device, electronic equipment and storage medium

    CN115601616A

  • Aerial image rotating target detection method based on annular smooth label

    CN116597324A

  • Method and system for identifying state of disconnecting switch

    CN117058572A

Cited By

  • Power equipment intelligent identification method and system based on deep learning

    CN120894640A