Equipment state lamp identification system and method based on Yolov4

Through the Yolov4-based device status light recognition system, the training process is optimized by using Mosaic data enhanced and improved loss function, which solves the problems of inefficiency and poor detection effects of traditional algorithms in device status light monitoring, and achieves efficient and accurate device status light recognition and early warning.

CN120374919AActive Publication Date: 2025-07-25BENXI IRON & STEEL (GROUP) INFORMATION AUTOMATION CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510300578.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-25
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

In the prior art, the monitoring efficiency of equipment status lights in industrial environments is low, manual monitoring consumes manpower and is prone to missed reports. Traditional algorithms such as RCNN and YOLO3 have problems with redundant calculations and poor target detection effects.

Method used

The Yolov4-based device status light recognition system is adopted to build data sets through Mosaic data augmentation and labeling, and feature extraction is combined with CSPDarknet53 structure, CBM and CSP modules. The loss function is improved using CIOU Loss and DIOU_NMS, and self-adversarial training and cross-small batch standardization are introduced, the training process is optimized, and the status light recognition model is deployed for real-time detection.

Benefits of technology

It improves the recognition efficiency and accuracy of equipment status lights, realizes efficient detection of multiple targets and small targets, promptly triggers early warning mechanisms, and ensures the normal operation of the equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374919A_ABST
    Figure CN120374919A_ABST
Patent Text Reader

Abstract

The invention discloses an equipment state lamp recognition system and method based on Yolov4, and relates to the technical field of equipment intelligent recognition, and the method comprises the steps: controlling a robot to collect an image of an equipment state lamp in a machine room, carrying out the Mosaic data enhancement and marking of the image, and constructing a data set; constructing a state lamp identification model, wherein the state lamp identification model comprises an input layer, a backbone network, a fusion network and a prediction output; the loss function is improved by using CIOU Loss and DIOUNMS, and the regression precision is improved; training the state lamp recognition model, and introducing self-confrontation training and cross-small-batch standardization to obtain a weight file; and deploying the trained state lamp identification model, and identifying an equipment state lamp image. According to the method, the state lamp recognition model is constructed, so that the problem that the detection effect of a current target detection algorithm on multiple targets and small targets is poor is solved, and the method has the advantages of being high in real-time recognition efficiency, good in small target recognition effect and high in recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of device intelligent recognition, and specifically to a device status light recognition system and method based on Yolov4. Background Art

[0002] There are a large number of indicator lights in the industrial environment to judge whether each device is working properly. When the indicator light shows a certain state, it indicates that a fault has occurred and needs to be dealt with by the staff in time. Therefore, it is necessary to monitor the status of the indicator lights in real time. However, in traditional manual monitoring, long-term monitoring is a heavy and boring task. When there are many indicator lights, it is almost impossible for manual work to achieve comprehensive and accurate monitoring. At the same time, abnormal situations are relatively few, so manual monitoring will cause huge waste of manpower and low efficiency. For the traditional industrial indicator light status monitoring, it not only consumes a lot of manpower but also is prone to missed reports due to the negligence of the monitoring personnel, so the monitoring efficiency is low.

[0003] With the development of deep learning in computer vision in recent years, many scholars have been using neural networks to solve related problems such as object detection and achieved good results. Currently, the main methods for detecting status lights are two algorithms: RCNN and YOLO3. R-CNN has redundant calculations. Because R-CNN first generates candidate regions and then performs convolution on the regions, there will be a certain degree of overlap in the candidate regions, resulting in repeated operations when the subsequent CNN extracts features. At the same time, since R-CNN stores the extracted features and then uses SVM for classification, it requires a larger storage space. Since CNN feature extraction is used, it is necessary to scale the candidate boxes, but in actual situations, the selected boxes have various sizes, which will cause target deformation; while the Yolo3 algorithm, as an object detection algorithm, has problems such as long training time, excessive hardware consumption, and poor detection effects for multi-targets and small targets. Summary of the Invention

[0004] The purpose of the present invention is to provide a device status light recognition system and method based on Yolov4 to solve the problems proposed in the prior art.

[0005] To achieve the above purpose, the present invention provides the following technical solution: A device status light recognition method based on Yolov4, the method includes:

[0006] Controlling a robot to collect images of the status lights of the equipment in the computer room, and performing Mosaic data augmentation and annotation on the images to construct a data set;

[0007] Build a status light recognition model, where the status light recognition model includes an input layer, a backbone network, a fusion network, and a prediction output; the input layer is used to input device status light images, the backbone network includes a CBM module and a CSP module for preliminary feature extraction; the fusion network includes an SPP module and a PANet module for multi-level feature fusion; the prediction output generates target detection results, and the target detection results include the position and category information of the status light;

[0008] Use CIOU Loss and DIOU_NMS to improve the loss function and enhance the regression accuracy;

[0009] Based on the dataset, train the status light recognition model, introduce self-adversarial training and cross-mini-batch normalization to optimize the training process and obtain a weight file;

[0010] Deploy the trained status light recognition model to recognize device status light images.

[0011] According to the above solution, the Mosaic data augmentation includes:

[0012] Randomly select four images from the dataset and obtain the annotation information of each image; the annotation information includes the target box coordinates and category labels;

[0013] Perform flipping, scaling, and color gamut change processing on each selected image respectively;

[0014] Stitch the four processed images into a large image in four directions to generate a composite image containing multiple targets and multiple backgrounds;

[0015] Adjust the annotation boxes of each image according to the position of the stitched large image to ensure the accurate position of the annotation boxes of each target in the stitched image and obtain the augmented image.

[0016] According to the above solution, the backbone network is based on the CSPDarknet53 structure and includes a CBM module and a CSP module;

[0017] The backbone network receives the image input by the input layer, and after preliminary feature processing by the CBM module, it successively passes through 5 CSP modules for progressive downsampling; a convolutional kernel with a fixed size is set in front of each CSP module, and the stride is set to a specific value to ensure that the size of the feature map gradually decreases while extracting deeper semantic information; after passing through the CBM module and the CSP module, preliminary feature extraction is completed, and a feature map with certain semantic information is output.

[0018] According to the above solution, the CBM module includes a convolutional layer, a batch normalization layer, and a Mish activation function; the convolutional layer performs a convolution operation on the input image through a convolution kernel to extract the basic features in the image; the batch normalization layer normalizes the features output by the convolutional layer; the Mish activation function introduces non-linear factors to enhance the gradient flow, enabling the status light recognition model to learn the complex relationships between features;

[0019] The formula of the Mish activation function is as follows:

[0020] Mish(x) = x · tanh(ln(1 + e x ));

[0021] where Mish(x) represents the output value after being processed by the Mish activation function; x represents the input of the activation function; tanh represents the hyperbolic tangent function; the Mish activation function can provide a small gradient when the input value is small to avoid gradient explosion; it can provide a large gradient when the input value is large to avoid gradient vanishing, which helps the neural network better learn complex data patterns and features;

[0022] The CSP module divides the feature map of the basic layer into two parts: one part continuously extracts high-level semantic features through the stacking of residual blocks, and the other part is directly connected to the final output through a small amount of processing; through the cross-stage hierarchical structure, the two parts of the features are merged to reduce the repeated calculation of gradient information;

[0023] On each feature map, multiple anchor boxes of different sizes are preset for detecting targets of different sizes; the sizes of the anchor boxes are obtained through clustering analysis based on the size distribution of the targets in the dataset to ensure that the model can effectively detect small, medium, and large targets; through the fusion and detection of multi-scale feature maps, the adaptability of the model to multi-scale targets is improved;

[0024] Using DropBlock, randomly mask continuous region blocks in the feature map, thereby forcing the network to rely on other parts for prediction, thus improving the generalization ability of the model. DropBlock is a regularization method to alleviate overfitting;

[0025] According to the above solution, the fusion network includes an SPP module and a PANet module;

[0026] The SPP module performs max-pooling operations on the input feature map at different scales, processes the feature map respectively using pooling kernels of various specifications to expand the perceptual range of the features and obtain context feature information covering different scales; splice the feature maps of different scales to generate a feature representation with rich context information;

[0027] The PANet module efficiently extracts and fuses multi-level feature maps through a top-down and bottom-up bidirectional feature transfer mechanism, combined with a feature fusion method of tensor connection.

[0028] According to the above solution, the top-down process includes that the PANet module receives multi-level feature maps from the backbone network and the SPP module, and the multi-level feature maps have different resolutions and semantic information; through upsampling operations, along the top-down path, the strong semantic information of the high-level feature maps is gradually transferred to the low-level feature maps, enabling the low-level feature maps to fuse high-level semantic information.

[0029] The bottom-up process includes that the PANet module starts from the low-level feature maps and gradually transfers strong localization information upward; through convolution operations, the localization information of the low-level feature maps is fused with the semantic information of the high-level feature maps to generate intermediate feature maps with strong semantic and strong localization capabilities.

[0030] The bidirectional feature transfer mechanism includes that the bottom-up path transfers low-level localization information to the high-level feature maps, which is combined with the semantic information transferred by the top-down path to form a bidirectional flow of multi-level features.

[0031] The tensor connection includes that the PANet module uses tensor connection to deeply fuse the feature maps generated by the top-down and bottom-up paths.

[0032] According to the above solution, the dataset is divided into a training set, a test set, and a validation set.

[0033] Using the training set, train the status light recognition model; using the validation set, combined with loss and accuracy, evaluate the performance of the status light recognition model; using the test set, for independent evaluation after the status light recognition model is trained to verify the generalization ability of the status light recognition model.

[0034] In each training iteration, generate adversarial samples through self-adversarial training and use the adversarial samples to optimize the detection ability of the status light recognition model; standardize the features through cross-mini-batch normalization to ensure the stability of the training process; combine the improved loss functions of CIOU Loss and DIOU_NMS to optimize the regression accuracy and prediction box screening effect of the status light recognition model.

[0035] The formula for CIOU is as follows:

[0036]

[0037] Among them, CIOU represents the Complete Intersection over Union, which is an index used to measure the overlapping relationship between the predicted bounding box and the ground truth box in object detection; IOU represents the Intersection over Union, which is used to represent the ratio of the intersection area to the union area of the predicted bounding box and the ground truth box; ρ represents the Euclidean distance; b represents the center point coordinates of the predicted bounding box; gt represents the ground truth value; b gt represents the center point coordinates of the ground truth box; c represents the diagonal distance of the smallest closed region that can simultaneously contain the predicted bounding box and the ground truth box; α represents the weight parameter; v represents the parameter for measuring the aspect ratio consistency, which is used to evaluate the similarity of the aspect ratios of the predicted bounding box and the ground truth box;

[0038] Calculate the loss function LOSS based on CIOU CIOU , and the formula is as follows:

[0039]

[0040] Among them, LOSS CIOU represents the loss function based on CIOU, which is used to optimize the regression relationship between the predicted bounding box and the ground truth box during the training of the object detection model, so that the speed and accuracy of the predicted bounding box regression are higher;

[0041] When the loss value no longer decreases significantly after being evaluated using the validation set, it is determined that the status light recognition model converges, the training iteration is terminated, and a weight file is obtained.

[0042] According to the above solution, the self-adversarial training includes a first stage and a second stage;

[0043] The first stage includes, during the training process of the status light recognition model, performing an adversarial attack on the input device status light image, and generating an adversarial sample by adjusting the image pixel values;

[0044] The second stage includes using the generated adversarial sample as the input and training the status light recognition model in a normal manner, so that the status light recognition model can detect the target in the adversarial sample;

[0045] The cross-mini-batch normalization includes dividing the training set into multiple mini-batch data, each mini-batch data containing partial device status light images and corresponding annotation information; calculating the mean and variance of the features in each mini-batch data, and collecting statistical data across multiple mini-batch data in the training set; using the collected statistical data to normalize the features to ensure the stability of the feature distribution and accelerate the convergence of the status light recognition model.

[0046] According to the above solution, the weight file stores the connection weights and biases of each part of the status light recognition model, which determines the feature extraction and prediction capabilities of the model for input data. By loading the weight file, the status light recognition model can efficiently recognize the device status light image and output the position and category information of the target.

[0047] A device status light recognition system based on Yolov4, which includes: a data acquisition module, a status light recognition module, a data storage module, and a user interaction module;

[0048] The data acquisition module includes an image capture module and a data preprocessing module. The image capture module is used to control the movement of the robot and the adjustment of the lifting rod, locate the device to be detected, and capture the status light image of the computer room equipment. The data preprocessing module is used to perform data enhancement and annotation on the captured image to construct a data set.

[0049] The data storage module is used to store and manage the captured image data, annotation files, data sets, training logs, weight files, and warning records, provide data backup and recovery, and ensure data security.

[0050] The status light recognition module includes a model construction module, a training optimization module, and a model deployment module. The model construction module is used to construct a status light recognition model, which includes an input layer, a backbone network, a fusion network, and a prediction output. The training optimization module uses the training set, improves the loss function using CIOU Loss and DIOU_NMS, introduces self-adversarial training and cross-mini-batch normalization, trains the status light recognition model, and generates a weight file. The model deployment module deploys the trained status recognition model, performs real-time detection on the input device status light image, and outputs the position and category information of the status light.

[0051] The user interaction module includes a user interface module and a warning module. The user interface module displays the recognition results and warning information of the device status light, provides an operation interface for training and deploying the status light recognition model, generates a visualization report, and displays the system operation status and recognition effect. The warning module receives the recognition results of the model deployment module, judges whether the color and status of the status light are normal, and when an abnormal status of the status light is detected, triggers a warning mechanism, sends a warning message to the user interface, records the warning information, and generates a log file.

[0052] Compared with the prior art, the beneficial effects of the present invention are:

[0053] 1. Based on the CSPDarknet53 structure, the present invention combines the CBM module and the CSP module, reduces the repeated calculation of gradient information, and improves the feature extraction efficiency;

[0054] 2. The present invention introduces the SPP module and adopts an improved PANet module, significantly improving the detection accuracy;

[0055] 3. The present invention uses CIOU Loss and DIOU_NMS to improve the loss function, enhancing the regression accuracy and the effect of predicting box screening;

[0056] 4. The present invention efficiently and accurately realizes the identification of the device status light through an automatic and intelligent process, improving the efficiency of status light identification;

[0057] 5. The present invention combines a warning module to judge the status light abnormality in real time and trigger a warning mechanism, effectively preventing equipment failures and ensuring the normal operation of the computer room equipment. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 is a step flow chart of a method for identifying device status lights based on Yolov4 according to the present invention;

[0059] Figure 2 is a structural schematic diagram of a system for identifying device status lights based on Yolov4 according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0061] Embodiment: As Figure 1 - Figure 2 shown, the present invention provides a technical solution, a method for identifying device status lights based on Yolov4, and the method includes:

[0062] S1. Control the robot to collect images of the device status lights in the computer room, and perform Mosaic data enhancement and annotation on the images to construct a data set;

[0063] Specifically, use a robot equipped with a high-resolution industrial camera to collect images of the device status lights in the computer room. For example, the resolution of the industrial camera is 1920×1080, the frame rate is 30fps, and it has good light adaptability and can clearly capture the device status lights under different lighting conditions; the robot moves to the target cabinet through a path planning algorithm, adjusts the lifting rod to the position of the device to be detected, and collects images of the device status lights; for example, about 13,000 images are collected, covering scenarios of different devices and different states (such as green lights, red lights, yellow lights, and unlit lights);

[0064] Furthermore, perform Mosaic data augmentation on the collected images; randomly select four images from the dataset, e.g., images A, B, C, and D; and obtain the annotation information for each image; the annotation information includes the target box coordinates and class labels; perform flipping, scaling, and color gamut change processing on each selected image respectively, e.g., the scaling ratio range is from 0.8 to 1.2, the brightness adjustment range is from -20% to +20%, and the contrast adjustment range is from -15% to +15%; splice the four processed images into a large image in the four directions of upper left, upper right, lower left, and lower right, e.g., image A is placed in the upper left corner, image B is placed in the upper right corner, image C is placed in the lower left corner, and image D is placed in the lower right corner; generate a composite image containing multiple targets and multiple backgrounds; adjust the annotation boxes of each image according to the position of the spliced large image to ensure the accurate position of the annotation box of each target in the spliced image, and obtain the enhanced image; for example, the annotation box coordinates of image A are adjusted from (x1, y1, x2, y2) to (x1’, y1’, x2’, y2’) to ensure the accurate position of the annotation box of each target in the spliced image; and use image annotation software to annotate the status lights in the image.

[0065] Furthermore, there is a folder VOC2007 containing the dataset package under the root directory of the YOLOV4 folder. VOC2007 contains three sub-folders, namely Annotations, ImageSets, and Images. Among them, the Images folder stores all the dataset photos of the status lights for training, the Annotations folder stores the.xml files corresponding to each photo after annotation, and the sub-folders in ImageSets store the training set and test set image information.

[0066] S2. Build a status light recognition model. The status light recognition model includes an input layer, a backbone network, a fusion network, and a prediction output; the input layer is used to input the device status light image, the backbone network includes a CBM module and a CSP module for preliminary feature extraction; the fusion network includes an SPP module and a PANet module for multi-level feature fusion; the prediction output generates the target detection result, and the target detection result includes the position and class information of the status light.

[0067] Specifically, the backbone network is based on the CSPDarknet53 structure and includes CBM modules and CSP modules. The backbone network receives the image input by the input layer, and after preliminary feature processing through the CBM modules, it undergoes progressive downsampling through 5 CSP modules in sequence. A fixed 3×3 convolutional kernel is provided before each CSP module, and the stride is set to a specific value to ensure that the size of the feature map gradually decreases while extracting deeper semantic information. For example, if the input image is 608×608, after passing through 5 CSP modules, the change rule of the feature map is: 608-304-152-76-38-19; finally, a 19×19 feature map is output. Through the CBM modules and CSP modules, preliminary feature extraction is completed, and a feature map with certain semantic information is output.

[0068] Further, the CBM module includes a convolutional layer, a batch normalization layer, and a Mish activation function. The convolutional layer performs a convolution operation on the input image through a convolutional kernel to extract the basic features in the image. The batch normalization layer normalizes the features output by the convolutional layer. For example, the convolutional kernel size is 3×3, the stride is 1, and the padding is 1. The Mish activation function introduces a non-linear factor to enhance the gradient flow, enabling the status light recognition model to learn the complex relationships between features.

[0069] The formula of the Mish activation function is as follows:

[0070] Mish(x) = x·tanh(ln(1 + e x ));

[0071] Among them, Mish(x) represents the output value after being processed by the Mish activation function; x represents the input of the activation function; tanh represents the hyperbolic tangent function. The Mish activation function can provide a small gradient when the input value is small to avoid gradient explosion; it can provide a large gradient when the input value is large to avoid gradient disappearance, which helps the neural network better learn complex data patterns and features.

[0072] Further, the CSP module divides the feature maps of the base layer into two parts: one part stacks residual blocks to continuously extract high-level semantic features, and the other part is directly connected to the final output through a small amount of processing; through a cross-stage hierarchical structure, the two parts of the features are merged to reduce the repeated calculation of gradient information; on each feature map, multiple anchor boxes of different sizes are preset for detecting targets of different sizes; the sizes of the anchor boxes are obtained through clustering analysis according to the size distribution of the targets in the dataset to ensure that the model can effectively detect small, medium, and large targets; through the fusion and detection of multi-scale feature maps, the adaptability of the model to multi-scale targets is improved; using DropBlock, continuously masked regions in the feature map are randomly masked, forcing the network to rely on other parts for prediction, thereby improving the generalization ability of the model;

[0073] Specifically, the fusion network includes an SPP module and a PANet module; the SPP module performs maximum pooling operations on the input feature maps at different scales, for example: pooling kernel sizes: 1×1, 5×5, 9×9, and 13×13; the feature maps are processed separately using pooling kernels of various specifications to expand the perceptual range of the features and obtain context feature information covering different scales. For example, for a 13×13 input feature map, a 5×5 pooling kernel is used for pooling with padding = 2, and the pooled feature map is still 13×13 in size; the feature maps of different scales are concatenated to generate a feature representation with rich context information; the PANet module efficiently extracts and fuses multi-level feature maps through a top-down and bottom-up two-way feature transfer mechanism combined with a feature fusion method of tensor connection.

[0074] Further, the top-down process includes that the PANet module receives multi-level feature maps from the backbone network and the SPP module, and the multi-level feature maps have different resolutions and semantic information; through the upsampling operation, along the top-down path, the strong semantic information of the high-level feature map is gradually transmitted to the low-level feature map, enabling the low-level feature map to fuse the high-level semantic information; the bottom-up process includes that the PANet module starts from the low-level feature map and gradually transmits strong localization information upward; through convolution operations, the localization information of the low-level feature map is fused with the semantic information of the high-level feature map to generate an intermediate feature map with strong semantic and strong localization capabilities; the two-way feature transfer mechanism includes that the bottom-up path transmits the low-level localization information to the high-level feature map and combines it with the semantic information transmitted by the top-down path to form a two-way flow of multi-level features; the tensor connection includes that the PANet module uses tensor connection to deeply fuse the feature maps generated by the top-down and bottom-up paths.

[0075] S3. Use CIOU Loss and DIOU_NMS to improve the loss function and enhance the regression accuracy;

[0076] Specifically, an improved loss function combining CIOU Loss and DIOU_NMS is used to optimize the regression accuracy and the prediction box screening effect of the status light recognition model; the positioning loss uses CIOU, and the NMS for prediction box screening becomes DIOU_NMS. The IOU is used as the regression optimization loss. CIOU takes into account factors such as the distance, overlap rate, and scale between the target and the anchor, making the target regression more stable.

[0077] The CIOU is as follows:

[0078]

[0079] Among them, CIOU is expressed as the Complete Intersection over Union, which is an index used to measure the overlap relationship between the prediction box and the ground truth box in object detection; IOU is expressed as the Intersection over Union, which is used to represent the ratio of the intersection area to the union area of the prediction box and the ground truth box; ρ is expressed as the Euclidean distance; b represents the center point coordinates of the prediction box; gt represents the ground truth value; b gt represents the center point coordinates of the ground truth box; c represents the diagonal distance of the smallest closed region that can simultaneously contain the prediction box and the ground truth box; α represents the weight parameter; v represents the parameter for measuring the aspect ratio consistency, which is used to evaluate the similarity of the aspect ratios of the prediction box and the ground truth box.

[0080] Furthermore, calculate the loss function LOSS CIOU based on CIOU, and the formula is as follows:

[0081]

[0082] Among them, LOSS CIOU represents the loss function based on CIOU, which is used to optimize the regression relationship between the prediction box and the ground truth box during the training of the object detection model, making the regression speed and accuracy of the prediction box higher.

[0083] S4. Based on the dataset, train the status light recognition model, introduce self-adversarial training and cross mini-batch normalization to optimize the training process and obtain the weight file;

[0084] Specifically, divide the dataset into a training set, a test set, and a validation set; use the training set to train the status light recognition model; use the validation set to evaluate the performance of the status light recognition model in combination with the loss and accuracy; use the test set for independent evaluation after the status light recognition model is trained to verify the generalization ability of the status light recognition model.

[0085] Further, in each training iteration, adversarial samples are generated through self-adversarial training, and the detection ability of the status light recognition model is optimized using the adversarial samples; the features are standardized through cross-mini-batch normalization to ensure the stability of the training process; the cross-mini-batch normalization includes dividing the training set into multiple mini-batch data, for example: each mini-batch contains 32 images; each mini-batch data contains partial device status light images and corresponding annotation information; calculating the mean and variance of the features in each mini-batch data, and collecting statistical data across multiple mini-batch data in the training set; using the collected statistical data to standardize the features to ensure the stability of the feature distribution and accelerate the convergence of the status light recognition model;

[0086] Further, the self-adversarial training includes a first stage and a second stage; the first stage includes, during the training process of the status light recognition model, performing adversarial attacks on the input device status light images, and generating adversarial samples by adjusting the image pixel values; for example: adjusting the pixel values of image A from (100, 150, 200) to (90, 135, 180); the second stage includes using the generated adversarial samples as input and training the status light recognition model in a normal manner so that the status light recognition model can detect the targets in the adversarial samples;

[0087] Further, when the loss value no longer significantly decreases after evaluation using the validation set, it is determined that the status light recognition model converges, the training iteration is terminated, and a weight file is obtained;

[0088] Further, the weight file stores the connection weights and biases of each part of the status light recognition model, which determines the feature extraction and prediction capabilities of the model for input data; by loading the weight file, the status light recognition model can efficiently recognize device status light images and output the position and category information of the targets.

[0089] S5. Deploy the trained status light recognition model to recognize device status light images;

[0090] Specifically, input device status light images, for example: with a size of 416×416, through model inference, and through normalization and non-maximum suppression ratio, the final rectangular prediction bounding box is obtained to frame out the device status light, and the prediction result is the position of the device status light that frames out the device light in the prediction image, and at the same time, the color of the bounding box is the color of the status light; for example: the detected position of the status light is (200, 300, 250, 350), the category is "green light", and the bounding box color is green. This is only for illustrative purposes and not for limitation.

[0091] The present invention provides another technical solution, a device status light recognition system based on Yolov4, which system comprises: a data acquisition module, a status light recognition module, a data storage module, and a user interaction module;

[0092] The data acquisition module includes an image capture module and a data preprocessing module; the image capture module is used for controlling the movement of the robot and the adjustment of the lifting rod, positioning the device to be detected, and capturing the images of the status lights of the equipment in the computer room; the data preprocessing module is used for performing data enhancement and annotation on the captured images to construct a data set;

[0093] The data storage module is used for storing and managing the captured image data, annotation files, data sets, training logs, weight files, and warning records, providing data backup and recovery to ensure data security;

[0094] The status light recognition module includes a model construction module, a training optimization module, and a model deployment module; the model construction module is used for constructing a status light recognition model, and the status light recognition model includes an input layer, a backbone network, a fusion network, and a prediction output; the training optimization module uses the training set, improves the loss function using CIOU Loss and DIOU_NMS, introduces self-adversarial training and cross-mini-batch normalization, trains the status light recognition model, and generates a weight file; the model deployment module deploys the trained status recognition model, performs real-time detection on the input device status light images, and outputs the position and category information of the status lights;

[0095] The user interaction module includes a user interface module and a warning module; the user interface module displays the recognition results and warning information of the device status lights, provides an operation interface for training and deploying the status light recognition model, generates a visualization report, and displays the system operation status and recognition effect; the warning module receives the recognition results of the model deployment module, judges whether the color and status of the status lights are normal, and when an abnormal status of the status lights is detected, triggers a warning mechanism, sends warning information to the user interface, records the warning information, and generates a log file.

[0096] The present invention provides another technical solution. In a computer room, a device status light recognition system based on Yolov4 is deployed to monitor the status lights of the equipment in the computer room in real time and detect equipment anomalies in a timely manner;

[0097] The trained model is deployed to a high-performance server in the computer room and is connected to the data acquisition module in real time. When a new device status light image is input, the status light recognition model can respond quickly and complete the analysis of the image within 0.8 seconds, accurately outputting the position and category information of the status lights;

[0098] The warning module constantly monitors the recognition results output by the model deployment module. The warning module determines whether the color and status of the status light are normal according to preset rules. For example, when the status light is green, it is in a normal state; when the status light is orange, it is in a slightly abnormal state; when the status light is red, it is moderately abnormal; when the status light is off, it indicates an abnormality of the status light or the device. This is only for illustrative purposes and not restrictive. When the status light of a certain server is recognized as red by the status light recognition model, the warning module immediately captures the abnormal information of the status light and triggers the warning mechanism;

[0099] After triggering the warning, the warning module quickly displays the abnormal information of the device in a prominent red pop-up window on the user interface module. The abnormal information includes the device location, the abnormal situation of the status light, etc.;

[0100] The warning module details and records this warning information, including the warning time, abnormal device information, recognition results, etc., and generates a complete log file, which is stored in the data storage module for convenient subsequent query and analysis.

[0101] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights.

Claims

1. A method for identifying device status lights based on Yolov4, characterized in that: The method includes: Controlling the robot to collect images of the status lights of the equipment in the machine room, performing Mosaic data augmentation and annotation on the images, and constructing a dataset; Constructing a status light recognition model, which includes an input layer, a backbone network, a fusion network, and a prediction output; the input layer is used to input the images of the equipment status lights, the backbone network includes CBM modules and CSP modules for preliminary feature extraction; the fusion network includes SPP modules and PANet modules for multi-level feature fusion; the prediction output generates target detection results, and the target detection results include the position and category information of the status lights; Using CIOU Loss and DIOU_NMS to improve the loss function and enhance the regression accuracy; Based on the dataset, training the status light recognition model, introducing self-adversarial training and cross-mini-batch normalization, optimizing the training process, and obtaining a weight file; Deploying the trained status light recognition model to recognize the images of the equipment status lights.

2. The device status light recognition method based on Yolov4 according to claim 1, wherein: The Mosaic data augmentation includes: Randomly selecting four images from the dataset and obtaining the annotation information of each image; the annotation information includes the coordinates of the target box and the category label; Performing flipping, scaling, and color gamut change processing on each of the selected images respectively; Stitching the four processed images into a large image in four directions to generate a composite image containing multiple targets and multiple backgrounds; Adjusting the annotation boxes of each image according to the position of the stitched large image to ensure the accurate position of the annotation boxes of each target in the stitched image, and obtaining the augmented image.

3. The method for recognizing the status lights of equipment based on Yolov4 according to claim 1, wherein: The backbone network is based on the CSPDarknet53 structure and includes CBM modules and CSP modules; The backbone network receives the images input by the input layer, and after preliminary feature processing through the CBM module, successively passes through 5 CSP modules for progressive downsampling; a convolutional kernel with a fixed size is provided before each CSP module, and the stride is set to a specific value to ensure that the size of the feature map gradually decreases while extracting deeper semantic information; after passing through the CBM module and CSP modules, the preliminary feature extraction is completed, and a feature map with certain semantic information is output.

4. The method for recognizing the status lights of equipment based on Yolov4 according to claim 3, wherein: The CBM module includes a convolutional layer, a batch normalization layer, and a Mish activation function; the convolutional layer performs a convolutional operation on the input image through a convolutional kernel to extract the basic features in the image; The batch normalization layer normalizes the features output by the convolutional layer; the Mish activation function introduces non-linear factors to enhance the gradient flow and enables the status light recognition model to learn the complex relationships between features; The CSP module divides the feature maps of the basic layer into two parts: one part continuously extracts high-level semantic features through the stacking of residual blocks, and the other part is directly connected to the final output through a small amount of processing; through the cross-stage hierarchy, the two parts of the features are merged to reduce the repeated calculation of gradient information.

5. The method for identifying the device status light based on Yolov4 according to claim 1, characterized in that: The fusion network includes an SPP module and a PANet module; The SPP module performs maximum pooling operations on the input feature maps at different scales, processes the feature maps respectively using pooling kernels of various specifications, expands the perceptual range of the features, and obtains context feature information covering different scales; The feature maps of different scales are concatenated to generate a feature representation with rich context information; The PANet module efficiently extracts and fuses multi-level feature maps through a top-down and bottom-up two-way feature transfer mechanism and a feature fusion method combining tensor connection.

6. The method for identifying the device status light based on Yolov4 according to claim 5, characterized in that: The top-down includes that the PANet module receives multi-level feature maps from the backbone network and the SPP module, and the multi-level feature maps have different resolutions and semantic information; through the upsampling operation, along the top-down path, the strong semantic information of the high-level feature maps is gradually transmitted to the low-level feature maps, so that the low-level feature maps fuse the high-level semantic information; The bottom-up includes that the PANet module starts from the low-level feature maps and transmits strong localization information upward level by level; through the convolution operation, the localization information of the low-level feature maps is fused with the semantic information of the high-level feature maps to generate intermediate feature maps with strong semantic and strong localization capabilities; The two-way feature transfer mechanism includes that the bottom-up path transmits the low-level localization information to the high-level feature maps and combines it with the semantic information transmitted by the top-down path to form a two-way flow of multi-level features; The tensor connection includes that the PANet module uses tensor connection to deeply fuse the feature maps generated by the top-down and bottom-up paths.

7. The method for identifying the device status light based on Yolov4 according to claim 1, characterized in that: The data set is divided into a training set, a test set, and a validation set; The training set is used to train the status light recognition model; the validation set is used to evaluate the performance of the status light recognition model in combination with the loss and accuracy; the test set is used for independent evaluation after the status light recognition model is trained to verify the generalization ability of the status light recognition model; In each training iteration, adversarial samples are generated through self-adversarial training, and the adversarial samples are used to optimize the detection ability of the status light recognition model; the features are normalized through cross-mini-batch normalization to ensure the stability of the training process; the regression accuracy and the prediction box screening effect of the status light recognition model are optimized by combining the improved loss function of CIOU Loss and DIOU_NMS; When the loss value no longer decreases significantly after being evaluated using the validation set, it is determined that the status light recognition model converges, the training iteration is terminated, and a weight file is obtained.

8. The method for identifying a device status light based on Yolov4 according to claim 7, characterized in that: The self-adversarial training includes a first stage and a second stage; The first stage includes, during the training process of the status light recognition model, performing adversarial attacks on the input device status light images, and generating adversarial samples by adjusting the image pixel values; The second stage includes using the generated adversarial samples as inputs and training the status light recognition model in a normal manner so that the status light recognition model can detect the targets in the adversarial samples; The cross-mini-batch normalization includes dividing the training set into multiple mini-batch data, and each mini-batch data contains partial device status light images and corresponding annotation information; Calculate the mean and variance of the features in each mini-batch data, and collect statistical data across multiple mini-batch data in the training set; use the collected statistical data to normalize the features to ensure the stability of the feature distribution and accelerate the convergence of the status light recognition model.

9. The method for identifying a device status light based on Yolov4 according to claim 1, characterized in that: The weight file stores the connection weights and biases of each part of the status light recognition model, and determines the feature extraction and prediction capabilities of the model for input data; by loading the weight file, the status light recognition model can efficiently recognize device status light images and output the position and category information of the targets.

10. A device status light recognition system based on Yolov4, which is applied to a device status light recognition method based on Yolov4 described in any one of claims 1-9, and is characterized in that: The system includes: a data acquisition module, a status light recognition module, a data storage module, and a user interaction module; The data acquisition module includes an image capture module and a data preprocessing module; the image capture module is used to control the movement of the robot and the adjustment of the lifting rod, locate the device to be detected, and capture the images of the equipment status lights in the computer room; the data preprocessing module is used to perform data augmentation and annotation on the captured images to construct a data set; The data storage module is used to store and manage the captured image data, annotation files, data sets, training logs, weight files, and warning records, provide data backup and recovery, and ensure data security; The status light recognition module includes a model construction module, a training optimization module, and a model deployment module; the model construction module is used to construct a status light recognition model, and the status light recognition model includes an input layer, a backbone network, a fusion network, and a prediction output; the training optimization module uses the training set, improves the loss function using CIOU Loss and DIOU_NMS, introduces self-adversarial training and cross-mini-batch normalization, trains the status light recognition model, and generates a weight file; the model deployment module deploys the trained status recognition model, performs real-time detection on the input device status light images, and outputs the position and category information of the status lights. The user interaction module includes a user interface module and an early warning module; the user interface module displays the recognition results of the device status lights and early warning information, provides an operation interface for training and deploying the status light recognition model, generates a visualization report, and shows the system operation status and recognition effect; the early warning module receives the recognition results of the model deployment module, judges whether the color and status of the status lights are normal, and when detecting an abnormal status of the status lights, triggers an early warning mechanism, sends early warning information to the user interface, records the early warning information, and generates a log file.

Citation Information

Patent Citations

  • Signal machine identification system, method and device based on image identification

    CN118097618A