Communication terminal remote monitoring method based on combination of visible light and infrared imaging
By combining visible light and infrared imaging technology, the optimized YOLOv8 model is used for remote monitoring of communication terminal equipment, which solves the problem of difficulty in operation and maintenance management of communication terminal equipment in outdoor environments, and realizes efficient automatic operation and maintenance and remote inspection, improving operation and maintenance efficiency and security.
Patent Information
- Application Number
- CN202411964810.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
AI Technical Summary
In outdoor environments, the operation and maintenance management of communication terminal equipment is difficult, and remote video inspection cannot be achieved, resulting in insufficient operation and maintenance efficiency and security.
The remote monitoring method of communication terminal based on the combination of visible light and infrared imaging is adopted. The visible light camera collects indicator light status pictures and infrared imaging devices to collect temperature pictures, pre-process and label them, and trains with the optimized YOLOv8 model to achieve automatic operation and maintenance and remote inspection.
It realizes automatic operation and maintenance of communication terminal equipment and visual remote inspection, improves operation and maintenance efficiency and security, and optimizes system performance through intelligent analysis and adapts to changes in different scenarios and needs, providing guarantees for the stable operation of the distribution automation system.
Smart Images

Figure CN119942435A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning target detection, and in particular to a communication terminal remote monitoring method based on the combination of visible light and infrared imaging. Background Art
[0002] Communication engineering is an important branch of electronic engineering. Communication engineering mainly studies the generation of signals, transmission, exchange and processing of information, as well as computer communication, digital communication, satellite communication, optical fiber communication, cellular communication, personal communication, stratospheric communication, multimedia technology, information highway, digital program-controlled exchange and other issues. In the process of communication engineering construction, communication terminal boxes are often used. Communication terminals are direct tools for people to enjoy different information applications (communication services). They are responsible for providing users with a good user interface, completing the required business functions and accessing communication networks. However, the outdoor environment is harsh, and after long-term work, it often leads to difficulties in operation and maintenance control, and remote video inspection cannot be achieved. Summary of the invention
[0003] The purpose of the present invention is to provide a remote monitoring method for communication terminals based on the combination of visible light and infrared imaging, to realize automatic operation and maintenance of communication equipment and visual remote inspection, to improve operation and maintenance efficiency and safety, and to further optimize system performance by continuously collecting operation and maintenance data and performing intelligent analysis, to adapt to changes in different scenarios and needs, and to provide strong guarantees for the stable operation of the distribution automation system.
[0004] To achieve the above object, the present invention is implemented through the following technical solutions:
[0005] A remote monitoring method for communication terminals based on the combination of visible light and infrared imaging, using a visible light camera to collect indicator light status pictures on a communication terminal device, and using an infrared imaging device to collect temperature pictures of the communication terminal device when it is working; preprocessing the collected indicator light status pictures and temperature pictures to form a data set; annotating the data set; and putting the annotated data set into an optimized YOLOv8 model for training.
[0006] YOLOv8 model, including:
[0007] S1, data preprocessing;
[0008] S2, backbone network structure;
[0009] S3, FPN-PAN structure;
[0010] S4, decoupling head Detection head structure;
[0011] S5. Label allocation strategy.
[0012] In S1, the data preprocessing adopts the YOLOv5 strategy, and the model training adopts mosaic enhancement, mixed enhancement, spatial perturbation, and color perturbation.
[0013] In S2, the backbone network structure is the MobileNetV3 network, including the attention mechanism, channel attention mechanism, and spatial attention mechanism. The contents are as follows:
[0014] 1) The attention mechanism is used in the convolutional neural network (CNN) to automatically focus on the characteristic parts of the input image data while suppressing irrelevant parts;
[0015] 2) The channel attention mechanism is used to focus on the correlation between different channels and enhance the expression of features by learning the importance of each channel. The steps are as follows:
[0016] S201, feature map input
[0017] Assume that the input feature map is represented as F in , whose shape is (C,H,W), where C is the number of channels, H and W are the height and width of the feature map respectively;
[0018] S202, global average pooling
[0019] Perform global average pooling on each channel to obtain a tensor with shape (C, 1, 1);
[0020] The global information of each channel is compressed into a scalar, and the formula is as follows:
[0021]
[0022] In formula ①, GAP(F in ) is the result of the global average pooling operation; H and W represent the height and width of the input feature map, respectively. in Represents the input feature map;
[0023] S203, fully connected layer or 1x1 convolution
[0024] The result of global average pooling is passed through a fully connected layer or 1x1 convolution to learn the interdependence between channels. Two fully connected layers are used, one for compression and the other for excitation.
[0025] S204, channel attention output
[0026] The attention weight α is combined with the input feature map F in Multiply channel by channel to get the channel attention weighted feature map F out , the formula is as follows:
[0027] F out=α⊙F in ②
[0028] In formula ②, α represents the attention weight, which is a value assigned to each channel of the feature map according to its importance; the formula recombines the features of different channels according to their importance;
[0029] 3) The spatial attention mechanism is used to focus on the importance of different spatial positions in the feature map and enhance the expression of features by learning the weight of each spatial position. The steps are as follows:
[0030] S211, feature map input
[0031] Assume that the feature map after channel attention is F out , whose shape is (C,H,W);
[0032] S212, channel average pooling and maximum pooling
[0033] The feature map is average pooled and max pooled in the channel dimension to obtain two tensors of shape (1, H, W). The formula is as follows:
[0034]
[0035] Formula ③ is average pooling, AvgPool(F out ) represents the feature map F out The result of the average pooling operation in the channel dimension; C represents the number of channels of the feature map; It means that the elements on all channels are summed and averaged;
[0036] Formula ④ is the maximum pooling, MaxPool(F out ) represents the feature map F out The result of the maximum pooling operation in the channel dimension; C represents the number of channels of the feature map;
[0037] S213, splicing
[0038] The results of average pooling and maximum pooling are concatenated in the channel dimension to obtain a tensor of shape (2, H, W). The formula is as follows:
[0039] F concat =Concat(AvgPool(F out ),MaxPool(F out )) ⑤
[0040] In formula ⑤, F concatRepresents the result of the concatenation operation, which represents a feature map with a shape of (2, Height, Width); Concat represents the concatenation operation, which is an operation that merges two or more tensors along a specified dimension; AvgPool(F out Represents the feature map F out The result of the average pooling operation; MaxPool(F out ) represents the feature map F out The result of the maximum pooling operation;
[0041] The concatenation operation is used for feature fusion in the network, using different information extracted by different pooling methods to enhance feature representation capabilities;
[0042] S214, 1x1 convolution
[0043] The concatenated feature map is passed through a 1x1 convolutional layer to reduce the number of channels to one channel, and a spatial attention weight tensor β with a shape of (1, H, W) is obtained;
[0044] S215, spatial attention output
[0045] The spatial attention weight β is combined with the feature map F after channel attention. out Multiply element by element to get the final weighted feature map F final , the formula is as follows:
[0046] F final =β⊙F out ⑥
[0047] In formula ⑥, β represents the spatial attention weight, which is used to redistribute weights to different spatial positions of the feature map; F out Represents the feature map after channel attention calculation;
[0048] Through the spatial attention weight β and the feature map F out This element-by-element multiplication operation is performed to strengthen or suppress features at different positions in the feature map for transfer to subsequent processing steps.
[0049] In S3, the FPN-PAN structure is used to construct the feature pyramid of YOLO, so that multi-scale information can be integrated.
[0050] In S4, the structure of the decoupled detection head is as follows: two parallel branches extract the category features and position features of the indicator light status image and temperature image respectively, and then use a layer of 1×1 convolution to complete the classification and positioning tasks of the indicator light status image and temperature image.
[0051] Two parallel branches are used to separately process different prediction targets in the object detection task.
[0052] In S5, the label assignment strategy includes classification loss VFLLoss (Varifocal Loss), which is as follows:
[0053] The classification loss VFLLoss (Varifocal Loss) is defined as follows:
[0054]
[0055] In formula ①, p is the predicted category score, p∈[0,1]; q is the predicted target score. If it is the real category, q is the loU between the prediction and the true value; if it is other categories, q is 0;
[0056] The classification loss VFL Loss uses an asymmetric parameter to weight the positive and negative samples. By only attenuating the negative samples, the unequal contribution of the foreground and background to the loss is completed. For the positive samples, q is used for weighting; for the negative samples, p is used. γ Demotion.
[0057] The images of the indicator light status on the communication terminal device include an indicator light image of a normal working status, an indicator light image of an abnormal working status, which is displayed in red, and an indicator light image of a stopped working status; the indicator light of the normal working status is green, the indicator light of the abnormal working status is red, and the indicator light of the stopped working status is not lit.
[0058] Compared with the prior art, the present invention has the following beneficial effects:
[0059] 1. The communication terminal box uses an intelligent computing unit (the intelligent computing unit is responsible for processing and analyzing data from various devices in the communication terminal device, such as temperature, humidity, equipment operating status and other data collected by the data acquisition assembly in the device, as well as image information collected by infrared and visible light binocular cameras, etc.) and an infrared imaging device and a visible light camera to provide operation and maintenance personnel with real-time equipment status monitoring and environmental monitoring, such as equipment temperature, operating status and abnormal conditions in the surrounding environment, etc., to improve operation and maintenance efficiency and safety;
[0060] 2. By continuously collecting operation and maintenance data and conducting intelligent analysis, the system performance can be further optimized, adapting to changes in different scenarios and needs, and providing strong guarantees for the stable operation of the distribution automation system. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 This is a schematic diagram of the overall network structure of YOLOv8.
[0062] Figure 2This is a schematic diagram of the MobileNetV3 network structure.
[0063] Figure 3 This is a schematic diagram of the YOLOv8-MobileNet-CSAM network structure. DETAILED DESCRIPTION
[0064] The present invention is described in detail below in conjunction with the accompanying drawings, but it should be noted that the implementation of the present invention is not limited to the following embodiments.
[0065] The following examples are implemented on the premise of the technical solution of the present invention, and provide detailed implementation methods and specific operation processes, but the protection scope of the present invention is not limited to the following examples. The methods used in the following examples are conventional methods unless otherwise specified.
[0066] Example 1
[0067] The present invention is applied to the remote communication and operation and maintenance of the power distribution automation system. Aiming at the protection and intelligent management of outdoor optical fiber communication equipment, the communication terminal box can replace the simple and single-function box of traditional communication equipment, and provide multiple functions such as modular design, intelligent operation and maintenance, and remote inspection for communication equipment. Figure 1 A remote monitoring method for communication terminals based on the combination of visible light and infrared imaging, using a visible light camera to collect indicator light status pictures on a communication terminal device, using an infrared imaging device to collect temperature pictures of the communication terminal device when it is working, and judging whether the communication terminal box is in a normal working state or an abnormal working state by the temperature; the indicator light status pictures on the communication terminal device include indicator light pictures of a normal working state, indicator light pictures of an abnormal working state, showing red, and indicator light pictures of a stopped working state; the indicator light of a normal working state is green, the indicator light of an abnormal working state is red, and the indicator light of a stopped working state is off; the collected indicator light status pictures and temperature pictures are preprocessed to form a data set; the data set is labeled; the labeled data set is put into the optimized YOLOv8 model for training.
[0068] The YOLOv8 algorithm was proposed by Glenn-Jocher. Compared with the YOLOv3 algorithm and the YOLOv5 algorithm, the main improvements include:
[0069] S1. Data preprocessing
[0070] YOLOv8’s data preprocessing adopts the YOLOv5 strategy, and uses four enhancement methods during training: mosaic enhancement (Mosaic), mixed enhancement (Mixup), spatial perturbation (random perspective), and color perturbation (HSV augment).
[0071] The YOLOv5 strategy specifically includes the following:
[0072] (1) Data preparation
[0073] -Collect and annotate datasets to ensure data quality and diversity;
[0074] -Data augmentation techniques, such as rotation, scaling, cropping, etc., to increase the generalization ability of the model;
[0075] (2) Model selection
[0076] -Choose the appropriate YOLOv5 model version according to task requirements, such as YOLOv5s, YOLOv5m, YOLOv5l, YOLOv5x, etc.;
[0077] - Consider the size of the model and the limitations of computing resources;
[0078] (3) Training configuration
[0079] -Set training parameters such as learning rate, batch size, number of training rounds, etc.;
[0080] -Select optimization algorithms, such as SGD, Adam, etc.;
[0081] -Use pre-trained models for transfer learning to speed up training and improve performance;
[0082] (4) Model training:
[0083] -Monitor the training process, including changes in the loss function, performance on the validation set, etc.;
[0084] -Adjust hyperparameters such as learning rate decay strategy, weight decay, etc. to optimize model performance;
[0085] (5) Model evaluation:
[0086] -Evaluate the model's accuracy, recall, mAP and other indicators on the test set;
[0087] -Analyze false detection cases and identify deficiencies of the model;
[0088] (6) Model deployment:
[0089] -Deploy the trained model to the target device, such as embedded systems, servers, etc.;
[0090] - Consider the model’s inference speed and hardware compatibility;
[0091] (7) Continuous Optimization:
[0092] -Continuously optimize models and training strategies based on actual application feedback;
[0093] - Follow the latest research progress and try to introduce new technologies and improvements.
[0094] Mosaic enhancement is a data enhancement technique widely used in computer vision, especially in target detection tasks; Mosaic enhancement is to improve the generalization ability and detection accuracy of computer vision models by increasing the diversity and complexity of training data; it is a simple and effective data enhancement method, which is of great significance for improving the performance of models in real-world scenarios. Mixup is a data enhancement technique widely used in deep learning, especially in the field of image classification; Mixup enhancement is to increase the diversity and complexity of training data, reduce overfitting, and improve the generalization ability and performance of the model; it is a simple and effective data enhancement method suitable for a variety of deep learning tasks. Spatial perturbation (random perspective) is a commonly used data enhancement technique in computer vision, used for image processing and target detection tasks. Spatial perturbation (Random Perspective) is to improve the robustness and performance of the model by increasing the diversity and complexity of training data; it is a simple and effective data enhancement method suitable for a variety of computer vision tasks; Color perturbation (HSV augment) is a commonly used data enhancement technique in computer vision; it is used for image processing and target detection tasks. Color perturbation (HSVAugment) improves the robustness and performance of the model by increasing the diversity and complexity of training data; it is a simple and effective data augmentation method suitable for a variety of computer vision tasks.
[0095] S2. Backbone network structure
[0096] The backbone network structure of YOLOv8 can be seen from YOLOv5. The architecture of the backbone network of YOLOv5 is very clear. Generally speaking, each layer of 3×3 convolution with a step size of 2 is used to downsample the feature map, and a C3 module is introduced to further enhance the features. The basic depth parameters of the C3 module are "3 / 6 / 9 / 3" respectively, which are used to make corresponding scaling according to models of different scales. In YOLOv8, this feature is also inherited. The original C3 modules are replaced by new C2f modules. The C2f module adds more branches to enrich the tributaries during gradient backpropagation.
[0097] Optimized YOLOv8 model
[0098] Achieving efficient and accurate detection on mobile devices has always attracted much attention. As one of the solutions, the YOLOv8 algorithm has made significant progress in detection accuracy and speed, but there is still room for improvement in specific application scenarios. Based on YOLOv8, this project replaces part of the backbone network of YOLOv8 with MobileNetV3 to reduce the amount of calculation and improve the lightweight degree of the YOLOv8 model, that is, it is necessary to ensure high computational efficiency while still being able to fully extract the rich features of the communication terminal device (the rich features of the communication terminal device are the diverse and complex information contained in the image of the mobile device, including the color, texture, shape, etc. of the communication terminal device. These features are crucial for target detection tasks because they can help the model better identify and classify objects in the image); in addition, the indicator lights of the communication terminal devices are special targets, and these special targets are often easily missed by the detection network. To address this problem, the study introduced the channel attention CSAM module to improve the detection effect of the improved YOLOv8 model on special targets.
[0099] MobileNetV3 is a lightweight network architecture proposed by the Google team. Compared with traditional convolutional neural networks, MobileNet network greatly reduces the number of model parameters and computational complexity at the expense of slightly lower accuracy, which makes it suitable for real-time image processing on devices with limited computing resources, such as mobile devices, embedded devices, and edge devices. After continuous improvements in the first two versions, the lightweight network MobileNetV3 has achieved significant improvements in performance and speed, and is currently widely sought after by academia and industry. The MobileNetV3 version mainly integrates the depth of MobileNetV1, separable convolution and the inverted linear bottleneck of MobileNetV2. In addition, new innovative features are introduced, including the use of the NetAdapt algorithm to automatically determine the optimal number of convolution kernels and channels, the SE channel attention mechanism to enhance the expression of features, and the new activation function h-swish, thereby comprehensively improving the performance of the network model; Up to now, MobileNetV3 provides two versions, Large and Small, the main difference between which lies in the requirements for computing resources and memory; The Small version strikes a balance between accuracy and speed, so the present invention selects MobileNetV3-Small as a model to replace the backbone of the YOLOv8 network, so as to meet the performance requirements for communication terminal detection; The modified network structure diagram is shown in Figure 2 ; Figure 2It consists of three parts: Backbone, Neck, and Head; Backbone consists of "Conv", "M-Netv3", and "SPPF"; it uses a series of convolution and deconvolution layers to extract features, and also uses residual connections and bottleneck structures to reduce the size of the network and improve performance; Neck consists of "Conv", "C2f", "Concat", and "Upsample"; it uses multi-scale feature fusion technology to fuse feature maps from different stages of Backbone to enhance feature representation capabilities; Head consists of "Detect"; it is responsible for the final target detection and classification tasks; "Conv" in the YOLOv8 network structure diagram is a convolution layer; it is the core component for processing image data; "M-Netv3" is a lightweight network architecture proposed by the Google team. Compared with traditional convolutional neural networks, the MobileNet network greatly reduces the number of model parameters and computational complexity at a slightly lower accuracy, which also makes it very suitable for real-time image processing on devices with limited computing resources; "SPPF" is a special variant of the spatial pyramid pooling module , which specifically optimizes the performance and efficiency of the pooling operation; the core function of the SPPF module is to quickly extract multi-scale features, enhance the detection capability of targets of different sizes, and optimize the computational efficiency; "Concat" means the operation of splicing multiple feature maps according to the channel dimension; this is a common deep learning operation, especially widely used in multi-scale feature fusion; through the Concat operation, YOLOv8 can better fuse shallow and deep features, thereby achieving higher accuracy and efficiency in target detection tasks; "Upsample" is upsampling, which is used to increase the spatial resolution (i.e., height and width) of the feature map; upsampling is a common operation in deep learning, especially in target detection tasks, it is used to restore low-resolution feature maps to higher resolutions in order to better detect small targets or restore more spatial details; through the Upsample operation, YOLOv8 can better handle multi-scale target detection tasks and improve the accuracy and robustness of the model; "Detect" usually refers to the detection head part of the network, which is responsible for generating the final target detection result from the input feature map; the detection head is an important part of the entire target detection model, which converts the feature information in the feature map into specific detection boxes and category probabilities.
[0100] MobileNetV3 network, including attention mechanism CSAM, channel attention mechanism, and spatial attention mechanism, the contents are as follows:
[0101] 1) The attention mechanism CSAM is used in the convolutional neural network CNN to automatically focus on the characteristic parts of the input image data while suppressing irrelevant parts; this mechanism can adaptively adjust the weights of different channels and spatial positions, thereby improving the ability of the improved YOLOv8 model to extract key information;
[0102] 2) The channel attention mechanism is used to focus on the correlation between different channels (feature maps) and enhance the expression of features by learning the importance of each channel;
[0103] The feature map is the output generated by the convolutional neural network in each layer; specifically, when the input image passes through the convolution layer, the convolution kernel will perform a convolution operation with the local area of the input to generate a new matrix, which is called a feature map; each point in the feature map represents the feature information of a small part of the input image; in the feature map, the channel is the third dimension of the three-dimensional tensor; each channel corresponds to an independent set of feature information; for example, in a feature map output by a convolution layer, there may be multiple channels, each channel captures a specific feature of the input image (such as edges, textures, colors, etc.); the feature information between channels may be independent of each other, but there may also be some correlation; the channel attention mechanism enhances the expressiveness of the feature map in the channel dimension by learning the importance of each channel in the feature map and adjusting the channel weights, so that the model can more effectively use useful feature information to complete the task.
[0104] Here are the steps:
[0105] S201, feature map input
[0106] Assume that the input feature map is represented as F in , whose shape is (C,H,W), where C is the number of channels, H and W are the height and width of the feature map respectively;
[0107] S202, global average pooling
[0108] Perform global average pooling (GAP) on each channel to obtain a tensor with a shape of (C, 1, 1). This step compresses the global information of each channel into a scalar, reflecting the importance of the channel in the overall feature map.
[0109] The global information of each channel is compressed into a scalar, and the formula is as follows:
[0110]
[0111] In formula ①, GAP(F in) is the result of the global average pooling operation; H and W represent the height and width of the input feature map, respectively. in Represents the input feature map;
[0112] S203, fully connected layer or 1x1 convolution
[0113] The result of global average pooling is passed through a fully connected layer or 1x1 convolution to learn the inter-channel dependencies. Usually, two fully connected layers are used in this step, which are called the squeeze and excitation processes respectively.
[0114] S204, channel attention output
[0115] The attention weight α is combined with the input feature map F in Multiply channel by channel to get the channel attention weighted feature map F out , the formula is as follows:
[0116] F out =α⊙F in ②
[0117] In formula ②, α represents the attention weight, which is a value assigned to each channel of the feature map according to its importance. These weights may be dynamically generated through the channel attention mechanism, and their main function is to highlight or suppress the features of specific channels in the input feature map. ⊙ represents element-by-element multiplication (Hadamard product), that is, the multiplication operation of the elements at corresponding positions of the matrix or tensor. This formula recombines the features of different channels according to their importance, which can make the model more sensitive to important information and help improve task performance, such as in image classification or target detection tasks;
[0118] 3) The spatial attention mechanism is used to focus on the importance of different spatial positions in the feature map and enhance the expression of features by learning the weight of each spatial position. The steps are as follows:
[0119] S211, feature map input
[0120] Assume that the feature map after channel attention is F out , whose shape is (C,H,W);
[0121] S212, channel average pooling and maximum pooling
[0122] The feature map is average pooled and max pooled in the channel dimension to obtain two tensors of shape (1, H, W). The formula is as follows:
[0123]
[0124] Formula ③ is average pooling, AvgPool(F out ) represents the feature map F out The result of the average pooling operation on the channel dimension. C represents the number of channels of the feature map. It means that the elements on all channels are summed and then averaged. Formula ④ is the maximum pooling, MaxPool(F out ) represents the feature map F out The result of the maximum pooling operation on the channel dimension. C represents the number of channels of the feature map;
[0125] S213, splicing
[0126] The results of average pooling and maximum pooling are concatenated in the channel dimension to obtain a tensor of shape (2, H, W). The formula is as follows:
[0127] F concat =Concat(AvgPool(F out ),MaxPool(F out )) ⑤
[0128] In formula ⑤, F concat Represents the result of the concatenation operation, which represents a feature map with a shape of (2, Height, Width). Concat represents the concatenation operation, which is an operation that merges two or more tensors along a specified dimension. AvgPool(F out Represents the feature map F out The result of the average pooling operation, MaxPool(F out ) represents the feature map F out The result of the maximum pooling operation. Such concatenation operations are often used for feature fusion in the network in order to utilize different information extracted by different pooling methods and enhance the ability of feature representation;
[0129] S214, 1x1 convolution
[0130] The concatenated feature map is passed through a 1x1 convolutional layer to reduce the number of channels to one channel, and a spatial attention weight tensor β with a shape of (1, H, W) is obtained;
[0131] S215, spatial attention output
[0132] The spatial attention weight β is combined with the feature map F after channel attention. out Multiply element by element to get the final weighted feature map F final , the formula is as follows:
[0133] F final =β⊙F out ⑥
[0134] In formula ⑥, β represents the spatial attention weight, which redistributes the weights of different spatial positions of the feature map to highlight or suppress the features of certain areas. out represents the feature map after channel attention calculation. ⊙ represents element-by-element multiplication, through the spatial attention weight β and the feature map F out This element-by-element multiplication operation can further strengthen or suppress features at different positions in the feature map so as to better transfer to subsequent processing steps;
[0135] The CSAM module combines the channel attention and spatial attention mechanisms to learn attention weights in the channel and spatial dimensions respectively, thereby enhancing the YOLOv8 model's ability to identify important features. This mechanism enables the YOLOv8 model to pay more attention to features that have an important impact on the task, thereby improving the model's performance.
[0136] Based on the YOLOv8-MobileNet network model, the present invention adds a CSAM attention mechanism module to the last layer of its backbone network, thereby realizing the model's ability to extract key information without increasing the complexity of the backbone network;
[0137] In summary, we finally get the network structure of YOLOv8-MobileNet-CSAM, see Figure 3 , Figure 3It consists of three parts: Backbone, Neck, and Head. It just replaces the backbone network with M-Netv3 and introduces the CSAM mechanism. Backbone consists of "Conv", "M-Netv3", "SPPF", and "CSAM". It uses a series of convolution and deconvolution layers to extract features, and also uses residual connections and bottleneck structures to reduce the size of the network and improve performance. Neck consists of "Conv", "C2f", "Concat", and "Upsample". It uses multi-scale feature fusion technology to fuse feature maps from different stages of Backbone to enhance feature representation. Head is composed of "Detect", which is responsible for the final target detection and classification tasks; "Conv" in the YOLOv8 network structure diagram is the convolution layer; it is the core component for processing image data; "M-Netv3" is a lightweight network architecture proposed by the Google team. Compared with traditional convolutional neural networks, the MobileNet network greatly reduces the number of model parameters and the amount of calculation under the premise of slightly lower accuracy, which also makes it very suitable for real-time image processing on devices with limited computing resources; "SPPF" is a special variant of the spatial pyramid pooling module, which specifically optimizes the performance and efficiency of the pooling operation; the core role of the SPPF module is to quickly extract multi-scale features and increase It has strong detection capabilities for targets of different sizes while optimizing computational efficiency. "CSAM" is an attention mechanism that is mainly used to automatically focus on the relevant feature parts of input data (such as images) in deep learning models (especially convolutional neural networks (CNNs)) while suppressing irrelevant parts. This mechanism can adaptively adjust the weights of different channels and spatial positions, thereby improving the model's ability to extract key information. "Concat" refers to the operation of concatenating multiple feature maps according to the channel dimension. This is a common deep learning operation, especially widely used in multi-scale feature fusion. Through the Concat operation, YOLOv8 can better fuse shallow and deep features, thereby achieving higher accuracy in target detection tasks. and efficiency; "Upsample" is upsampling, which is used to increase the spatial resolution (i.e., height and width) of the feature map; upsampling is a common operation in deep learning, especially in target detection tasks, it is used to restore low-resolution feature maps to higher resolutions in order to better detect small targets or restore more spatial details; through the Upsample operation, YOLOv8 can better handle multi-scale target detection tasks and improve the accuracy and robustness of the model; "Detect" usually refers to the detection head part of the network, which is responsible for generating the final target detection result from the input feature map; the detection head is an important part of the entire target detection model, which converts the feature information in the feature map into specific detection boxes and category probabilities.
[0138] The optimized YOLOv8-MobileNet-CSAM network model has high efficiency and lightweight structure. Compared with the original YOLOv8 network model, it improves the model's ability to extract key information and can remotely monitor the status of communication terminal devices in real time.
[0139] S3, FPN-PAN structure
[0140] YOLOv8 still uses the FPN-PAN structure to construct the feature pyramid of YOLO, so that multi-scale information can be fully integrated; in computer vision, "multi-scale information" refers to feature information of different scales (or different resolutions). Such as feature maps of indicator light status pictures, temperature pictures, etc. This information is very important in image processing and analysis because it allows the algorithm to capture different levels of features of objects in details of different resolutions;
[0141] Except that the C3 module in FPN-PAN is replaced by the C2f module, the rest of the structure is consistent with the FPN-PAN structure of YOLOv5.
[0142] S4, decoupling head detection head structure
[0143] From YOLOv3 to YOLOv5, the detection head has always been "coupled", that is, a layer of convolution is used to complete the two tasks of classification and positioning at the same time. It was not until the advent of YOLOX that the YOLO series was equipped with a "decoupled head" for the first time; YOLOv8 also adopts the structure of a decoupled head. Two parallel branches extract the category features and position features of the indicator light status image and the temperature image respectively, and then use a layer of 1×1 convolution to complete the classification and positioning tasks of the indicator light status image and the temperature image; the main purpose of the two parallel branches is to process different prediction targets (such as the position and category of the object) in the target detection task separately to improve the detection accuracy and efficiency; it is usually composed of several convolutional layers for feature extraction, and finally outputs the category probability through a fully connected layer or convolutional layer; the output tensor shape is (N, C), where N is the number of candidate boxes and C is the number of categories.
[0144] The overall network structure of YOLOv8 is shown in Figure 1, the YOLOv8 network structure is mainly composed of the following three parts: Backbone, Neck, and Head; Backbone is composed of "Conv", "C2f", and "SPPF"; it uses a series of convolution and deconvolution layers to extract features, and also uses residual connections and bottleneck structures to reduce the size of the network and improve performance; Neck is composed of "Conv", "C2f", "Concat", and "Upsample"; it uses multi-scale feature fusion technology to fuse feature maps from different stages of Backbone to enhance features Representation capability; Head is composed of "Detect"; it is responsible for the final target detection and classification tasks; "Conv" in the YOLOv8 network structure diagram is the convolution layer; it is the core component for processing image data; "C2f" is the name of a specific module, representing a specific convolution structure; the C2f module is composed of a convolution layer and a residual connection; the role of the C2f module in YOLOv8 is to effectively extract and fuse features by combining convolution layers and residual connections to improve the accuracy and efficiency of target detection; this design method helps the network better process complex image data, thereby achieving better performance in target detection tasks; " "SPPF" is a special variant of the spatial pyramid pooling module that specifically optimizes the performance and efficiency of the pooling operation; the core function of the SPPF module is to quickly extract multi-scale features, enhance the detection capability of targets of different sizes, and optimize computational efficiency; "Concat" means the operation of concatenating multiple feature maps according to the channel dimension; this is a common deep learning operation, especially widely used in multi-scale feature fusion; through the Concat operation, YOLOv8 can better fuse shallow and deep features, thereby achieving higher accuracy and efficiency in target detection tasks; "Upsample" is upsampling, which is used to increase the feature map spatial resolution (i.e. height and width); upsampling is a common operation in deep learning, especially in target detection tasks, it is used to restore low-resolution feature maps to higher resolutions in order to better detect small targets or recover more spatial details; through the Upsample operation, YOLOv8 can better handle multi-scale target detection tasks and improve the accuracy and robustness of the model; "Detect" usually refers to the detection head part of the network, which is responsible for generating the final target detection result from the input feature map; the detection head is an important part of the entire target detection model, which converts the feature information in the feature map into specific detection boxes and category probabilities.
[0145] S5. Label allocation strategy
[0146] Although YOLOv5 has designed some functions for automatically clustering candidate boxes, clustering candidate boxes depends on the data set; if the data set is not sufficient and cannot accurately reflect the distribution characteristics of the data itself, the clustered candidate boxes will also be too different from the real object size ratio; YOLOv8 does not adopt the candidate box strategy, so the problem to be solved is the multi-scale allocation of positive and negative sample matching; different from SimOTA used by YOLOX, YOLOv8 adopts the same TOOD strategy as YOLOv6 in the label allocation problem, which is a dynamic label allocation strategy; YOLOv8 only uses targetbboxes and targetscores, and does not include whether there is an object prediction, so the loss of YOLOv8 mainly includes two parts: category loss and regression loss; for YOLOv8, its classification loss is VFLLoss (Varifocal Loss), and its regression loss is in the form of CIoU Loss and DFL Loss; the classification loss Varifocal Loss content is as follows:
[0147]
[0148] In formula ①, p is the predicted category score, p∈[0,1]; q is the predicted target score. If it is the real category, q is the loU between the prediction and the true value; if it is other categories, q is 0;
[0149] The true category refers to the target detection task, for a specific candidate box (such as YOLO's detection cell), if the object category it detects matches a pre-defined category, then the candidate box is considered to contain an object of this specific category; in this case, the target score q for the candidate box will be defined as the loU value between the predicted box and the corresponding true box. The loU value is an indicator of the similarity between two bounding boxes, with a value between 0 and 1, where 1 means that the two bounding boxes completely overlap. Other categories refer to all other categories except a specific category; in the target detection task, if a candidate box does not detect a real object, or it detects an object that does not belong to the specified category, the corresponding target score q is usually set to 0; this approach is intended to make the predictor give extremely low scores to these "non-matching" situations, so that the model gradually corrects the prediction during the optimization process to better match the true category.
[0150] VFL Loss uses asymmetric parameters to weight positive and negative samples. By only attenuating negative samples, it achieves unequal processing of the contribution of foreground and background to the loss. For positive samples, q is used for weighting. If the GT of the positive sample IoUWhen p is very high, it contributes more to the loss, allowing the network to focus on high-quality samples, that is, training high-quality positive examples improves AP more than low-quality ones; for negative samples, use p γ The weighting is downgraded to reduce the contribution of negative examples to the loss, because the prediction p of negative samples will become smaller after taking the power, which can reduce the overall contribution of negative samples to the loss.
[0151] The communication terminal device adopts the announcement number CN212571823U, the patent name is "A distribution network communication terminal box", and other types of communication terminal boxes can also be used; the infrared imaging device adopts the announcement number CN219956722U, the patent name is "A infrared imaging device", and other types of infrared imaging devices can also be used.
[0152] The communication terminal box of the present invention adopts an intelligent computing unit (the intelligent computing unit is responsible for processing and analyzing data from various devices in the communication terminal device, such as temperature, humidity, equipment operating status and other data collected by the data acquisition assembly in the device, as well as image information collected by infrared and visible light binocular cameras, etc.) and an infrared imaging device and a visible light camera, which provide real-time equipment status monitoring and environmental monitoring for operation and maintenance personnel, such as equipment temperature, operating status and abnormal conditions of the surrounding environment, etc., to improve operation and maintenance efficiency and safety; by continuously collecting operation and maintenance data and performing intelligent analysis, it can further optimize system performance, adapt to changes in different scenarios and needs, and provide strong guarantees for the stable operation of the distribution automation system.
Claims
1. A communication terminal remote monitoring method based on the combination of visible light and infrared imaging, characterized in that: Using a visible light camera to collect indicator light status pictures on the communication terminal device, and using an infrared imaging device to collect temperature pictures of the communication terminal device when it is working; pre-processing the collected indicator light status pictures and temperature pictures to form a data set; Label the data set; put the labeled data set into the optimized YOLOv8 model for training.
2. A communication terminal remote monitoring method based on the combination of visible light and infrared imaging according to claim 1, characterized in that: The YOLOv8 model includes: S1, data preprocessing; S2, backbone network structure; S3, FPN-PAN structure; S4, decoupling head Detection head structure; S5. Label allocation strategy.
3. The communication terminal remote monitoring method based on the combination of visible light and infrared imaging according to claim 2 is characterized in that: In S1, the data preprocessing adopts the YOLOv5 strategy, and the model training adopts mosaic enhancement, mixed enhancement, spatial perturbation, and color perturbation.
4. The communication terminal remote monitoring method based on the combination of visible light and infrared imaging according to claim 2 is characterized in that: In S2, the backbone network structure is the MobileNetV3 network, including the attention mechanism, channel attention mechanism, and spatial attention mechanism, as follows: 1) The attention mechanism is used in the convolutional neural network (CNN) to automatically focus on the characteristic parts of the input image data while suppressing irrelevant parts; 2) The channel attention mechanism is used to focus on the correlation between different channels and enhance the expression of features by learning the importance of each channel. The steps are as follows: S201, feature map input Assume that the input feature map is represented as F in , whose shape is (C,H,W), where C is the number of channels, H and W are the height and width of the feature map respectively; S202, global average pooling Perform global average pooling on each channel to obtain a tensor with shape (C, 1, 1); The global information of each channel is compressed into a scalar, and the formula is as follows: In formula ①, GAP(F in ) is the result of the global average pooling operation; H and W represent the height and width of the input feature map, respectively. in Represents the input feature map; S203, fully connected layer or 1x1 convolution The result of global average pooling is passed through a fully connected layer or 1x1 convolution to learn the interdependence between channels. Two fully connected layers are used, one for compression and the other for excitation. S204, channel attention output The attention weight α is combined with the input feature map F in Multiply channel by channel to get the channel attention weighted feature map F out , the formula is as follows: F out =α⊙F in ② In formula ②, α represents the attention weight, which is a value assigned to each channel of the feature map according to its importance; the formula recombines the features of different channels according to their importance; 3) The spatial attention mechanism is used to focus on the importance of different spatial positions in the feature map and enhance the expression of features by learning the weight of each spatial position. The steps are as follows: S211, feature map input Assume that the feature map after channel attention is F out , whose shape is (C,H,W); S212, channel average pooling and maximum pooling The feature map is average pooled and max pooled in the channel dimension to obtain two tensors of shape (1, H, W). The formula is as follows: Formula ③ is average pooling, AvgPool(F out ) represents the feature map F out The result of the average pooling operation in the channel dimension; C represents the number of channels of the feature map; It means that the elements on all channels are summed and then averaged; Formula ④ is the maximum pooling, MaxPool(F out ) represents the feature map F out The result of the maximum pooling operation in the channel dimension; C represents the number of channels of the feature map; S213, splicing The results of average pooling and maximum pooling are concatenated in the channel dimension to obtain a tensor of shape (2, H, W). The formula is as follows: F concat =Concat(AvgPool(F out ),MaxPool(F out )) ⑤ In formula ⑤, F concat Represents the result of the concatenation operation, which represents a feature map with a shape of (2, Height, Width); Concat represents the concatenation operation, which is an operation that merges two or more tensors along a specified dimension; AvgPool(F out Represents the feature map F out The result of the average pooling operation; MaxPool(F out ) represents the feature map F out The result of the maximum pooling operation; The concatenation operation is used for feature fusion in the network, using different information extracted by different pooling methods to enhance feature representation capabilities; S214, 1x1 convolution The concatenated feature map is passed through a 1x1 convolutional layer to reduce the number of channels to one channel, and a spatial attention weight tensor β with a shape of (1, H, W) is obtained; S215, spatial attention output The spatial attention weight β is combined with the feature map F after channel attention. out Multiply element by element to get the final weighted feature map F final , the formula is as follows: F final =β⊙F out ⑥ In formula ⑥, β represents the spatial attention weight, which is used to redistribute weights to different spatial positions of the feature map; F out Represents the feature map after channel attention calculation; Through the spatial attention weight β and the feature map F out This element-by-element multiplication operation is performed to strengthen or suppress features at different positions in the feature map for transfer to subsequent processing steps.
5. The communication terminal remote monitoring method based on the combination of visible light and infrared imaging according to claim 2 is characterized in that: In S3, the FPN-PAN structure is used to construct the feature pyramid of YOLO, so that multi-scale information can be integrated.
6. The communication terminal remote monitoring method based on the combination of visible light and infrared imaging according to claim 2 is characterized in that: In S4, the decoupling head Detection head structure is: two parallel branches extract the category features and position features of the indicator light status image and the temperature image respectively, and then use a layer of 1×1 convolution to complete the classification and positioning tasks of the indicator light status image and the temperature image.
7. A communication terminal remote monitoring method based on the combination of visible light and infrared imaging according to claim 6, characterized in that: Two parallel branches are used to separately process different prediction targets in the object detection task.
8. The communication terminal remote monitoring method based on the combination of visible light and infrared imaging according to claim 2 is characterized in that: In S5, the label assignment strategy includes classification loss VFLLoss (Varifocal Loss), which is as follows: The classification loss VFLLoss (Varifocal Loss) is defined as follows: In formula ①, p is the predicted category score, p∈[0,1]; q is the predicted target score. If it is the real category, q is the loU between the prediction and the true value; if it is other categories, q is 0; The classification loss VFL Loss uses an asymmetric parameter to weight the positive and negative samples. By only attenuating the negative samples, the unequal contribution of the foreground and background to the loss is completed. For the positive samples, q is used for weighting; for the negative samples, p is used. γ Demotion.
9. The communication terminal remote monitoring method based on the combination of visible light and infrared imaging according to claim 2 is characterized in that: The pictures of the indicator light status on the communication terminal device include pictures of the indicator light in normal working status, pictures of the indicator light in abnormal working status, which are displayed in red, and pictures of the indicator light in stopped working status; the indicator light in normal working status is green, the indicator light in abnormal working status is red, and the indicator light in stopped working status is not lit.
Citation Information
Patent Citations
Distribution network communication terminal box
CN212571823U
Infrared imaging device
CN219956722U