Method and device for detecting appearance of acquisition terminal based on binary feature aggregation network
Through the appearance detection method of the acquisition terminal based on the dual feature aggregation network, the problems of low detection efficiency and poor accuracy in the prior art acquisition terminal are solved, and higher detection accuracy and stability are achieved, and the error detection rate and labor cost are reduced.
Patent Information
- Application Number
- CN202510290821.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
AI Technical Summary
The existing acquisition terminal appearance detection methods cannot accurately and effectively identify and detect the acquisition terminal, and there are problems of inefficiency, accuracy and consistency.
The appearance detection method of the acquisition terminal based on the bivariate feature aggregation network is adopted to extract the low-level features of the standardized image and perform feature aggregation through at least one bivariate feature aggregation network to integrate local details and global structure information.
It significantly improves the accuracy and stability of the appearance detection of the acquisition terminal, enhances the model's understanding of complex scenarios, reduces the rate of false detection and missed detection, simplifies the operation process and reduces labor costs.
Smart Images

Figure CN120220059A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for detecting the appearance of a collection terminal in the technical field of power system detection, in particular to a method for detecting the appearance of a collection terminal based on a dual-feature aggregation network, and also relates to a device for detecting the appearance of a collection terminal based on a dual-feature aggregation network. Background Art
[0002] With the accelerating advancement of the construction of smart grids and the continuous improvement of power management informatization, as a key device for power metering and data collection, the appearance quality and functional integrity of the collection terminal are crucial for the stable operation of the power system and the electricity consumption experience of users. The collection terminal is not only the metering basis for the transaction between the power supply and demand sides, but also an important node for realizing remote monitoring, data analysis, and fault warning in the smart grid. Therefore, ensuring that the appearance of the collection terminal is undamaged, the display screen is clearly readable, and the internal functional modules operate normally is the basis for ensuring stable power supply, improving power service quality, and achieving the goal of energy conservation and emission reduction. In the framework of the smart grid, the collection terminal not only undertakes the traditional power metering task, but also incorporates more intelligent functions, such as two-way communication, real-time data upload, remote control, and diagnosis. The realization of these functions depends on the complex electronic components and precise manufacturing processes inside the collection terminal, and also poses higher requirements for the appearance design and material selection of the collection terminal. Once the appearance of the collection terminal is damaged or the function is abnormal, it may not only affect the accuracy of metering, but also interfere with the normal operation of the smart grid, and even pose potential safety hazards.
[0003] Currently, the following methods mainly exist in the existing field of collection terminal appearance detection. The first is the manual off-line observation method. This traditional method requires staff to take out the collection terminal one by one and check the appearance details such as the data displayed on the liquid crystal screen and the shell with the naked eye. This method not only has low efficiency, consumes a large amount of human and material resources, but also the detection results are easily affected by personal subjective judgment, and there are problems in terms of accuracy and consistency. The second is the automated line verification mode. However, in actual operation, this mode encounters problems of image gray value fluctuations caused by external environmental factors (such as light changes, air dust) and liquid crystal screen process differences. These factors directly affect the detection accuracy, resulting in a high misjudgment rate of the collection terminal, laying hidden dangers for subsequent use and maintenance. The third is the collection terminal verification method based on feature recognition and automatic correction technology. This method finely preprocesses the image of the collection terminal and uses a high-precision region positioning algorithm and differential operation technology to accurately identify the appearance defects of the collection terminal. Nevertheless, the popularization and application of this method still face certain limitations. Its applicability may be limited by the specific type and specifications of the collection terminal. At the same time, for some specific or complex defect types, this method may not be able to achieve completely effective identification and correction. Summary of the Invention
[0004] To solve the technical problem that the existing appearance detection methods for acquisition terminals cannot accurately and effectively identify and detect acquisition terminals, the present invention provides an appearance detection method and device for acquisition terminals based on a dual-feature aggregation network.
[0005] The present invention is implemented by the following technical solutions: An appearance detection method for an acquisition terminal based on a dual-feature aggregation network, which includes the following steps:
[0006] Obtain a standardized image of the appearance of the acquisition terminal;
[0007] Extract low-level features of the standardized image to obtain an input feature map;
[0008] Perform feature aggregation through at least one dual-feature aggregation network; the first dual-feature aggregation network aggregates the input feature map, and other dual-feature aggregation networks aggregate the fused feature maps output by the previous dual-feature aggregation network; each dual-feature aggregation network performs: first adjust the number of channels of the input feature map or the fused feature map and evenly divide it into two parts, one part is used for global feature extraction through a stack of multiple neck network modules one, and the other part is used for local feature extraction through a stack of multiple neck network modules two, and then perform feature fusion on the extracted global features and local features and integrate and output them as the fused feature map;
[0009] Separate the regression branch and the classification branch of the fused feature map;
[0010] First calculate the regression loss and classification loss of a loss function, and then optimize the training model used to represent and detect the appearance of the acquisition terminal through the loss function.
[0011] By introducing a dual-feature aggregation network, the present invention effectively fuses the local detailed information and global structural information in the appearance features of the acquisition terminal. This dual-feature aggregation mechanism enhances the model's ability to capture the appearance features of the acquisition terminal through two parallel feature extraction channels, resulting in a significant improvement in detection accuracy compared to traditional neural network models. At the same time, this method shows stronger robustness in the face of complex scenarios such as the diversity, occlusion, or illumination changes in the appearance of the acquisition terminal, not only improving the accuracy and stability of the acquisition terminal appearance detection, but also enhancing the model's understanding ability of complex scenarios, providing more comprehensive support for the acquisition terminal appearance detection, thereby improving the practicality and application scope of the acquisition terminal appearance detection technology, and solving the technical problem that the existing acquisition terminal appearance detection methods cannot accurately and effectively identify and detect the acquisition terminal.
[0012] As a further improvement of the above solution, the method for obtaining the standardized image includes the following steps:
[0013] Collect the appearance image data of the acquisition terminal and perform annotation; the appearance image data of the terminal includes the appearance image of the normal acquisition terminal and the appearance image of the abnormal acquisition terminal;
[0014] Preprocess the appearance image data of the terminal to obtain the standardized image; the preprocessing method of the appearance image data of the terminal includes: adjusting the size of the appearance image of the terminal; normalizing the pixel values of the appearance image of the terminal and performing standardization processing; performing data augmentation on the appearance image data of the terminal.
[0015] As a further improvement of the above solution, extract the low-level features of the standardized image through at least one convolutional module one; the convolutional kernel of the convolutional module one is 3x3, and the stride is set to 2; the convolutional module one also introduces non-linearity through an activation function and performs batch normalization.
[0016] Furthermore, the activation function is expressed as:
[0017] silu(x) = x * σ(x)
[0018] silu(x)' = silu(x) + σ(x) * (1 - silu(x))
[0019] where σ(x) is the logistic function, x is the input value, and silu(x) is the activation function.
[0020] As a further improvement of the above solution, each dual feature aggregation network performs the following steps: first, adjust the number of channels of the input feature map or the fused feature map through convolutional module two and evenly divide it into two parts, one part of which performs global feature extraction through multiple stacked neck network modules one, and the other part performs local feature extraction through multiple stacked neck network modules two, then perform feature fusion on the extracted global features and local features, and integrate and output through convolutional module three; each neck network module one or neck network module two includes two 3x3 convolutional layers and adopts residual connection.
[0021] As a further improvement of the above solution, the separation method of the regression branch and the classification branch includes the following steps: perform convolution and upsampling on the fused feature map to generate detection layers of multiple scales; each detection layer includes a regression branch for predicting the target bounding box and a classification branch for predicting the target category.
[0022] As a further improvement of the above solution, process the fused feature map through a head network to separate the regression branch and the classification branch; the head network adopts a double decoupled head structure and generates detection results including the target bounding box, the target category, and the target confidence.
[0023] As a further improvement of the above solution, the calculation formula of the classification loss of the loss function is:
[0024]
[0025] where L is the classification loss, N is the number of samples, y i is the true label of the i-th sample, and p i is the predicted probability of the i-th sample.
[0026] As a further improvement of the above solution, the calculation formula of the classification loss of the loss function is:
[0027] DFL(S i , S i+1 ) = -((y i+1 - y) log(S i ) + (y - y i ) log(S i+1 ))
[0028] where S i is the predicted value output by the network, S i+1 is the adjacent predicted value output by the network, y is the adjacent predicted value output by the network, y i is the integral value of the label, and y i+1 is the adjacent integral value of the label.
[0029] The present invention also provides an appearance detection device for a collection terminal based on a dual-feature aggregation network, which includes:
[0030] A preprocessing unit for obtaining a standardized image of the appearance of the collection terminal;
[0031] An input unit for extracting low-level features of the standardized image to obtain an input feature map;
[0032] A dual-feature aggregation unit for performing feature aggregation through at least one dual-feature aggregation network; the first dual-feature aggregation network performs feature aggregation on the input feature map, and other dual-feature aggregation networks perform feature aggregation on the fused feature map output by the previous dual-feature aggregation network; each dual-feature aggregation network performs: first, adjusting the number of channels of the input feature map or the fused feature map and evenly dividing it into two parts, one part of which performs global feature extraction through a stacked plurality of neck network modules one, and the other part of which performs local feature extraction through a stacked plurality of neck network modules two, and then performing feature fusion on the extracted global features and local features and integrating and outputting them as the fused feature map;
[0033] A separation unit for separating the regression branch and the classification branch of the fused feature map;
[0034] A loss calculation unit, which is configured to first calculate the regression loss and classification loss of a loss function, and then optimize a training model for representing and detecting the appearance of the acquisition terminal through the loss function.
[0035] Compared with the existing acquisition terminal appearance detection methods and devices, the acquisition terminal appearance detection method and device based on the dual feature aggregation network of the present invention have the following beneficial effects:
[0036] 1. The acquisition terminal appearance detection method based on the dual feature aggregation network effectively fuses the local detail information and global structure information in the acquisition terminal appearance features by introducing the dual feature aggregation network. This dual feature aggregation mechanism enhances the model's ability to capture the appearance features of the acquisition terminal through two parallel feature extraction channels, resulting in a significant improvement in detection accuracy compared to traditional neural network models. At the same time, when facing complex scenarios such as the diversity, occlusion, or illumination changes of the acquisition terminal appearance, this method shows stronger robustness, not only improving the accuracy and stability of the acquisition terminal appearance detection, but also enhancing the model's understanding ability of complex scenarios, providing more comprehensive support for the acquisition terminal appearance detection, thereby improving the practicality and application scope of the acquisition terminal appearance detection technology, and solving the technical problem that the existing acquisition terminal appearance detection methods cannot accurately and effectively identify and detect the acquisition terminal.
[0037] 2. The acquisition terminal appearance detection method based on the dual feature aggregation network simplifies the operation process and reduces the labor cost. On the basis of maintaining high accuracy, this method optimizes the neural network, effectively improving the detection speed by streamlining the network structure and reducing redundant calculations. This enables the acquisition terminal appearance detection to be carried out in real-time or near real-time conditions, meeting the requirements for efficient and rapid detection in industrial fields. At the same time, this method enhances the generalization ability of the model. The design of the dual feature aggregation network not only improves the sensitivity of the model to the appearance features of the acquisition terminal, but also enhances the generalization ability of the model. This means that even when facing unseen acquisition terminal models or appearance changes, the model can still maintain a high detection accuracy, thereby reducing the limitations of the model in practical applications.
[0038] 3. The acquisition terminal appearance detection method based on the dual - feature aggregation network reduces the false detection and missed detection rates. Thanks to the accurate capture of the acquisition terminal appearance features by the dual - feature aggregation network and the powerful performance of the neural network in the field of object detection, this method has been improved in feature extraction and detection strategies, effectively reducing the false detection and missed detection rates in the acquisition terminal appearance detection. At the same time, this method realizes automated detection, greatly simplifies the operation process, simplifies the operation process and reduces labor costs, providing strong support for the intelligent and automated transformation of the power industry. In addition, this method improves the maintenance and management level of the acquisition terminal, can timely detect abnormal conditions on the appearance of the acquisition terminal, such as damage, pollution, etc., thus providing an important basis for the timely maintenance and replacement of the acquisition terminal. This helps to extend the service life and safety of the acquisition terminal, and further improves the stability and reliability of the entire power system.
[0039] 4. The acquisition terminal appearance detection method based on the dual - feature aggregation network has brought significant beneficial effects in terms of detection accuracy, robustness, detection speed, model generalization ability, false detection and missed detection rates, and simplification of the operation process. These effects not only improve the performance of the acquisition terminal appearance detection, but also provide strong support for the intelligent and automated development of the power industry. Description of the Drawings
[0040] Figure 1 It is the flowchart of the acquisition terminal appearance detection method based on the dual - feature aggregation network in Embodiment 1 of the present invention;
[0041] Figure 2 It is the framework diagram of the backbone network in the acquisition terminal appearance detection method based on the dual - feature aggregation network in Embodiment 1;
[0042] Figure 3 It is the aggregation flowchart of the dual - feature aggregation network in the acquisition terminal appearance detection method based on the dual - feature aggregation network in Embodiment 1;
[0043] Figure 4 It is the schematic diagram of the head network in the acquisition terminal appearance detection method based on the dual - feature aggregation network in Embodiment 1. Detailed Embodiment
[0044] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described here are only used to explain the present invention, and are not used to limit the present invention.
[0045] Embodiment 1
[0046] Please refer to Figures 1 to 4, this embodiment provides a method for detecting the appearance of a collection terminal based on a dual - feature aggregation network. This detection method uses a traditional neural network as the baseline model, and on this basis, a dual - feature aggregation network is designed for detection. Among them, the method for detecting the appearance of the collection terminal mainly includes the following steps.
[0047] (1) Obtain a standardized image of the collection terminal appearance. In this embodiment, the method for obtaining the standardized image includes the following steps: collect the appearance image data of the collection terminal and perform annotation; the appearance image data of the terminal includes the appearance image of a normal collection terminal and the appearance image of an abnormal collection terminal; pre - process the appearance image data of the terminal to obtain a standardized image. Among them, the pre - processing method of the appearance image data of the terminal includes: adjust the size of the appearance image of the terminal; normalize the pixel values of the appearance image of the terminal and perform standardization processing; perform data augmentation on the appearance image data of the terminal.
[0048] This step is mainly for pre - processing the input of the appearance image of the collection terminal. The input image will first go through pre - processing steps to meet the input requirements of the model. Usually, the input image will be adjusted to a fixed size (640x640) to ensure input consistency. Then, the pixel values of the image will be normalized from [0, 255] to [0, 1] and standardized (subtract the mean and divide by the standard deviation) to accelerate model convergence and improve training stability. In the training stage, various data augmentation techniques (such as Mosaic, MixUp, random cropping, flipping, etc.) are applied to increase data diversity and improve the generalization ability of the model.
[0049] (2) Extract low - level features of the standardized image to obtain an input feature map. In this embodiment, at least one convolutional module one (convolutional module 1) is used to extract low - level features of the standardized image, that is, edges and textures. The convolutional kernel of convolutional module one is 3x3, and the stride is set to 2 to achieve down - sampling of the feature map, thereby reducing the spatial size and computational complexity. Convolutional module one also introduces non - linearity through an activation function to enhance the expression ability of the model. Convolutional module one also performs batch normalization to accelerate the training process and improve the stability of the model. Among them, the activation function is the SiLU activation function, expressed as:
[0050] silu(x) = x * σ(x)
[0051] silu(x) ′ = silu(x)+σ(x)*(1 - silu(x))
[0052] Among them, σ(x) is the logistic function, x is the input value, and silu(x) is the activation function.
[0053] (3) Feature aggregation is performed through at least one dual - feature aggregation network. The first dual - feature aggregation network aggregates the input feature map, and other dual - feature aggregation networks aggregate the fused feature maps output by the previous dual - feature aggregation network. Each dual - feature aggregation network performs the following: First, the number of channels of the input feature map or the fused feature map is adjusted and evenly divided into two parts. One part is used for global feature extraction through a plurality of stacked neck network modules one, and the other part is used for local feature extraction through a plurality of stacked neck network modules two. Then, the extracted global features and local features are fused and integrated to output a fused feature map.
[0054] In some embodiments, each dual - feature aggregation network performs the following steps: First, the number of channels of the input feature map or the fused feature map is adjusted and evenly divided into two parts through a second convolutional module (corresponding to Figure 2 convolutional module 2, convolutional module 3, convolutional module 4, convolutional module 5, etc. in it). One part is used for global feature extraction through a plurality of stacked neck network modules one, and the other part is used for local feature extraction through a plurality of stacked neck network modules two. Then, the extracted global features and local features are fused through a feature fusion module, and the output is integrated through a third convolutional module; each neck network module one or neck network module two includes two 3x3 convolutional layers and uses residual connections. Among them, each neck network module is composed of two 3x3 convolutional layers, and residual connections are used to alleviate gradient vanishing. This design enhances the gradient propagation efficiency by forcing feature reuse. The splitting operation reduces the number of processing channels of the subsequent neck network modules. The dual - feature aggregation network aggregates features of different depths, retaining both shallow - layer detailed information and integrating deep - layer semantic features. Especially when dealing with multi - scale targets, its efficient feature fusion ability significantly improves the detection sensitivity of the model to small targets and the localization accuracy of large targets, while maintaining a relatively low computational complexity, providing an efficient solution for the real - time detection task of the appearance of the acquisition terminal. Finally, the spatial pyramid pooling module solves the problem of inconsistent input image sizes through multi - scale pooling operations and enhances the multi - scale feature extraction ability of the model.
[0055] Since there are significant differences in the morphology, size, etc. of the appearance defect images of the acquisition terminal, these phenomena require the acquisition terminal appearance detection model to have strong adaptability and robustness to accurately identify and classify abnormal situations in different environments. This dual - feature aggregation network is one of the core components of the acquisition terminal appearance detection model, integrating the idea of the cross - stage partial network and introducing a more flexible multi - branch structure. The goal is to use two branches operating in parallel, each focusing on the accurate extraction of global features and local features, thereby strengthening the model's ability to analyze and fuse various feature channel information, such as Figure 3As shown in the figure. The above steps also design a feature fusion and optimization mechanism, which can automatically select and fuse the most valuable features while suppressing irrelevant or redundant features. This feature fusion and optimization mechanism further improves the detection performance of the model, enabling it to maintain stable and accurate detection results even in the face of complex scenarios. In summary, this step addresses the problems of diverse appearance defects, complex backgrounds, and varying target sizes in the acquisition terminal. By using a dual-feature aggregation network to compensate for the limitations of the original model's receptive field and fixed convolutional feature capture mode, the average detection accuracy and performance of the acquisition terminal's appearance detection can be effectively improved.
[0056] (4) Separate the regression branch and the classification branch of the fused feature map. In this embodiment, the method for separating the regression branch and the classification branch includes the following steps: performing convolution and upsampling on the fused feature map to generate detection layers of multiple scales; each detection layer includes a regression branch for predicting the target bounding box and a classification branch for predicting the target category.
[0057] As Figure 4 shown, in some other embodiments, a head network processes the fused feature map to separate the regression branch and the classification branch; the head network adopts a double decoupled head structure and generates detection results including the target bounding box, the target category, and the target confidence. The head network is a key component in the acquisition terminal appearance detection model. It adopts an advanced double decoupled head structure to achieve the separation of the regression branch and the classification branch, thereby improving the convergence speed and detection effect of the model. This design enables the head network to more effectively process the fused feature map from the neck network and generate accurate detection results. The network simplifies the training process and reduces the hyperparameter settings. By directly predicting the center point and other related attributes of the target, the head network can generate the final detection results including the bounding box, the category, and the confidence. This change not only improves the flexibility of the model but also significantly enhances the accuracy and efficiency of target detection. During the inference process, the head network receives the fused feature map and, through a series of convolution and upsampling operations, generates detection layers of multiple scales. Each detection layer contains a regression branch and a classification branch, which are respectively used to predict the bounding box and the category of the target.
[0058] (5) First calculate the regression loss and the classification loss of a loss function, and then optimize the training model used to represent and detect the appearance of the acquisition terminal through the loss function. In this embodiment, the calculation formula for the classification loss of the loss function is:
[0059]
[0060] In the above formula, L is the classification loss, N is the number of samples, y i is the true label of the i-th sample, and p i is the predicted probability of the i-th sample.
[0061] The calculation formula for the classification loss of the loss function is as follows:
[0062] DFL(S i ,S i+1 ) = -((y i+1 -y)log(S i )+(y - y i )log(S i+1 ))
[0063] In the above formula, S i is the predicted value output by the network, S i+1 is the adjacent predicted value output by the network, y is the adjacent predicted value output by the network, y i is the integral value of the label, and y i+1 is the adjacent integral value of the label.
[0064] The above loss function is mainly composed of classification loss and regression loss. First is the classification loss, which is mainly used to measure the difference between the class probability distribution predicted by the model and the actual class label. The regression loss is mainly used to measure the difference between the predicted bounding box of the model and the actual bounding box, combining multiple factors such as the distance between the center points of the predicted box and the label box, the length of the diagonal of the circumscribed rectangle, and the aspect ratio, to more comprehensively measure the difference between the predicted bounding box and the actual bounding box. This embodiment adopts advanced model training and optimization strategies, including the selection of the loss function, the adjustment of the learning rate, data augmentation techniques, etc. These strategies help the model converge faster during the training process and at the same time improve the generalization ability of the model, enabling it to better adapt to different acquisition terminal appearance detection tasks.
[0065] Compared with the existing acquisition terminal appearance detection methods, the acquisition terminal appearance detection method based on the dual - feature aggregation network in this embodiment has the following beneficial effects:
[0066] 1. The appearance detection method of the acquisition terminal based on the dual - feature aggregation network effectively integrates the local detail information and the global structure information in the appearance features of the acquisition terminal by introducing the dual - feature aggregation network. This dual - feature aggregation mechanism enhances the model's ability to capture the appearance features of the acquisition terminal through two parallel feature extraction channels, resulting in a significant improvement in detection accuracy compared to traditional neural network models. At the same time, when facing complex scenarios such as the diversity, occlusion, or illumination changes in the appearance of the acquisition terminal, this method shows stronger robustness, not only improving the accuracy and stability of the acquisition terminal appearance detection, but also enhancing the model's understanding ability of complex scenarios, providing more comprehensive support for the acquisition terminal appearance detection, thus improving the practicality and application scope of the acquisition terminal appearance detection technology, and solving the technical problem that the existing acquisition terminal appearance detection methods cannot accurately and effectively identify and detect the acquisition terminal.
[0067] 2. The appearance detection method of the acquisition terminal based on the dual - feature aggregation network simplifies the operation process, reduces labor costs, and optimizes the detection speed. On the basis of maintaining high accuracy, this method optimizes the neural network, effectively improving the detection speed by streamlining the network structure and reducing redundant calculations. This enables the acquisition terminal appearance detection to be carried out in real - time or near - real - time conditions, meeting the requirements of the industrial field for efficient and rapid detection.
[0068] 3. The appearance detection method of the acquisition terminal based on the dual - feature aggregation network enhances the generalization ability of the model. The design of the dual - feature aggregation network not only improves the sensitivity of the model to the appearance features of the acquisition terminal, but also enhances the generalization ability of the model. This means that even when facing unseen acquisition terminal models or appearance changes, the model can still maintain a high detection accuracy, thus reducing the limitations of the model in practical applications.
[0069] 4. The appearance detection method of the acquisition terminal based on the dual - feature aggregation network reduces the false detection and missed detection rates. Thanks to the accurate capture of the appearance features of the acquisition terminal by the dual - feature aggregation network and the powerful performance of the neural network in the field of object detection, this method improves the feature extraction and detection strategies, effectively reducing the false detection and missed detection rates in the acquisition terminal appearance detection.
[0070] 5. The appearance detection method of the acquisition terminal based on the dual - feature aggregation network realizes automated detection, greatly simplifies the operation process, simplifies the operation process and reduces labor costs, providing strong support for the intelligent and automated transformation of the power industry.
[0071] 6. The appearance detection method of the acquisition terminal based on the dual - feature aggregation network improves the maintenance and management level of the acquisition terminal, and can timely detect abnormal conditions of the acquisition terminal appearance, such as damage, pollution, etc., thus providing an important basis for the timely maintenance and replacement of the acquisition terminal. This helps to improve the service life and safety of the acquisition terminal, and further enhances the stability and reliability of the entire power system.
[0072] In summary, the appearance detection method of the acquisition terminal based on the dual - feature aggregation network has brought significant beneficial effects in terms of detection accuracy, robustness, detection speed, model generalization ability, false detection and missed detection rates, and simplification of the operation process. These effects not only improve the performance of the acquisition terminal appearance detection, but also provide strong support for the intelligent and automated development of the power industry.
[0073] Example 2
[0074] This example provides an appearance detection method of the acquisition terminal based on the dual - feature aggregation network, which conducts model training on the basis of Example 1, as follows.
[0075] 1. Data preparation: Collect a large amount of acquisition terminal appearance image data, including images of normal acquisition terminals and abnormal acquisition terminals (such as damaged, polluted, etc.). Label the image data, including information such as the position, size of the acquisition terminal, and abnormal types.
[0076] 2. Model construction: On the basis of the traditional neural network, introduce the dual - feature aggregation network module. The dual - feature aggregation network module includes two parallel backbone networks, one for extracting local detail features of the acquisition terminal appearance, and the other for extracting global structural features. Fuse the features of the two backbone networks to form a more comprehensive and rich feature representation.
[0077] 3. Model training: Use the prepared acquisition terminal appearance image data to train the model. During the training process, adopt an appropriate loss function to optimize the model parameters. Through multiple iterative trainings, make the model gradually learn the feature representation and detection strategy of the acquisition terminal appearance.
[0078] 4. Model evaluation and optimization: Use the test data set to evaluate the trained model, including indicators such as detection accuracy and robustness. Optimize the model according to the evaluation results, such as adjusting the network structure, loss function parameters, etc.
[0079] Example 3
[0080] This example provides an appearance detection method of the acquisition terminal based on the dual - feature aggregation network, which is designed and implemented on the basis of Example 1 or Example 2, as follows.
[0081] 1. Design and implement modules or units for each step, such as an image acquisition module, a preprocessing module, a detection module, and a postprocessing module. The image acquisition module is responsible for obtaining the appearance image of the acquisition terminal, which can be achieved through methods such as a camera or an image database. The preprocessing module preprocesses the acquired image, such as denoising and enhancing the contrast. The detection module uses the trained dual feature aggregation network acquisition terminal appearance detection model to detect the preprocessed image. The postprocessing module processes the detection results, such as removing duplicate detection frames and merging similar detection frames.
[0082] 2. System implementation: Use a programming language and a deep learning framework to implement the functions of each module of the system. During the implementation process, pay attention to optimizing the code performance and improving the detection speed. Provide a friendly user interface for the system to facilitate users to operate and view the detection results.
[0083] 3. System testing and deployment: Conduct a comprehensive test on the system, including functional testing, performance testing, and compatibility testing, etc. Optimize and improve the system according to the test results. Deploy the system to the actual scenario for actual application and verification.
[0084] Example 4
[0085] This example provides a method for detecting the appearance of an acquisition terminal based on a dual feature aggregation network, which integrates the acquisition terminal appearance detection and the intelligent maintenance system on the basis of any one of Example 1, Example 2, and Example 3, as follows.
[0086] 1. System integration: Integrate the acquisition terminal appearance detection system with the intelligent maintenance system of the power company. Through interfaces or data exchange methods, transmit the detection results of the detection system to the intelligent maintenance system in real time.
[0087] 2. Intelligent maintenance strategy formulation: According to the detection results of the detection system, the intelligent maintenance system automatically formulates the maintenance plan for the acquisition terminal. For the detected abnormal acquisition terminals, the system can automatically trigger maintenance tasks and dispatch maintenance personnel to handle them.
[0088] 3. System monitoring and management: The intelligent maintenance system monitors and manages the running status of the acquisition terminal appearance detection system. System administrators can view the real-time data, historical records, and the execution status of maintenance tasks of the detection system through the interface.
[0089] 4. Data analysis and visualization: Conduct in-depth analysis on the data of the detection system to mine the laws and trends of abnormal acquisition terminal appearances. Display the analysis results through visualization tools (such as charts, reports, etc.) to provide support for the decision-making of the power company.
[0090] Example 4
[0091] This embodiment provides an acquisition terminal appearance detection device based on a dual - feature aggregation network. The device includes a pre - processing unit, an input unit, a dual - feature aggregation unit, a separation unit, and a loss calculation unit, and may also include some other units set according to actual requirements.
[0092] The pre - processing unit is used to obtain a standardized image of the acquisition terminal appearance. The pre - processing unit is mainly used to implement the pre - processing steps in Embodiment 1. Through pre - processing, the input image can be adjusted to a fixed size, the pixel values of the image can be standardized, and at the same time, during the training stage, a variety of data augmentation techniques (such as Mosaic, MixUp, random cropping, flipping, etc.) are applied to increase the diversity of data and improve the generalization ability of the model.
[0093] The input unit is used to extract low - level features of the standardized image to obtain an input feature map. The input unit is mainly used to implement the step of obtaining the input feature map in Embodiment 1, which is achieved through Convolution Module 1 here.
[0094] The dual - feature aggregation unit is used to perform feature aggregation through at least one dual - feature aggregation network. The first dual - feature aggregation network aggregates the input feature map, and other dual - feature aggregation networks aggregate the fused feature maps output by the previous dual - feature aggregation network. Each dual - feature aggregation network performs the following: first, adjust the number of channels of the input feature map or fused feature map and evenly divide it into two parts. One part performs global feature extraction through a stack of multiple Neck Network Modules 1, and the other part performs local feature extraction through a stack of multiple Neck Network Modules 2. Then, fuse the extracted global features and local features and integrate them to output a fused feature map. Similarly, the dual - feature aggregation unit is also used to implement step (3) in Embodiment 1, and will not be elaborated here.
[0095] The separation unit is used to separate the regression branch and the classification branch of the fused feature map. The separation unit is mainly used to implement step (4) in Embodiment 1. Through processing by the Head Network, it can more effectively process the fused feature map from the Neck Network and generate accurate detection results.
[0096] The loss calculation unit is used to first calculate the regression loss and classification loss of a loss function, and then optimize the training model used to represent and detect the acquisition terminal appearance through the loss function. The loss calculation unit is mainly used to implement step (5) in Embodiment 1. Through relevant strategies, the model can converge faster during the training process, and at the same time improve the generalization ability of the model, enabling it to better adapt to different acquisition terminal appearance detection tasks.
[0097] Embodiment 5
[0098] This embodiment provides a computer terminal, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the acquisition terminal appearance detection method based on the dual-feature aggregation network in any one of Embodiment 1, Embodiment 2, Embodiment 3, and Embodiment 4.
[0099] When the method in any one of Embodiment 1, Embodiment 2, Embodiment 3, and Embodiment 4 is applied, it can be applied in the form of software. For example, it can be designed as an independently running program and installed on a computer terminal. The computer terminal can be a computer, a smart phone, a control system, and other Internet of Things devices, etc. The method in any one of Embodiment 1, Embodiment 2, Embodiment 3, and Embodiment 4 can also be designed as an embedded running program and installed on a computer terminal, such as installed on a single-chip microcomputer.
[0100] Embodiment 5
[0101] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of the acquisition terminal appearance detection method based on the dual-feature aggregation network in any one of Embodiment 1, Embodiment 2, Embodiment 3, and Embodiment 4.
[0102] When the method in any one of Embodiment 1, Embodiment 2, Embodiment 3, and Embodiment 4 is applied, it can be applied in the form of software. For example, it can be designed as an independently running program on a computer-readable storage medium. The computer-readable storage medium can be a USB flash drive, designed as a USB key, and designed as a program that starts the whole method through external triggering through the USB flash drive.
[0103] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for detecting the appearance of a collection terminal based on a binary feature aggregation network, characterized in that: It includes the following steps: Obtaining a standardized image of the appearance of the acquisition terminal; Extracting low-level features of the standardized image to obtain an input feature map; Feature aggregation is performed through at least one binary feature aggregation network; the first binary feature aggregation network feature aggregates the input feature map, and other binary feature aggregation networks feature aggregate the fused feature map output by the previous binary feature aggregation network; each binary feature aggregation network performs: firstly adjusting the number of channels of the input feature map or the fused feature map and evenly dividing it into two parts, one part of which is subjected to global feature extraction through a plurality of stacked neck network modules 1, and the other part of which is subjected to local feature extraction through a plurality of stacked neck network modules 2, and then the extracted global features and local features are subjected to feature fusion and integrated output as the fused feature map; Separating the regression branch and the classification branch of the fused feature map; First, a regression loss and a classification loss of a loss function are calculated, and then a training model for representing and detecting the appearance of the acquisition terminal is optimized by the loss function.
2. The method for detecting appearance of a collection terminal based on a binary feature aggregation network according to claim 1, characterized in that: The method for obtaining the standardized image comprises the following steps: Collecting terminal appearance image data and marking it; the terminal appearance image data includes normal collection terminal appearance images and abnormal collection terminal appearance images; The terminal appearance image data is preprocessed to obtain the standardized image; the preprocessing method of the terminal appearance image data includes: adjusting the size of the terminal appearance image; normalizing the pixel value of the terminal appearance image and performing standardization processing; and performing data enhancement on the terminal appearance image data.
3. The method for detecting appearance of a collection terminal based on a binary feature aggregation network according to claim 1, characterized in that: The low-level features of the standardized image are extracted through at least one convolution module 1; the convolution kernel of the convolution module 1 is 3x3, and the stride is set to 2; the convolution module 1 also introduces nonlinearity through an activation function and performs batch normalization.
4. The method for detecting appearance of a collection terminal based on a binary feature aggregation network according to claim 3, characterized in that: The activation function is expressed as: silu(x)=x*σ(x) silu(x)′=silu(x)+σ(x)*(1-silu(x)) Among them, σ(x) is the logic function, x is the input value, and silu(x) is the activation function.
5. The method for detecting appearance of a collection terminal based on a binary feature aggregation network according to claim 1, characterized in that: Each binary feature aggregation network performs the following steps: first, the number of channels of the input feature map or the fused feature map is adjusted through the convolution module 2 and evenly divided into two parts, one of which is subjected to global feature extraction through multiple stacked neck network modules 1, and the other part is subjected to local feature extraction through multiple stacked neck network modules 2, and then the extracted global features and local features are feature fused and integrated and outputted through the convolution module 3; each neck network module 1 or neck network module 2 includes two 3x3 convolution layers and adopts residual connection.
6. The method for detecting appearance of a collection terminal based on a binary feature aggregation network according to claim 1, characterized in that: The method for separating the regression branch and the classification branch comprises the following steps: performing convolution and upsampling on the fused feature map to generate detection layers of multiple scales; each detection layer comprises a regression branch for predicting a target bounding box and a classification branch for predicting a target category.
7. The method for detecting appearance of a collection terminal based on a binary feature aggregation network according to claim 1, characterized in that: The fused feature map is processed through a head network to separate a regression branch and a classification branch; the head network adopts a double decoupled head structure and generates a detection result including a target bounding box, a target category and a target confidence.
8. The method for detecting appearance of a collection terminal based on a binary feature aggregation network according to claim 1, characterized in that: The calculation formula of the classification loss of the loss function is: Where L is the classification loss, N is the number of samples, and y i is the true label of the i-th sample, p i is the predicted probability of the ith sample.
9. The method for detecting appearance of a collection terminal based on a binary feature aggregation network according to claim 1, characterized in that: The calculation formula of the classification loss of the loss function is: DFL(S i ,S i+1 )=-((y i+1 -y)log(S i )+(y-y i )log(S i+1 )) Among them, S i is the predicted value output by the network, S i+1 is the near prediction value output by the network, y is the near prediction value output by the network, y i is the integral value of the label, y i+1 is the proximity integral value of the label.
10. A collection terminal appearance detection device based on a binary feature aggregation network, characterized in that: It includes: A preprocessing unit, which is used to obtain a standardized image of the appearance of the acquisition terminal; An input unit, which is used to extract low-level features of the standardized image to obtain an input feature map; A binary feature aggregation unit, which is used to perform feature aggregation through at least one binary feature aggregation network; the first binary feature aggregation network feature aggregates the input feature map, and other binary feature aggregation networks feature aggregate the fused feature map output by the previous binary feature aggregation network; each binary feature aggregation network performs: firstly adjusting the number of channels of the input feature map or the fused feature map and evenly dividing it into two parts, one part of which is subjected to global feature extraction through a plurality of stacked neck network modules 1, and the other part of which is subjected to local feature extraction through a plurality of stacked neck network modules 2, and then the extracted global features and local features are feature fused and integrated to output as the fused feature map; A separation unit, which is used to separate the regression branch and the classification branch of the fused feature map; A loss calculation unit is used to first calculate the regression loss and classification loss of a loss function, and then optimize the training model used to represent and detect the appearance of the acquisition terminal through the loss function.