Street lamp brightness lack detection method and street lamp brightness lack detection device based on deep learning

Through the deep learning-based street light lack of light detection method, the combination of backbone network and detection network is used to solve the problems of low street light lack of light detection efficiency and high leakage detection rate, and efficient and accurate street light lack of light detection is achieved.

CN120472381AActive Publication Date: 2025-08-12STREAMAP TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510947435.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-08-12
Estimated Expiration
2045-07-10

AI Technical Summary

Technical Problem

In the management of smart city infrastructure, street lights are under-light detection problems with low detection efficiency and high missed detection rate.

Method used

The street light lack of light detection method based on deep learning is adopted. Through the combination of the backbone network and the detection network, the multi-scale hollow packet fusion module and the airspace frequency domain hybrid enhanced attention module are used to detect street light lack of light, and realize high-precision positioning and morphological determination.

Benefits of technology

It improves the efficiency of street light detection, reduces the missed detection rate, and achieves high-precision detection without manual participation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472381A_ABST
    Figure CN120472381A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of computer vision, and provides a street lamp brightness lack detection method and a street lamp brightness lack detection device based on deep learning. The method comprises the steps of obtaining a to-be-detected street lamp image; inputting the to-be-detected street lamp image into a street lamp brightness lack detection network comprising a backbone network and a detection network, and performing down-sampling on the to-be-detected street lamp image through a first convolution module in the backbone network to obtain a first feature map, a second convolution module in each feature extraction network in the backbone network is used for carrying out down-sampling on the feature map obtained by the previous layer of network to obtain a second feature map, and a multi-scale cavity grouping fusion module in each feature extraction network is used for carrying out different-scale feature extraction on the second feature map to obtain a third feature map; and the detection network carries out street lamp brightness lack detection based on the third feature map obtained by the at least two feature extraction networks so as to determine a target detection result. According to the invention, the efficiency of street lamp brightness lack detection can be improved, and the omission factor of street lamp brightness lack can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of computer vision technology, and in particular relates to a street lamp under-lighting detection method and a street lamp under-lighting detection device based on deep learning. Background Art

[0002] In smart city infrastructure management, streetlight outage detection requires systems to achieve dual-target detection, including high-precision positioning and morphological determination, in complex environments. However, traditional detection solutions, such as fixed threshold segmentation and manual inspections, suffer from low detection efficiency and high missed detection rates. Summary of the Invention

[0003] The embodiments of the present application provide a street lamp under-lighting detection method and a street lamp under-lighting detection device based on deep learning, which can improve the detection efficiency of street lamp under-lighting and reduce the missed detection rate of street lamp under-lighting.

[0004] In a first aspect, an embodiment of the present application provides a method for detecting streetlight under-lighting based on deep learning, comprising: Obtain the image of the street lamp to be detected; Inputting the street lamp image to be detected into a street lamp under-lighting detection network to obtain a target detection result, the target detection result including: an identification result of the street lamp in the street lamp image to be detected, and a position of the street lamp in the street lamp image to be detected when the identification result indicates that the street lamp is under-lighting; The street lamp under-lighting detection network includes a backbone network and a detection network; The backbone network sequentially includes a first convolution module and N sequentially connected feature extraction networks, and the N feature extraction networks sequentially include a second convolution module and a multi-scale hole group fusion module, where N is an integer greater than 2; The first convolution module is used to: downsample the street lamp image to be detected to obtain a first feature map; The second convolution module in each of the feature extraction networks is used to: downsample the feature map obtained by the previous layer network to obtain a second feature map; The multi-scale hole grouping fusion module in each feature extraction network is used to: perform feature extraction of different scales on the second feature map obtained by the second convolution module in the same feature extraction network to obtain a third feature map; The detection network is used to determine the target detection result based on the third feature map obtained by at least two of the feature extraction networks.

[0005] In an embodiment of the present application, by obtaining an image of a street lamp to be detected and inputting the image of the street lamp to be detected into a street lamp lack-lighting detection network including a backbone network and a detection network, the image of the street lamp to be detected can be downsampled by the first convolution module in the backbone network to obtain a first feature map, and the feature map obtained by the previous layer network can be downsampled by the second convolution module in each feature extraction network in the backbone network to obtain a second feature map, and features of different scales are extracted from the second feature map by the multi-scale hole grouping fusion module in each feature extraction network to obtain a third feature map, which can make full use of the correlation between features of different scales and significantly improve the feature extraction street lamp lack-lighting detection network's ability to capture street lamp lack-lighting in complex scenes. On this basis, when the detection network performs street lamp lack-lighting detection based on the third feature map obtained by at least two feature extraction networks, the missed detection rate of street lamp lack-lighting can be reduced, and the street lamp lack-lighting detection network in this scheme is constructed based on deep learning, and no human participation is required in the above-mentioned street lamp lack-lighting detection process, thereby improving the detection efficiency of street lamp lack-lighting.

[0006] In a second aspect, an embodiment of the present application provides a street lamp under-lighting detection device based on deep learning, comprising: An image acquisition module, used to acquire an image of a street lamp to be detected; a target detection module, configured to input the street lamp image to be detected into a street lamp under-lighting detection network to obtain a target detection result, the target detection result comprising: an identification result of the street lamp in the street lamp image to be detected, and a position of the street lamp in the street lamp image to be detected when the identification result indicates that the street lamp is under-lighting; The street lamp under-lighting detection network includes a backbone network and a detection network; The backbone network sequentially includes a first convolution module and N sequentially connected feature extraction networks, and the N feature extraction networks sequentially include a second convolution module and a multi-scale hole group fusion module, where N is an integer greater than 2; The first convolution module is used to: downsample the street lamp image to be detected to obtain a first feature map; The second convolution module in each of the feature extraction networks is used to: downsample the feature map obtained by the previous layer network to obtain a second feature map; The multi-scale hole grouping fusion module in each feature extraction network is used to: perform feature extraction of different scales on the second feature map obtained by the second convolution module in the same feature extraction network to obtain a third feature map; The detection network is used to determine the target detection result based on the third feature map obtained by at least two of the feature extraction networks.

[0007] In a third aspect, an embodiment of the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic device implements a street light under-lighting detection method based on deep learning as described in any one of the first aspects above.

[0008] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a computer, it implements the street lamp under-lighting detection method based on deep learning as described in the first aspect above.

[0009] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program. When the computer program is run, the street light under-lighting detection method based on deep learning as described in any one of the above-mentioned first aspects is executed.

[0010] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0012] Figure 1 This is a flowchart of a street lamp under-lighting detection method based on deep learning provided in an embodiment of the present application; Figure 2 This is another flowchart of the street lamp under-lighting detection method based on deep learning provided in an embodiment of the present application; Figure 3 This is a structural example diagram of a multi-scale hole group fusion module provided in an embodiment of the present application; Figure 4 This is a structural example diagram of the spatial-frequency hybrid enhanced attention module provided in an embodiment of the present application; Figure 5 This is a diagram illustrating an example of the structure of a streetlight under-lighting detection network provided in an embodiment of the present application; Figure 6 Schematic diagram of the structure of a street lamp under-lighting detection device based on deep learning provided in an embodiment of the present application; Figure 7 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0013] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0014] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0015] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0016] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0017] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0018] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0019] The deep learning-based street lamp under-lighting detection method provided in the embodiments of the present application can be applied to electronic devices such as mobile phones, tablet computers, wearable devices, vehicle-mounted devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). The embodiments of the present application do not impose any restrictions on the specific types of electronic devices.

[0020] See also Figure 1 , Figure 1 The following is a flowchart of a method for detecting streetlight under-lighting based on deep learning provided in an embodiment of the present application. As an example and not a limitation, the method includes the following steps: Step 101: Acquire an image of a street lamp to be detected.

[0021] The street lamp image to be detected may refer to an image containing street lamps. The number of street lamps in the street lamp image to be detected may be one or more, which is not limited in this application.

[0022] As an example but not a limitation, the streetlight image to be detected may be a normalized 3-channel RGB image with a resolution of 640*640, and the RGB image includes the streetlight.

[0023] As an example and not a limitation, a vehicle is equipped with a camera, and the vehicle can collect street light images through the camera while driving, or collect street light images through a camera installed at a street light pole, traffic signal pole, etc., and determine the collected street light image as the street light image to be detected and transmit it to the electronic device. The electronic device can realize street light lack detection of the street light image to be detected through step 102.

[0024] In step 102, the street lamp image to be detected is input into a street lamp under-lighting detection network to obtain a target detection result. The target detection result includes: a recognition result of the street lamp in the street lamp image to be detected, and a position of the street lamp in the street lamp image to be detected when the recognition result indicates that the street lamp is under-lighting.

[0025] The streetlight absence detection network in step 102 is a pre-trained streetlight absence detection network. After the electronic device inputs the image of the streetlight to be detected into the pre-trained streetlight absence detection network, the network can output a streetlight absence detection result (i.e., a target detection result) for the streetlight in the image, achieving high-precision streetlight location and morphology determination.

[0026] In one embodiment, the Figure 2 Steps 201 to 205 are shown to train the streetlight under-lighting detection network.

[0027] Step 201: Acquire multiple streetlight sample images, multiple streetlight sample images, predicted probabilities of streetlights being off in the multiple streetlight sample images, and position information of real frames and predicted frames in the multiple streetlight sample images.

[0028] The position information of the real frame includes but is not limited to the width, height, coordinates of the four sides, and coordinates of the center point of the real frame. The position information of the predicted frame includes but is not limited to the width, height, coordinates of the four sides, and coordinates of the center point of the predicted frame.

[0029] The electronic device can first obtain multiple street light sample images and the position information of the real frames in the multiple street light sample images, and then input the multiple street light sample images into the street light lack-lighting detection network to obtain the predicted probability of street light lack-lighting in the multiple street light sample images and the position information of the predicted frames in the multiple street light sample images.

[0030] The multiple sample streetlight images can be sourced from a high-quality, representative dataset of streetlight under-lighting images. This dataset can include a large number of nighttime urban road images (e.g., 3,000 images), covering a wide range of urban environments, weather conditions, streetlight types, and shooting angles, offering both high scene diversity and practicality. Given the low frequency of streetlight under-lighting in natural scenes and the scarcity of relevant data, the streetlight under-lighting image dataset emphasizes a combination of rarity and widespread coverage during its collection and annotation process, striving to cover as many under-lighting types (e.g., completely off, partially lit, flickering) as possible, as well as normal samples, to enhance the generalization capabilities and practical application of the streetlight under-lighting detection network.

[0031] This embodiment uses the PyTorch deep learning framework to build and optimize a streetlight detection network. Before selecting multiple streetlight sample images from the streetlight image dataset, the streetlight image dataset can be collected and cleaned, its labeled labels converted to the format required for streetlight detection network training, and the training and validation sets divided into a specific ratio (e.g., 8:2) to ensure that the streetlight detection network demonstrates good generalization capabilities even on untrained or unvalidated data.

[0032] To enhance the robustness and adaptability of the streetlight detection network, a series of data augmentation techniques can be implemented before the sample streetlight images are fed into the network. These techniques include, but are not limited to, randomly resizing the sample streetlight images, adding random noise, and simulating different weather conditions to simulate various scenarios encountered in the real world. These augmented images are then fed into the network, enabling it to learn stably under a wider range of conditions.

[0033] Step 202 : For the t-th streetlight sample image, determine the classification loss value of the t-th streetlight sample image based on the prediction probability of the t-th streetlight sample image and the prediction probabilities of the remaining streetlight sample images.

[0034] Among them, the remaining street light sample images are images other than the t-th street light sample image in the multiple street light sample images, the t-th street light sample image is any street light sample image in the multiple street light sample images, and t is a positive integer less than or equal to the number of the multiple street light sample images.

[0035] In a possible implementation, step 202 may include: Determine the average of the predicted probabilities for the remaining streetlight sample images; The predicted probability of the t-th street light sample image and the average of the predicted probabilities of the remaining street light sample images are weighted and summed to obtain the weighted predicted probability; Based on the weighted prediction probability and the prediction probability of the t-th street lamp sample image, the classification loss value of the t-th street lamp sample image is determined.

[0036] Optionally, the weight used in performing the weighted summation of the predicted probability of the t-th street lamp sample image and the average value of the predicted probabilities of the remaining street lamp sample images can be set according to an empirical value.

[0037] The electronic device uses the average value of the predicted probabilities of the remaining street light sample images to adjust the classification loss value generated by the t-th street light sample image, and can adaptively adjust the hyperparameters to solve the problems of unsatisfactory detection effect and large time consumption of the street light lack-of-light detection network caused by adjusting the hyperparameters.

[0038] The classification loss value of the t-th street light sample image can be expressed as follows:

[0039]

[0040] in, represents the classification loss value of the t-th street lamp sample image, represents the predicted probability of the t-th street light sample image, represents the average value of the predicted probabilities of the remaining street light sample images, represents the hyperparameter, 1 represents the weight of the average value of the predicted probability of the remaining street light sample images, represents the weight of the predicted probability of the t-th street light sample image, Indicates the number of street light sample images, represents the i-th remaining street light sample image.

[0041] Step 203: Based on the position information of the real frame and the position information of the predicted frame in the t-th street light sample image, determine the intersection-over-union loss value of the t-th street light sample image, and the L1 norm of the difference between the coordinates of the four sides of the real frame in the t-th street light sample image and the coordinates of the corresponding sides in the predicted frame.

[0042] Taking the real box as an example, the four sides of the real box can include the left boundary, right boundary, upper boundary and lower boundary. For the real box represented by (x1, y1, x2, y2), the coordinates of the left boundary are x = x1, the coordinates of the right boundary are x = x2, the coordinates of the upper boundary are y = y1, and the coordinates of the lower boundary are y = y2.

[0043] Step 204 : Determine the regression loss value of the tth streetlight sample image based on the intersection-over-union loss value, the L1 norm, and the width and height of the true box in the tth streetlight sample image.

[0044] The regression loss value of the t-th street light sample image can be calculated using the following formula:

[0045]

[0046]

[0047]

[0048] in, represents the intersection-over-union loss value of the t-th streetlight sample image, and Indicates the L1 norm of the two edges corresponding to the width, and Indicates the L1 norm of the two edges corresponding to the height, and Represent the width and height of the real box respectively, Represents the regression loss value of the t-th street light sample image.

[0049] By taking the difference between the four sides of the true box and the predicted box, the position information of the predicted box in the street lamp image to be tested is introduced, which suppresses the regression box that may appear in the middle and lower part of the street lamp image to be tested, thereby improving the accuracy and robustness of detection. At the same time, the position information is modified by exponential modification to avoid the phenomenon of optimization stagnation when the gradient reaches 0.

[0050] Step 205 : training a streetlight under-lighting detection network based on the classification loss values and / or regression loss values of the plurality of streetlight sample images.

[0051] In one embodiment, a streetlight under-lighting detection network may be trained based on classification loss values of a plurality of streetlight sample images.

[0052] In another embodiment, a streetlight under-lighting detection network may be trained based on regression loss values of a plurality of streetlight sample images.

[0053] In another embodiment, the classification loss value and regression loss value of the tth street light sample image can be weightedly summed to obtain the target loss value of the tth street light sample image; based on the target loss values of multiple street light sample images, the street light under-lighting detection network is trained.

[0054] The target loss value of the t-th street light sample image can be calculated using the following formula.

[0055]

[0056] in, represents the target loss value of the t-th street light sample image, Represents the weight of the classification loss value of the t-th street lamp sample image, The weight representing the regression loss value of the t-th street light sample image.

[0057] The above target loss value improves the positioning accuracy by introducing distance constraints of spatial positions and adaptive mining of difficult samples.

[0058] Considering the significant uncertainty in the shape, position and boundary definition of unlit street lamps, training the street lamp unlit detection network with the above-mentioned target loss value can enhance the sensitivity to small street lamp areas and long street lamps while retaining the basic overlap evaluation index, thereby improving the boundary fitting quality and positioning accuracy.

[0059] In one embodiment, when the recognition result indicates that the street lamp is not lacking in brightness, the above-mentioned target detection result may include the position of the street lamp in the street lamp image to be detected, or may not include the position of the street lamp in the street lamp image to be detected. This application does not limit this.

[0060] As an example and not a limitation, if the recognition result for a streetlight indicates that it is not bright, a prediction box with a first color can be used to select the streetlight in the image of the streetlight to be detected. If the recognition result for a streetlight indicates that it is not bright, a prediction box with a second color can be used to select the streetlight in the image of the streetlight to be detected, or no prediction box can be used. The first color and the second color are different colors to facilitate distinguishing whether the streetlight selected by the prediction box is bright or not.

[0061] In one embodiment, the street lamp under-lighting detection network includes a backbone network and a detection network; The backbone network includes a first convolutional module and N sequentially connected feature extraction networks, and the N feature extraction networks include a second convolutional module and a multi-scale hollow grouping fusion module (MSHG), where N is an integer greater than 2. The first convolution module is used to downsample the street lamp image to be detected to obtain a first feature map; The second convolution module in each feature extraction network is used to downsample the feature map obtained by the previous layer network to obtain a second feature map; The multi-scale hole grouping fusion module in each feature extraction network is used to extract features of different scales on the second feature map obtained by the second convolution module in the same feature extraction network to obtain a third feature map; The detection network is used to determine a target detection result based on a third feature map obtained by at least two feature extraction networks.

[0062] Among them, the first convolution module, the second convolution module, and the third and fourth convolution modules below are mainly composed of convolution, batch normalization, and activation functions to downsample the input image or feature map.

[0063] As a feature extraction module, the multi-scale hole grouping fusion module is mainly used to extract features, obtain richer feature maps, enhance the expression ability of the street light under-lighting detection network, and solve the problems of variable street light target size and weak semantic expression in night scenes.

[0064] The street light under-lighting detection network gradually extracts features through multiple convolution modules such as the first convolution module and the second convolution module in N feature extraction networks, which can ensure that more features are extracted.

[0065] Optionally, the at least two feature extraction networks may be N feature extraction networks, or may be a portion of the N feature extraction networks. As an example and not a limitation, N is 4, and the detection network may determine the target detection result based on the third feature graphs of the three feature extraction networks that are ranked second, third, and fourth in execution order among the four feature extraction networks.

[0066] In one embodiment, the multi-scale dilation grouping fusion module includes: a plurality of dilation convolution groups with different grouping rates and dilation rates, a first splicing module, a first channel attention module, a first spatial attention module, and a multi-scale dilation fusion module; The dilated convolution group is used to extract features from the second feature map to obtain a fourth feature map. The first splicing module is used to: splice the fourth feature maps obtained by multiple hole convolution groups to obtain a first spliced feature map; The first channel attention module is used to perform channel attention enhancement on the first spliced feature map to obtain a first enhanced feature map; The first spatial attention module is used to perform spatial attention enhancement on the first spliced feature map to obtain a second enhanced feature map; The multi-scale hole fusion module is used to sum the first enhanced feature map and the second enhanced feature map to obtain a third feature map.

[0067] The multi-scale hole group fusion module introduces hole convolution groups with different grouping rates and hole rates, which can extract features at different spatial scales. First, it can increase the receptive field. The use of hole convolution groups, especially different hole rates and grouping rates, can significantly expand the receptive field and reduce the amount of computation without losing resolution. This is crucial for capturing information with global distribution characteristics (such as continuous unlit road sections) in complex night scenes. Secondly, it can capture multi-scale information. In image processing tasks, especially target detection objects, they can appear in different sizes and shapes. By using multiple hole convolution groups in parallel, features of different scales can be captured simultaneously, which helps improve the adaptability and recognition ability of the street light detection network to various scale features in the image to be detected, and enhances the ability of the street light detection network to obtain local details of street lights (such as broken street lights) at different distances.

[0068] As an example and not a limitation, the multi-scale spatial grouping fusion module implements feature extraction at different scales through four parallel dilated convolution groups, each of which is configured with a different grouping rate and dilation rate. The first dilated convolution group has a grouping rate of 1 and a dilation rate of 0, and uses a 1x1 convolution kernel to directly extract features without changing the spatial scale. The second dilated convolution group uses a 3x3 convolution kernel, a grouping rate of 8, and a dilation rate of 2 to moderately expand the receptive field; the third dilated convolution group uses a 5x5 convolution kernel, a grouping rate of 4, and a dilation rate of 3 to further expand the receptive field to capture a wider range of contextual information; the fourth dilated convolution group uses a 7x7 convolution kernel, a grouping rate of 2, and a dilation rate of 4 to provide the widest receptive field.

[0069] Among them, the first splicing module splices the fourth feature maps obtained by multiple hole convolution groups in the channel dimension to synthesize a comprehensive feature map (i.e., the first splicing feature map). The first splicing feature map is calibrated by the first channel attention module and the first spatial attention module in parallel to finally obtain the enhanced feature map (i.e., the third feature map).

[0070] Optionally, the first channel attention module and the second channel attention module described below may include an average pooling module, a maximum pooling module, an addition module, a multilayer perceptron (MLP), and a multiplication module. The first spatial attention module may include a fifth convolution module and a multiplication module. The convolution kernel of the fifth convolution module may be 1x1. Based on this, the first spatial attention module has a lightweight design, adapted for edge computing devices, and enables real-time processing. The multilayer perceptron may include two fully connected layers and a sigmoid activation function.

[0071] The first channel attention module and the second channel attention module described below are dual-pooled channel attention modules. By simultaneously performing global average pooling and global maximum pooling on the input feature map, they can capture the mean and corresponding extreme values along the channel dimension. A fully connected layer then compresses and restores the channels to generate channel attention weights, which are then applied to each channel to dynamically calibrate their importance.

[0072] like Figure 3 The figure shows an example structure diagram of the multi-scale hole grouping fusion module. Figure 3 AvgPool represents the average pooling module, MaxPool represents the maximum pooling module, ɡ represents the grouping rate, and r represents the void rate. Figure 3 and the following text Figure 4 B represents the batch size, C represents the number of channels, 4C represents 4 times of C, H represents height, and W represents width.

[0073] If the second feature map is , then the implementation process of the multi-scale hole grouping fusion module can be expressed by the following formula.

[0074] First splicing feature map It is expressed as follows:

[0075] in, Represents the splicing module, represents the fourth feature map of the j-th dilated convolution group, Represents the sigmoid activation function.

[0076] First enhanced feature map It is expressed as follows:

[0077]

[0078]

[0079]

[0080] in, Represents the feature map output by the average pooling module, Represents the feature map output by the maximum pooling module, Represents the feature map output by the multilayer perceptron, represents two fully connected layers, Represents element-wise dot product.

[0081] Second enhanced feature map It is expressed as follows:

[0082] in, Represents the convolution module in the first spatial attention module.

[0083] The third characteristic map It is expressed as follows:

[0084] In one embodiment, the detection network includes a feature fusion network and a detection head; The feature fusion network is used to: fuse the third feature maps obtained by at least two feature extraction networks to obtain a fused feature map; The detection head is used to determine the target detection results based on the fused feature map.

[0085] The feature fusion network can fuse features at multiple levels, so that the street light under-lighting detection network takes into account the semantic information of different feature layers, enhances the detection ability of street lights of different sizes, and improves the recall rate.

[0086] In order to ensure that more features are extracted while avoiding a surge in computational complexity and ensuring the real-time performance of streetlight under-lighting detection, in one embodiment, N is 4, and the four feature extraction networks are arranged in an execution order, and the feature fusion network includes: a first upsampling module, a second splicing module, a first cross-stage partial (CSP) module, a second upsampling module, a third splicing module, a second cross-stage partial module, a first spatial-frequency hybrid enhanced attention (SFHEA) module, a third convolution module, a fourth splicing module, a third cross-stage partial module, a second spatial-frequency hybrid enhanced attention module, a fourth convolution module, a fifth splicing module, a fourth cross-stage partial module, and a third spatial-frequency hybrid enhanced attention module; The first upsampling module is used to upsample the third feature map obtained by the fourth feature extraction network to obtain a fifth feature map; The second splicing module is used to: splice the fifth feature map and the third feature map obtained by the third feature extraction network to obtain a second spliced feature map; The first cross-stage local module is used to: perform channel fusion on the second spliced feature map to obtain a sixth feature map; The second upsampling module is used to upsample the sixth feature map to obtain a seventh feature map; The third splicing module is used to: splice the seventh feature map and the third feature map obtained by the second feature extraction network to obtain a third spliced feature map; The second cross-stage local module is used to: perform channel fusion on the third spliced feature map to obtain an eighth feature map; The first spatial-frequency hybrid enhanced attention module is used to focus the eighth feature map on the unlit streetlight to obtain a first fused feature map; The third convolution module is used to downsample the eighth feature map to obtain a ninth feature map; The fourth splicing module is used to: splice the ninth feature map and the sixth feature map to obtain a fourth spliced feature map; The third cross-stage local module is used to: perform channel fusion on the fourth spliced feature map to obtain a tenth feature map; The second spatial-frequency hybrid enhanced attention module is used to focus the tenth feature map on the unlit streetlight to obtain a second fused feature map; The fourth convolution module is used to downsample the tenth feature map to obtain an eleventh feature map; The fifth splicing module is used to: splice the eleventh feature map with the third feature map obtained by the fourth feature extraction network to obtain a fifth spliced feature map; The fourth cross-stage local module is used to: perform channel fusion on the fifth spliced feature map to obtain a twelfth feature map; The third spatial-frequency hybrid enhanced attention module is used to focus the twelfth feature map on the unlit streetlight to obtain a third fused feature map; Among them, the first fused feature map, the second fused feature map and the third fused feature map are all fused feature maps obtained by the feature fusion network.

[0087] This embodiment uses multiple spatial-frequency hybrid enhanced attention modules, such as the first spatial-frequency hybrid enhanced attention module, the second spatial-frequency hybrid enhanced attention module, and the third spatial-frequency hybrid enhanced attention module, to fuse spatial attention and frequency-domain transformation features, guide the street light under-lighting detection network to focus on abnormal brightness areas in space, suppress high-frequency noise interference in the frequency domain, and achieve efficient perception of faint under-lighting targets in night images.

[0088] Optionally, the first spatial-frequency hybrid enhanced attention module, the second spatial-frequency hybrid enhanced attention module, and the third spatial-frequency hybrid enhanced attention module each include: a second channel attention module, a frequency-domain hybrid filtering enhancement module, and a second spatial attention module; The second channel attention module is used to perform channel attention enhancement on the input feature map to obtain a third enhanced feature map; The frequency domain hybrid filtering enhancement module is used to: perform frequency domain enhancement on the third enhanced feature map to obtain a fourth enhanced feature map; The second spatial attention module is used to: perform spatial attention enhancement on the fourth enhanced feature map to obtain a fifth enhanced feature map; The feature fusion module is used to perform element-by-element point multiplication on the third enhanced feature map and the fifth enhanced feature map to obtain a fused feature map.

[0089] The second spatial attention module may include the sixth convolution module and The convolution kernel of the sixth convolution module can be 1x1. On this basis, the second spatial attention module is lightweight and adapts to edge computing devices to achieve real-time processing.

[0090] The second spatial attention module performs a 1x1 convolution on the frequency-domain enhanced features (i.e., the fourth enhanced feature map) to generate a spatial response map, corresponding to the spatial attention weight matrix, focusing on the key areas after frequency-domain enhancement. Finally, the feature fusion module performs an element-by-element dot product on the third and fifth enhanced feature maps to achieve feature fusion, enabling joint optimization of channels, frequency domains, and space.

[0091] In one embodiment, the frequency domain hybrid filtering enhancement module includes, in sequence: a Fourier transform module, a high pass filter (HPF), an inverse Fourier transform module, a subtraction module, and a weighting module; The Fourier transform module is used to: perform Fourier transform on the third enhanced feature map to obtain a thirteenth feature map; The high-pass filter is used to remove the signal below the cutoff frequency in the thirteenth feature map to obtain the fourteenth feature map; The inverse Fourier transform module is used to: perform an inverse Fourier transform on the fourteenth feature map to obtain a fifteenth feature map; The subtraction module is configured to: subtract the fifteenth feature map from the third enhanced feature map to obtain a sixteenth feature map; The weighting module is used to perform weighted summation on the fifteenth feature map and the sixteenth feature map to obtain a fourth enhanced feature map.

[0092] The Fourier transform module can be implemented using the FFT2 function, which implements a two-dimensional fast Fourier transform (2D FFT). The inverse Fourier transform module can be implemented using the IFFT2 function, which implements a two-dimensional inverse fast Fourier transform.

[0093] The high-pass filter has an adjustable cutoff frequency. This high-pass filter enhances streetlight edge details (such as cracked lampshades and broken filaments) in low-light nighttime scenes. The weighting module also preserves low-frequency brightness information to avoid over-sharpening noise. The second channel attention module suppresses irrelevant high-frequency noise channels such as headlight reflections and moonlight interference. The second spatial attention module dynamically focuses on potentially dimmed areas (such as streetlights obscured by trees), accurately distinguishing between normal illumination and partially dark areas.

[0094] The implementation process of the three spatial-frequency hybrid enhanced attention modules, namely the first spatial-frequency hybrid enhanced attention module, the second spatial-frequency hybrid enhanced attention module and the third spatial-frequency hybrid enhanced attention module, can be expressed by the following formula. The input of the spatial-frequency hybrid enhanced attention module can be expressed as .

[0095] The third enhanced feature map output by the second channel attention module The calculation formula of can refer to the calculation formula of the first enhanced feature map, which will not be repeated here.

[0096] Thirteenth characteristic graph It can be expressed as follows:

[0097] Fourteenth characteristic diagram It can be expressed as follows:

[0098] in, and represents the two-dimensional coordinates in the Fourier frequency domain, Indicates the cutoff frequency ratio, which is used to control the proportion of high-frequency signals retained. and Represent the height and width of the thirteenth feature map respectively, The weights of feature maps of different scales can be dynamically adjusted.

[0099] Fifteenth characteristic graph It can be expressed as follows:

[0100] Fourth enhanced feature map It can be expressed as follows:

[0101] in, represents the weight of the fifteenth feature map, Represents the weight of the sixteenth feature map.

[0102] Fifth enhanced feature map It can be expressed as follows:

[0103] in, Represents the sixth convolutional module.

[0104] Fusion feature map It can be expressed as follows:

[0105] like Figure 4 Shown is a structural example diagram of the spatial-frequency hybrid enhanced attention module.

[0106] Figure 4 The d=(-2, -1) in the figure indicates that Fourier transform or inverse Fourier transform is performed on the height and width channels of the feature map.

[0107] In one embodiment, the detection head includes: a first detection head, a second detection head, a third detection head and a processing module; The first detection head is used to: determine a first detection result based on the first fused feature map; The second detection head is used to: determine a second detection result based on the second fused feature map; The third detection head is used to: determine a third detection result based on the third fused feature map; The processing module is used to process the first detection result, the second detection result and the third detection result to obtain a target detection result.

[0108] This embodiment performs street lamp under-lighting detection through multiple detection heads, such as a first detection head, a second detection head, and a third detection head, and can perform under-lighting detection on street lamps of different sizes.

[0109] like Figure 5 The following is an example diagram of the structure of a street light under-lighting detection network. Figure 5 The processing module in the detection head is not shown. Figure 5As shown, the number of channels of the second feature map output by the second convolution module in the first feature extraction network in the execution order is 64, and the resolution is 160*160; the number of channels of the second feature map output by the second convolution module in the second feature extraction network is 128, and the resolution is 80*80; the number of channels of the second feature map output by the second convolution module in the third feature extraction network is 256, and the resolution is 40*40; the number of channels of the second feature map output by the second convolution module in the fourth feature extraction network is 256, and the resolution is 20*20. The number of channels of the fifth feature map is 256, and the resolution is 40*40; the number of channels of the second spliced feature map is 512, and the resolution is 40*40; the number of channels of the sixth feature map is 256, and the resolution is 40*40; the number of channels of the seventh feature map is 256, and the resolution is 80*80; the number of channels of the third spliced feature map is 384, and the resolution is 80*80; the number of channels of the eighth feature map is 128, and the resolution is 80*80; the number of channels of the ninth feature map is 128, and the resolution is 40*40; the number of channels of the fourth spliced feature map is 384, and the resolution is 40*40; the number of channels of the tenth feature map is 128, and the resolution is 40*40; the number of channels of the eleventh feature map is 128, and the resolution is 20*20; the number of channels of the fifth spliced feature map is 384, and the resolution is 20*20; the number of channels of the twelfth feature map is 128, and the resolution is 20*20.

[0110] The streetlight detection network provided in this embodiment is an end-to-end deep neural network that not only efficiently adapts to complex scenarios but also exhibits enhanced robustness and generalization capabilities. By improving several key aspects, including feature extraction, attention mechanisms, and loss function design, this embodiment enhances the network's perception capabilities and positioning accuracy in complex scenarios.

[0111] The end-to-end learnable feature enables the street light under-lighting detection network to adapt to different street light source distributions (such as the spectrum differences between LED arrays and sodium lamps), significantly improving the detection rate and positioning accuracy of under-lighting street lights in complex environments, and providing highly robust visual analysis support for smart city lighting systems.

[0112] The embodiment of the present application obtains an image of a street lamp to be detected and inputs the image of the street lamp to be detected into a street lamp lack-lighting detection network including a backbone network and a detection network. The first convolution module in the backbone network can downsample the image of the street lamp to be detected to obtain a first feature map. The second convolution module in each feature extraction network in the backbone network can downsample the feature map obtained by the previous layer network to obtain a second feature map. The multi-scale void grouping fusion module in each feature extraction network can extract features of different scales from the second feature map to obtain a third feature map. The third feature map can make full use of the correlation between features of different scales and significantly improve the ability of the feature extraction street lamp lack-lighting detection network to capture street lamp lack-lighting in complex scenes. On this basis, when the detection network performs street lamp lack-lighting detection based on the third feature map obtained by at least two feature extraction networks, the missed detection rate of street lamp lack-lighting can be reduced. The street lamp lack-lighting detection network in this solution is constructed based on deep learning. No human participation is required in the above-mentioned street lamp lack-lighting detection process, thereby improving the detection efficiency of street lamp lack-lighting.

[0113] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0114] Corresponding to the street lamp under-lighting detection method based on deep learning described in the above embodiment, Figure 6 A structural schematic diagram of a street lamp under-lighting detection device based on deep learning provided in an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.

[0115] Reference Figure 6 , the device comprises: Image acquisition module 601, used to acquire the image of the street lamp to be detected; The target detection module 602 is configured to input the street lamp image to be detected into a street lamp under-lighting detection network to obtain a target detection result, wherein the target detection result includes: an identification result of the street lamp in the street lamp image to be detected, and a position of the street lamp in the street lamp image to be detected when the identification result indicates that the street lamp is under-lighting; The street lamp under-lighting detection network includes a backbone network and a detection network; The backbone network sequentially includes a first convolution module and N sequentially connected feature extraction networks, and the N feature extraction networks sequentially include a second convolution module and a multi-scale hole group fusion module, where N is an integer greater than 2; The first convolution module is used to: downsample the street lamp image to be detected to obtain a first feature map; The second convolution module in each of the feature extraction networks is used to: downsample the feature map obtained by the previous layer network to obtain a second feature map; The multi-scale hole grouping fusion module in each feature extraction network is used to: perform feature extraction of different scales on the second feature map obtained by the second convolution module in the same feature extraction network to obtain a third feature map; The detection network is used to determine the target detection result based on the third feature map obtained by at least two of the feature extraction networks.

[0116] Optionally, the multi-scale hole grouping fusion module includes: a plurality of hole convolution groups with different grouping rates and hole rates, a first splicing module, a first channel attention module, a first spatial attention module and a multi-scale hole fusion module; The dilated convolution group is used to: perform feature extraction on the second feature map to obtain a fourth feature map; The first splicing module is used to: splice the fourth feature maps obtained by the multiple dilated convolution groups to obtain a first spliced feature map; The first channel attention module is used to: perform channel attention enhancement on the first spliced feature map to obtain a first enhanced feature map; The first spatial attention module is used to: perform spatial attention enhancement on the first spliced feature map to obtain a second enhanced feature map; The multi-scale hole fusion module is used to: sum the first enhanced feature map and the second enhanced feature map to obtain the third feature map.

[0117] Optionally, the detection network includes a feature fusion network and a detection head; The feature fusion network is used to: fuse the third feature maps obtained by at least two feature extraction networks to obtain a fused feature map; The detection head is used to determine the target detection result based on the fused feature map.

[0118] Optionally, N is 4, the four feature extraction networks are sorted in order of execution, and the feature fusion network includes: a first upsampling module, a second splicing module, a first cross-stage local module, a second upsampling module, a third splicing module, a second cross-stage local module, a first spatial-frequency hybrid enhanced attention module, a third convolution module, a fourth splicing module, a third cross-stage local module, a second spatial-frequency hybrid enhanced attention module, a fourth convolution module, a fifth splicing module, a fourth cross-stage local module and a third spatial-frequency hybrid enhanced attention module; The first upsampling module is used to: upsample the third feature map obtained by the fourth feature extraction network to obtain a fifth feature map; The second splicing module is used to: splice the fifth feature map and the third feature map obtained by the third feature extraction network to obtain a second spliced feature map; The first cross-stage local module is used to: perform channel fusion on the second spliced feature map to obtain a sixth feature map; The second upsampling module is configured to: upsample the sixth feature map to obtain a seventh feature map; The third splicing module is used to: splice the seventh feature map and the third feature map obtained by the second feature extraction network to obtain a third spliced feature map; The second cross-stage local module is used to: perform channel fusion on the third spliced feature map to obtain an eighth feature map; The first spatial-frequency hybrid enhanced attention module is used to focus the eighth feature map on the dimly lit streetlight to obtain a first fused feature map; The third convolution module is used to: downsample the eighth feature map to obtain a ninth feature map; The fourth splicing module is used to: splice the ninth feature map and the sixth feature map to obtain a fourth spliced feature map; The third cross-stage local module is used to: perform channel fusion on the fourth spliced feature map to obtain a tenth feature map; The second spatial-frequency hybrid enhanced attention module is used to focus the tenth feature map on the unlit streetlight to obtain a second fused feature map; The fourth convolution module is used to: downsample the tenth feature map to obtain an eleventh feature map; The fifth splicing module is used to: splice the eleventh feature map with the third feature map obtained by the fourth feature extraction network to obtain a fifth spliced feature map; The fourth cross-stage local module is used to: perform channel fusion on the fifth spliced feature map to obtain a twelfth feature map; The third spatial-frequency hybrid enhanced attention module is used to focus the twelfth feature map on the unlit streetlight to obtain a third fused feature map; Among them, the first fused feature map, the second fused feature map and the third fused feature map are all the fused feature maps obtained by the feature fusion network.

[0119] Optionally, the first spatial-frequency hybrid enhanced attention module, the second spatial-frequency hybrid enhanced attention module, and the third spatial-frequency hybrid enhanced attention module each include: a second channel attention module, a frequency-domain hybrid filtering enhancement module, and a second spatial attention module; The second channel attention module is used to: perform channel attention enhancement on the input feature map to obtain a third enhanced feature map; The frequency domain hybrid filtering enhancement module is used to: perform frequency domain enhancement on the third enhanced feature map to obtain a fourth enhanced feature map; The second spatial attention module is used to: perform spatial attention enhancement on the fourth enhanced feature map to obtain a fifth enhanced feature map; A feature fusion module is used to perform element-by-element point multiplication on the third enhanced feature map and the fifth enhanced feature map to obtain the fused feature map.

[0120] Optionally, the frequency domain hybrid filtering enhancement module includes, in sequence: a Fourier transform module, a high-pass filter, an inverse Fourier transform module, a subtraction module and a weighting module; The Fourier transform module is used to: perform Fourier transform on the third enhanced feature map to obtain a thirteenth feature map; The high-pass filter is used to remove signals below a cutoff frequency in the thirteenth characteristic graph to obtain a fourteenth characteristic graph; The inverse Fourier transform module is used to: perform an inverse Fourier transform on the fourteenth feature map to obtain a fifteenth feature map; The subtraction module is configured to: subtract the fifteenth feature map from the third enhanced feature map to obtain a sixteenth feature map; The weighting module is used to perform weighted summation on the fifteenth feature map and the sixteenth feature map to obtain the fourth enhanced feature map.

[0121] Optionally, the detection head includes: a first detection head, a second detection head, a third detection head and a processing module; The first detection head is used to: determine a first detection result based on the first fusion feature map; The second detection head is used to: determine a second detection result based on the second fused feature map; The third detection head is used to: determine a third detection result based on the third fusion feature map; The processing module is used to process the first detection result, the second detection result and the third detection result to obtain the target detection result.

[0122] Optionally, the above device further includes: A data acquisition module is used to acquire a plurality of streetlight sample images and a plurality of the streetlight sample images, a predicted probability of a streetlight being unlit in the plurality of the streetlight sample images, and position information of a real frame and a predicted frame in the plurality of the streetlight sample images; a first determining module, configured to determine, for the tth street light sample image, a classification loss value of the tth street light sample image based on the prediction probability of the tth street light sample image and the prediction probabilities of the remaining street light sample images, where the remaining street light sample images are images other than the tth street light sample image among the plurality of street light sample images, the tth street light sample image is any one of the plurality of street light sample images, and t is a positive integer less than or equal to the number of the plurality of street light sample images; a second determination module, configured to determine, based on position information of the true frame and position information of the predicted frame in the t-th street lamp sample image, an intersection-over-union loss value of the t-th street lamp sample image, and an L1 norm of differences between coordinates of four sides of the true frame in the t-th street lamp sample image and coordinates of corresponding sides in the predicted frame; a third determining module, configured to determine a regression loss value of the tth street lamp sample image based on the intersection-over-union loss value of the tth street lamp sample image, the L1 norm, and the width and height of the true box in the tth street lamp sample image; A network training module is used to train the street lamp under-lighting detection network based on the classification loss values and / or regression loss values of the plurality of street lamp sample images.

[0123] Optionally, the first determining module is specifically configured to: Determining an average value of the predicted probabilities of the remaining street lamp sample images; Performing a weighted summation on the predicted probability of the t-th street lamp sample image and the average of the predicted probabilities of the remaining street lamp sample images to obtain a weighted predicted probability; Based on the weighted prediction probability and the prediction probability of the tth street lamp sample image, a classification loss value of the tth street lamp sample image is determined.

[0124] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.

[0125] Figure 7 This is a schematic diagram of the structure of the electronic device provided in the embodiment of the present application. Figure 7 As shown, the electronic device 7 of this embodiment includes: at least one processor 70 ( Figure 7Only one is shown in the figure), a memory 71, and a computer program 72 stored in the memory 71 and executable on the at least one processor 70, wherein the processor 70 implements the steps of any of the above-mentioned method embodiments when executing the computer program 72.

[0126] The electronic device may include, but is not limited to, a processor 70 and a memory 71. Those skilled in the art will understand that Figure 7 It is only an example of the electronic device 7 and does not constitute a limitation on the electronic device 7. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include input and output devices, network access devices, etc.

[0127] The processor 70 may be a central processing unit (CPU), or other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor.

[0128] In some embodiments, the memory 71 may be an internal storage unit of the electronic device 7, such as a hard drive or memory of the electronic device 7. In other embodiments, the memory 71 may also be an external storage device of the electronic device 7, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash card, etc. equipped on the electronic device 7. Furthermore, the memory 71 may include both an internal storage unit of the electronic device 7 and an external storage device. The memory 71 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory 71 may also be used to temporarily store data that has been output or is about to be output.

[0129] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0130] If the integrated unit is implemented as a software functional unit and sold or used as a standalone product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process steps in the above-mentioned method embodiments by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a device / electronic device, a recording medium, computer memory, read-only memory (ROM), random access memory (RAM), an electrical carrier signal, a telecommunications signal, and a software distribution medium. Examples include a USB flash drive, a removable hard drive, a magnetic disk, or an optical disk.

[0131] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0132] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0133] In the embodiments provided in the present application, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0134] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0135] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A streetlight under-lighting detection method based on deep learning, characterized in that: include: Obtain the image of the street lamp to be detected; Inputting the street lamp image to be detected into a street lamp under-lighting detection network to obtain a target detection result, the target detection result including: an identification result of the street lamp in the street lamp image to be detected, and a position of the street lamp in the street lamp image to be detected when the identification result indicates that the street lamp is under-lighting; The street lamp under-lighting detection network includes a backbone network and a detection network; The backbone network sequentially includes a first convolution module and N sequentially connected feature extraction networks, and the N feature extraction networks sequentially include a second convolution module and a multi-scale hole group fusion module, where N is an integer greater than 2; The first convolution module is used to: downsample the street lamp image to be detected to obtain a first feature map; The second convolution module in each of the feature extraction networks is used to: downsample the feature map obtained by the previous layer network to obtain a second feature map; The multi-scale hole grouping fusion module in each feature extraction network is used to: perform feature extraction of different scales on the second feature map obtained by the second convolution module in the same feature extraction network to obtain a third feature map; The detection network is used to determine the target detection result based on the third feature map obtained by at least two of the feature extraction networks.

2. The street lamp under-lighting detection method based on deep learning according to claim 1, characterized in that: The multi-scale hole grouping fusion module includes: a plurality of hole convolution groups with different grouping rates and hole rates, a first splicing module, a first channel attention module, a first spatial attention module and a multi-scale hole fusion module; The dilated convolution group is used to: perform feature extraction on the second feature map to obtain a fourth feature map; The first splicing module is used to: splice the fourth feature maps obtained by the multiple dilated convolution groups to obtain a first spliced feature map; The first channel attention module is used to: perform channel attention enhancement on the first spliced feature map to obtain a first enhanced feature map; The first spatial attention module is used to: perform spatial attention enhancement on the first spliced feature map to obtain a second enhanced feature map; The multi-scale hole fusion module is used to: sum the first enhanced feature map and the second enhanced feature map to obtain the third feature map.

3. The street lamp under-lighting detection method based on deep learning according to claim 1, characterized in that: The detection network includes a feature fusion network and a detection head; The feature fusion network is used to: fuse the third feature maps obtained by at least two feature extraction networks to obtain a fused feature map; The detection head is used to determine the target detection result based on the fused feature map.

4. The street lamp under-lighting detection method based on deep learning according to claim 3, characterized in that: N is 4, the four feature extraction networks are sorted in order of execution, and the feature fusion network includes: a first upsampling module, a second splicing module, a first cross-stage local module, a second upsampling module, a third splicing module, a second cross-stage local module, a first spatial-frequency hybrid enhanced attention module, a third convolution module, a fourth splicing module, a third cross-stage local module, a second spatial-frequency hybrid enhanced attention module, a fourth convolution module, a fifth splicing module, a fourth cross-stage local module, and a third spatial-frequency hybrid enhanced attention module; The first upsampling module is used to: upsample the third feature map obtained by the fourth feature extraction network to obtain a fifth feature map; The second splicing module is used to: splice the fifth feature map and the third feature map obtained by the third feature extraction network to obtain a second spliced feature map; The first cross-stage local module is used to: perform channel fusion on the second spliced feature map to obtain a sixth feature map; The second upsampling module is configured to: upsample the sixth feature map to obtain a seventh feature map; The third splicing module is used to: splice the seventh feature map and the third feature map obtained by the second feature extraction network to obtain a third spliced feature map; The second cross-stage local module is used to: perform channel fusion on the third spliced feature map to obtain an eighth feature map; The first spatial-frequency hybrid enhanced attention module is used to focus the eighth feature map on the dimly lit streetlight to obtain a first fused feature map; The third convolution module is used to: downsample the eighth feature map to obtain a ninth feature map; The fourth splicing module is used to: splice the ninth feature map and the sixth feature map to obtain a fourth spliced feature map; The third cross-stage local module is used to: perform channel fusion on the fourth spliced feature map to obtain a tenth feature map; The second spatial-frequency hybrid enhanced attention module is used to focus the tenth feature map on the unlit streetlight to obtain a second fused feature map; The fourth convolution module is used to: downsample the tenth feature map to obtain an eleventh feature map; The fifth splicing module is used to: splice the eleventh feature map with the third feature map obtained by the fourth feature extraction network to obtain a fifth spliced feature map; The fourth cross-stage local module is used to: perform channel fusion on the fifth spliced feature map to obtain a twelfth feature map; The third spatial-frequency hybrid enhanced attention module is used to focus the twelfth feature map on the unlit streetlight to obtain a third fused feature map; Among them, the first fused feature map, the second fused feature map and the third fused feature map are all the fused feature maps obtained by the feature fusion network.

5. The street lamp under-lighting detection method based on deep learning according to claim 4, characterized in that: The first spatial-frequency hybrid enhanced attention module, the second spatial-frequency hybrid enhanced attention module, and the third spatial-frequency hybrid enhanced attention module each include: a second channel attention module, a frequency-domain hybrid filtering enhancement module, and a second spatial attention module; The second channel attention module is used to: perform channel attention enhancement on the input feature map to obtain a third enhanced feature map; The frequency domain hybrid filtering enhancement module is used to: perform frequency domain enhancement on the third enhanced feature map to obtain a fourth enhanced feature map; The second spatial attention module is used to: perform spatial attention enhancement on the fourth enhanced feature map to obtain a fifth enhanced feature map; A feature fusion module is used to perform element-by-element point multiplication on the third enhanced feature map and the fifth enhanced feature map to obtain the fused feature map.

6. The street lamp under-lighting detection method based on deep learning according to claim 5, characterized in that: The frequency domain hybrid filtering enhancement module includes: a Fourier transform module, a high-pass filter, an inverse Fourier transform module, a subtraction module and a weighting module in sequence; The Fourier transform module is used to: perform Fourier transform on the third enhanced feature map to obtain a thirteenth feature map; The high-pass filter is used to remove signals below a cutoff frequency in the thirteenth characteristic graph to obtain a fourteenth characteristic graph; The inverse Fourier transform module is used to: perform an inverse Fourier transform on the fourteenth feature map to obtain a fifteenth feature map; The subtraction module is configured to: subtract the fifteenth feature map from the third enhanced feature map to obtain a sixteenth feature map; The weighting module is used to perform weighted summation on the fifteenth feature map and the sixteenth feature map to obtain the fourth enhanced feature map.

7. The street lamp under-lighting detection method based on deep learning according to claim 4, characterized in that: The detection head includes: a first detection head, a second detection head, a third detection head and a processing module; The first detection head is used to: determine a first detection result based on the first fusion feature map; The second detection head is used to: determine a second detection result based on the second fused feature map; The third detection head is used to: determine a third detection result based on the third fusion feature map; The processing module is used to process the first detection result, the second detection result and the third detection result to obtain the target detection result.

8. The street lamp under-lighting detection method based on deep learning according to any one of claims 1 to 7, characterized in that: Before inputting the image of the street lamp to be detected into the street lamp lack-of-light detection network, the method further includes: Acquire multiple streetlight sample images and multiple streetlight sample images, predicted probabilities of streetlights being off in the multiple streetlight sample images, and position information of real frames and predicted frames in the multiple streetlight sample images; For the t-th street light sample image, determining a classification loss value of the t-th street light sample image based on the prediction probability of the t-th street light sample image and the prediction probabilities of the remaining street light sample images, where the remaining street light sample images are images other than the t-th street light sample image among the multiple street light sample images, the t-th street light sample image is any one of the multiple street light sample images, and t is a positive integer less than or equal to the number of the multiple street light sample images; Based on the position information of the real frame and the position information of the predicted frame in the t-th street lamp sample image, determining the intersection-over-union loss value of the t-th street lamp sample image, and the L1 norm of the difference between the coordinates of the four sides of the real frame in the t-th street lamp sample image and the coordinates of the corresponding sides in the predicted frame; Determine the regression loss value of the tth street lamp sample image based on the intersection-over-union loss value of the tth street lamp sample image, the L1 norm, and the width and height of the real box in the tth street lamp sample image; The street lamp under-lighting detection network is trained based on the classification loss values and / or regression loss values of the plurality of street lamp sample images.

9. The street lamp under-lighting detection method based on deep learning according to claim 8, characterized in that: The determining the classification loss value of the tth street lamp sample image based on the prediction probability of the tth street lamp sample image and the prediction probabilities of the remaining street lamp sample images includes: Determining an average value of the predicted probabilities of the remaining street lamp sample images; Performing a weighted summation on the predicted probability of the t-th street lamp sample image and the average of the predicted probabilities of the remaining street lamp sample images to obtain a weighted predicted probability; Based on the weighted prediction probability and the prediction probability of the tth street lamp sample image, a classification loss value of the tth street lamp sample image is determined.

10. A street lamp under-lighting detection device based on deep learning, characterized in that: include: An image acquisition module, used to acquire an image of a street lamp to be detected; a target detection module, configured to input the street lamp image to be detected into a street lamp under-lighting detection network to obtain a target detection result, the target detection result comprising: an identification result of the street lamp in the street lamp image to be detected, and a position of the street lamp in the street lamp image to be detected when the identification result indicates that the street lamp is under-lighting; The street lamp under-lighting detection network includes a backbone network and a detection network; The backbone network sequentially includes a first convolution module and N sequentially connected feature extraction networks, and the N feature extraction networks sequentially include a second convolution module and a multi-scale hole group fusion module, where N is an integer greater than 2; The first convolution module is used to: downsample the street lamp image to be detected to obtain a first feature map; The second convolution module in each of the feature extraction networks is used to: downsample the feature map obtained by the previous layer network to obtain a second feature map; The multi-scale hole grouping fusion module in each feature extraction network is used to: perform feature extraction of different scales on the second feature map obtained by the second convolution module in the same feature extraction network to obtain a third feature map; The detection network is used to determine the target detection result based on the third feature map obtained by at least two of the feature extraction networks.

Citation Information

Patent Citations

  • Abnormal street lamp detection system and method based on YOLOv8 model

    CN119672629A

  • Road traffic sign target detection method based on YOLOv8

    CN119785312A

  • Lane line segmentation method and device, electronic equipment and storage medium

    CN119942128A

  • Traffic lane line detection method and apparatus, and terminal device and readable storage medium

    WO2022126377A1