Transmission tower detection method, electronic equipment, storage medium and program product
By extracting and fusing feature maps of different scales, the problem of low detection accuracy of transmission pole tower category in the prior art is solved, and the detection effect of high accuracy is achieved.
Patent Information
- Application Number
- CN202510186994.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-17
AI Technical Summary
The prior art is difficult to accurately identify the types of transmission pole towers of different scales, resulting in low detection accuracy.
By obtaining the image frame to be detected, the first feature map of multiple scales is extracted, and the feature map of the remaining scales is characterized by fusing the feature map of the scale to obtain the second feature map, and then the detection is performed.
The accuracy of detection of transmission pole tower categories of different scales is achieved, and the feature maps containing features of each scale can be accurately extracted.
Smart Images

Figure CN120164097A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of power grid inspection, and particularly to a method for detecting transmission towers, an electronic device, a storage medium, and a program product. Background Art
[0002] In order to ensure the safe and stable operation of the power grid, different maintenance strategies need to be adopted for transmission towers of different categories (such as horn towers, cathead towers, etc.). Therefore, it is necessary to identify the categories of transmission towers.
[0003] In the related art, feature extraction can be performed on an image frame to obtain a feature map of the image frame, and then the category of the transmission tower in the image frame can be identified according to the feature map to obtain the category of the transmission tower in the image frame.
[0004] However, due to the different positions and perspectives of the cameras during the acquisition of the image frame, the scales of the transmission towers in the image frame are different, and it is impossible to accurately extract a feature map that contains the features of the transmission towers of various scales from the image frame; furthermore, the accuracy of detecting the categories of transmission towers of different scales in the image frame based on the feature map is low. Summary of the Invention
[0005] Embodiments of this application provide a method for detecting transmission towers, an electronic device, a storage medium, and a program product, so as to achieve the effect of high accuracy in detecting the categories of transmission towers of different scales in an image frame.
[0006] In a first aspect, an embodiment of this application provides a method for detecting a transmission tower, including:
[0007] Obtain an image frame to be detected, where the image frame to be detected is an image including a transmission tower;
[0008] Extract first feature maps of multiple scales of the image frame to be detected;
[0009] For each scale of the first feature maps, fuse the first feature maps of the remaining scales with the first feature maps of this scale to obtain second feature maps of this scale;
[0010] Detect each scale of the second feature maps to obtain a detection result, where the detection result includes the category of the transmission tower.
[0011] In a possible implementation manner, the extracting first feature maps of multiple scales of the image frame to be detected includes:
[0012] Input the image frame to be detected into a backbone network unit in a target detection model for feature extraction, and respectively obtain the first feature maps output by each of the multiple feature extraction layers cascaded by the backbone network unit to obtain the first feature maps of multiple scales.
[0013] In a possible implementation, for each first feature map of each scale, fusing the first feature maps of the remaining scales with the first feature map of this scale to obtain the second feature map of this scale includes:
[0014] For any first feature map of the remaining scales, determining a target sampling module corresponding to the remaining scale; processing the first feature map of the remaining scale by using the target sampling module to obtain the target first feature map of the remaining scale; the resolution of the target first feature map is the same as the resolution of the first feature map of this scale;
[0015] Fusing the target first feature maps corresponding to the remaining scales respectively and the first feature map of this scale to obtain the second feature map of this scale.
[0016] In a possible implementation, the target sampling module includes an upsampling module for upsampling or a downsampling module for downsampling.
[0017] In a possible implementation, for each first feature map of each scale, fusing the first feature maps of the remaining scales with the first feature map of this scale to obtain the second feature map of this scale includes:
[0018] Inputting the first feature maps of each scale into the neck network unit in the target detection model, and the neck network unit fuses the first feature map of each scale with the first feature maps of the remaining scales to obtain second feature maps of multiple scales.
[0019] In a possible implementation, the neck network unit includes convolution streams corresponding to multiple scales respectively; wherein, each convolution stream includes at least one cascaded convolution layer; wherein, the resolution of each first feature map remains unchanged after the convolution operation of any convolution layer;
[0020] And the neck network unit fusing the first feature map of each scale with the first feature maps of the remaining scales includes:
[0021] For each scale, determining the target resolution of the first feature map corresponding to this scale, the first target convolution layer in the convolution stream corresponding to this scale, and the transition first feature map corresponding to this convolution stream to be input into the first target convolution layer, and determining the second target convolution layer in other convolution streams; wherein, the transition first feature map is the output of any convolution layer in the convolution stream;
[0022] The output of the second target convolutional layer in other convolutional streams is processed by the target sampling module and then input into the first target convolutional layer, and the corresponding transition first feature map of this convolutional stream is input into the first target convolutional layer, and the first target convolutional layer outputs the fused intermediate first feature map corresponding to this scale.
[0023] In a possible implementation manner, the neck network unit includes deep feature extraction modules corresponding to multiple scales respectively;
[0024] For the first feature map of each scale, fusing the first feature maps of the remaining scales with the first feature map of this scale to obtain the second feature map of this scale, includes:
[0025] In the order from large to small scale, the second feature maps of each scale are determined step by step based on the following first operation:
[0026] Obtain the first convolution result obtained by convolving the first feature map of this scale by the convolutional stream of this scale, the second convolution results obtained by convolving the first feature maps of each scale smaller than this scale by their respective convolutional streams, and the candidate second feature maps output by the deep feature extraction modules of other scales larger than this scale;
[0027] Input the first convolution result, each of the second convolution results, and the candidate second feature maps into the deep feature extraction module of this scale, and the deep feature extraction module fuses the first convolution result, each of the second convolution results, and each of the candidate second feature maps to obtain the second feature map of this scale;
[0028] Take the next scale smaller than the current scale as the updated current scale, and repeat the execution of the first operation.
[0029] In a possible implementation manner, detecting each scale of the second feature map to obtain a detection result, includes:
[0030] Input each scale of the second feature map into the head network unit in the target detection model, and the head network unit outputs the detection result.
[0031] In a possible implementation manner, the head network unit includes detection modules corresponding to multiple scales respectively, and each scale of the detection module is used to perform object detection on the second feature map of this scale.
[0032] In a possible implementation manner, the image frame to be detected includes two continuously captured image frames; the method further includes:
[0033] Compare the number and category of transmission towers detected in the front and rear two image frames;
[0034] In response to the comparison result meeting a preset condition, output the categories of each transmission tower in the subsequent image frame.
[0035] In a second aspect, an embodiment of the present application provides a detection device for transmission towers, including:
[0036] An acquisition module, configured to acquire an image frame to be detected, where the image frame to be detected is an image including a transmission tower;
[0037] An extraction module, configured to extract first feature maps of multiple scales of the image frame to be detected;
[0038] A fusion module, configured to perform feature fusion on the first feature maps of the remaining scales with the first feature map of each scale to obtain a second feature map of each scale;
[0039] A detection module, configured to detect each scale of the second feature map to obtain a detection result, where the detection result includes the category of the transmission tower.
[0040] In a possible implementation manner, the extraction module is specifically configured to:
[0041] Input the image frame to be detected into a backbone network unit in a target detection model for feature extraction, and respectively obtain the first feature maps output by each of the multiple feature extraction layers cascaded by the backbone network unit to obtain the first feature maps of multiple scales.
[0042] In a possible implementation manner, the fusion module is specifically configured to:
[0043] For any first feature map of the remaining scales, determine a target sampling module corresponding to the remaining scale; use the target sampling module to process the first feature map of the remaining scale to obtain a target first feature map of the remaining scale; the resolution of the target first feature map is the same as the resolution of the first feature map of this scale;
[0044] Fuse the target first feature maps corresponding to the respective remaining scales and the first feature map of this scale to obtain the second feature map of this scale.
[0045] In a possible implementation manner, the target sampling module includes an upsampling module for upsampling or a downsampling module for downsampling.
[0046] In a possible implementation manner, the fusion module is specifically configured to:
[0047] Input the first feature maps of each scale into the neck network unit in the object detection model. The neck network unit fuses the first feature maps of each scale with the first feature maps of the remaining scales to obtain second feature maps of multiple scales.
[0048] In a possible implementation manner, the neck network unit includes convolution streams corresponding to multiple scales respectively; wherein, each convolution stream includes at least one cascaded convolution layer; wherein, the resolution of each first feature map remains unchanged after the convolution operation of any convolution layer;
[0049] And for "the neck network unit fuses the first feature maps of each scale with the first feature maps of the remaining scales" in the fusion module, it specifically is used for:
[0050] For each scale, determine the target resolution of the first feature map corresponding to this scale, the first target convolution layer in the convolution stream corresponding to this scale, and the transitional first feature map corresponding to this convolution stream to be input to the first target convolution layer, and determine the second target convolution layer in other convolution streams; wherein, the transitional first feature map is the output of any convolution layer in the convolution stream;
[0051] Input the output of the second target convolution layer in other convolution streams into the first target convolution layer after being processed by the target sampling module, and input the transitional first feature map corresponding to this convolution stream into the first target convolution layer. The first target convolution layer outputs the fused intermediate first feature map corresponding to this scale.
[0052] In a possible implementation manner, the neck network unit includes deep feature extraction modules corresponding to multiple scales respectively.
[0053] In a possible implementation manner, the fusion module is specifically used for:
[0054] In the order from large to small according to the scale, determine the second feature maps of each scale step by step based on the following first operation:
[0055] Obtain the first convolution result obtained by convolving the first feature map of the current scale by the convolution stream of the current scale, the second convolution results obtained by convolving the first feature maps of the scales smaller than the current scale by the convolution streams of their respective scales, and the candidate second feature maps output by the deep feature extraction modules of the other scales larger than the current scale;
[0056] Input the first convolution result, each of the second convolution results, and the candidate second feature maps into the deep feature extraction module of this scale. The deep feature extraction module fuses the first convolution result, each of the second convolution results, and each of the candidate second feature maps to obtain the second feature map of the current scale;
[0057] Take the next scale smaller than the current scale as the updated current scale, and repeatedly execute the first operation.
[0058] In a possible implementation manner, the detection module is specifically configured to:
[0059] Input the second feature map of each scale into the head network unit in the target detection model, and output the detection result by the head network unit.
[0060] In a possible implementation manner, the head network unit includes detection modules corresponding to multiple scales respectively, and the detection module of each scale is used to perform target detection on the second feature map of that scale.
[0061] In a possible implementation manner, the image frame to be detected includes two continuously captured image frames; the device further includes:
[0062] Compare the number and category of transmission towers detected in the front and rear two image frames respectively;
[0063] In response to the comparison result meeting a preset condition, output the category of each transmission tower in the latter image frame.
[0064] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor;
[0065] The memory stores computer execution instructions;
[0066] The processor executes the computer execution instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementation manners of the first aspect.
[0067] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer execution instructions are stored, and when the computer execution instructions are executed by a processor, they are used to implement the above first aspect and / or various possible implementation manners of the first aspect.
[0068] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the above first aspect and / or various possible implementation manners of the first aspect.
[0069] The detection method, electronic device, storage medium, and program product for transmission towers provided by the embodiments of the present application first extract first feature maps of multiple scales from the acquired image frames to be detected; secondly, for each scale of the first feature maps, the features in the first feature maps of the remaining scales are fused into the first feature maps of this scale to obtain the second feature maps of this scale; finally, detections are performed based on the features in the second feature maps of each scale to obtain the categories of the transmission towers in the image frames to be detected. It is realized that in the second feature maps of each scale, the features and resolution in the first feature maps of this scale are retained, and the features in the first feature maps of the remaining scales are also fused. The features of transmission towers of various scales can be accurately extracted from the image frames to be detected, and further, the accuracy of detecting the categories of transmission towers of different scales based on the features in the second feature maps of each scale is high. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0071] Figure 1 It is a schematic flowchart of a detection method for a transmission tower provided by the present application;
[0072] Figure 2 It is a schematic flowchart of another detection method for a transmission tower provided by the present application;
[0073] Figure 3 It is a schematic diagram of the network structure of the object detection model provided by the present application;
[0074] Figure 4 It is a schematic diagram of the structure of a detection device for a transmission tower provided by the present application;
[0075] Figure 5 It is a schematic diagram of the structure of an electronic device provided by the present application.
[0076] Through the above-mentioned accompanying drawings, the specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and text descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0077] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0078] First, the nouns involved in the present application are explained:
[0079] Transmission tower: It refers to a structure used to support and erect transmission lines in a power transmission network.
[0080] For different types of transmission towers (such as horn towers, cathead towers, standard iron towers, etc.), different maintenance strategies need to be adopted. For example, different defect detection algorithms are used for defect detection of different types of transmission towers. Therefore, in order to ensure the safe and stable operation of the power grid, it is necessary to identify the types of transmission towers.
[0081] In one example, feature extraction of the transmission tower can be performed on the image frame to obtain a feature map of the image frame, and then the type of the transmission tower in the image frame can be identified based on the feature map to obtain the type of the transmission tower in the image frame.
[0082] However, due to the different camera positions and perspectives during image frame acquisition, the scales of the transmission towers in the image frame are different (for example, when a drone flies low to collect video, the scale of the transmission tower in the image frame is large, but when the drone flies high to collect video, the scale of the transmission tower in the image frame is small). It is impossible to accurately extract a feature map containing the features of transmission towers of various scales from the image frame; furthermore, the accuracy of detecting the types of transmission towers of different scales in the image frame based on the feature map is low.
[0083] The present application provides a detection method, an electronic device, a storage medium, and a program product for transmission towers to solve the above technical problems.
[0084] The technical solutions of the present application and how the technical solutions of the present application solve the above technical problems will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.
[0085] Figure 1 It is a schematic flowchart of a detection method for transmission towers provided by the present application, as Figure 1 shown, the method includes:
[0086] S101. Obtain an image frame to be detected, where the image frame to be detected is an image including a transmission tower.
[0087] Exemplarily, the execution subject of this embodiment can be any device such as a terminal device, other electronic devices / computer devices, a server, a distributed system, and other devices that can implement the solution of this application, and there is no limitation thereto. Among them, the terminal device can include, but is not limited to, any form such as a drone, a robot, a smart phone, a tablet computer, a mobile monitoring vehicle, etc.; the server can be an independent server or a server cluster, such as can include, but is not limited to, any form such as a cloud server, a distributed server, a blockchain server, etc.
[0088] This embodiment will be introduced with the execution subject being a terminal device.
[0089] Exemplarily, the image frame to be detected is, for example, an image frame obtained by sampling a video collected by a drone for inspecting a transmission tower based on a preset rule, or an image frame obtained by sampling a video collected by a ground vehicle for inspecting a transmission tower based on a preset rule, or an image frame obtained by other means, etc. The method for obtaining the image frame to be detected in this embodiment is not limited, as long as the obtained image frame to be detected includes a transmission tower. Among them, the preset rule can be determined according to experience, and can include, but is not limited to, uniform sampling, random sampling, key frame sampling, etc., and there is no limitation thereto in this embodiment.
[0090] S102. Extract first feature maps of multiple scales of the image frame to be detected.
[0091] Exemplarily, after the terminal device obtains the image frame to be detected (assuming a resolution of 640 pixels × 640 pixels), it can perform feature extraction on the image frame to be detected to obtain first feature maps of multiple scales of the image frame to be detected.
[0092] Among them, the first feature maps represent the basic image features in the image frame to be detected. The first feature maps of different scales correspond to different resolutions.
[0093] The high-resolution first feature maps include more image detail features in the image frame to be detected, such as edges, textures, colors, etc. The low-resolution first feature maps include more global structures and context information, such as the overall shape and category features of the transmission tower, etc. The resolution of the medium-resolution first feature maps is between the high resolution and the low resolution, and the medium-resolution first feature maps include both certain local features and certain global context information. The number of high-resolution first feature maps, medium-resolution first feature maps, and low-resolution first feature maps can be determined according to requirements, and there is no limitation thereto in this embodiment.
[0094] In one example, it is assumed that after the terminal device obtains an image frame to be detected with a resolution of 640 pixels × 640 pixels, multiple-scale first feature maps are extracted from the image frame to be detected. Among them, the resolutions of the multiple-scale first feature maps can be, for example, 80 pixels × 80 pixels (high resolution), 40 pixels × 40 pixels (medium resolution), and 20 pixels × 20 pixels (low resolution) respectively. Among them, the first feature map with a resolution of 80 pixels × 80 pixels can include, for example, the outline, material texture, color, etc. of the transmission tower in the image frame to be detected; the first feature map with a resolution of 40 pixels × 40 pixels can include, for example, the local shape information of the transmission tower in the image frame to be detected, such as the shapes of some components of the transmission tower (such as insulators and cross arms); the first feature map with a resolution of 20 pixels × 20 pixels can include, for example, the overall form of the transmission tower in the image frame to be detected and the relationship between its components.
[0095] S103. For each scale of the first feature map, fuse the first feature maps of the remaining scales with the first feature map of this scale to obtain the second feature map of this scale.
[0096] Exemplarily, after obtaining the first feature maps of multiple scales of the image frame to be detected, for each scale of the first feature map, the features in the first feature maps of the remaining scales can be fused into the first feature map of this scale to obtain the second feature map of this scale. Among them, the second feature map represents the cross-resolution image basic features and image semantic features in the image frame to be detected. Different second feature maps correspond to different resolutions.
[0097] The fusion method in each embodiment of the present application can be to perform an addition operation on the corresponding elements at the same position of each feature map or to splice each feature map along the channel dimension. This embodiment does not limit the fusion method, as long as the features in each feature map can be fused.
[0098] In one example, still referring to the example in step S102, it is assumed that the resolutions of the obtained first feature maps of multiple scales are 80 pixels × 80 pixels, 40 pixels × 40 pixels, and 20 pixels × 20 pixels respectively. Then, for the first feature map with a resolution of 80 pixels × 80 pixels, the features in the first feature map with a resolution of 40 pixels × 40 pixels and the features in the first feature map with a resolution of 20 pixels × 20 pixels are fused into the first feature map with a resolution of 80 pixels × 80 pixels to obtain the second feature map with a resolution of 80 pixels × 80 pixels.
[0099] For the first feature map with a resolution of 40 pixels × 40 pixels, the features in the first feature map with a resolution of 80 pixels × 80 pixels and the features in the first feature map with a resolution of 20 pixels × 20 pixels are fused into the first feature map with a resolution of 40 pixels × 40 pixels to obtain a second feature map with a resolution of 40 pixels × 40 pixels.
[0100] For the first feature map with a resolution of 20 pixels × 20 pixels, the features in the first feature map with a resolution of 80 pixels × 80 pixels and the features in the first feature map with a resolution of 40 pixels × 40 pixels are fused into the first feature map with a resolution of 20 pixels × 20 pixels to obtain a second feature map with a resolution of 20 pixels × 20 pixels.
[0101] S104. Detect each scale of the second feature map to obtain a detection result, where the detection result includes the category of the transmission tower.
[0102] Exemplarily, after obtaining the second feature maps of multiple scales, detection can be performed based on the features in each scale of the second feature map to obtain a detection result, where the detection result includes the category of the transmission tower in the image frame to be detected. Among them, the category of the transmission tower includes, for example, but is not limited to, horn towers, cathead towers, standard iron towers, etc.
[0103] In one example, still referring to the example in step S103, assuming that the resolutions of the second feature maps of multiple scales obtained are 80 pixels × 80 pixels, 40 pixels × 40 pixels, and 20 pixels × 20 pixels respectively, then detection is performed according to the features in the second feature map with a resolution of 80 pixels × 80 pixels, the features in the second feature map with a resolution of 40 pixels × 40 pixels, and the features in the second feature map with a resolution of 20 pixels × 20 pixels to obtain the category of the transmission tower in the image frame to be detected.
[0104] The detection method of the transmission tower provided by the embodiment of the present application first extracts the first feature maps of multiple scales from the obtained image frame to be detected; secondly, for each scale of the first feature map, the features in the first feature maps of the other scales are fused into the first feature map of this scale to obtain the second feature map of this scale; finally, detection is performed according to the features in the second feature maps of each scale to obtain the category of the transmission tower in the image frame to be detected, achieving that in the second feature map of each scale, both the features and resolution of the first feature map of this scale are retained, and the features in the first feature maps of the other scales are fused, and further, the accuracy of detecting the categories of transmission towers of different scales based on the features in the second feature maps of each scale is high.
[0105] In some embodiments, the above step S103 includes:
[0106] First, for any first feature map of the remaining scales, determine the target sampling module corresponding to the remaining scale; use the target sampling module to process the first feature map of the remaining scale to obtain the target first feature map of the remaining scale; the resolution of the target first feature map is the same as the resolution of the first feature map of this scale.
[0107] Second, fuse the target first feature map corresponding to each remaining scale and the first feature map of this scale to obtain the second feature map of this scale.
[0108] In one example, the target sampling module includes an upsampling module for upsampling or a downsampling module for downsampling.
[0109] Exemplarily, for any first feature map of the remaining scales, determining the target sampling module corresponding to the remaining scale can be understood as, for each scale of the first feature map, for any first feature map of the remaining scales, if the resolution of the first feature map of the remaining scale is lower than the resolution of the first feature map of this scale, then determine the target sampling module corresponding to the remaining scale as the upsampling module; if the resolution of the first feature map of the remaining scale is higher than the resolution of the first feature map of this scale, then determine the target sampling module corresponding to the remaining scale as the downsampling module.
[0110] The upsampling module can perform upsampling, for example, by nearest neighbor interpolation or bilinear interpolation, etc., and align the number of channels by connecting a 1×1 convolution. The downsampling module can perform downsampling, for example, by one or more 3×3 convolutions with a stride of 2, or by a pooling layer for downsampling, etc. The implementation manner of the upsampling module or the downsampling module in this embodiment is not limited, as long as the resolution of the target first feature map can be made the same as the resolution of the first feature map of the corresponding scale.
[0111] In one example, still referring to the example in step S102, for the first feature map with a resolution of 40 pixels × 40 pixels, for the first feature map with a resolution of 80 pixels × 80 pixels, since the resolution of 80 pixels × 80 pixels is higher than the resolution of 40 pixels × 40 pixels, then determine the target sampling module corresponding to the resolution of 80 pixels × 80 pixels as the downsampling module. Then, the first feature map with a resolution of 80 pixels × 80 pixels can be input into the downsampling module for downsampling to obtain the target first feature map with a resolution of 40 pixels × 40 pixels corresponding to the resolution scale of 80 pixels × 80 pixels.
[0112] For the first feature map with a resolution of 20 pixels × 20 pixels, since the resolution of 20 pixels × 20 pixels is lower than that of 40 pixels × 40 pixels, the target sampling module corresponding to the resolution of 20 pixels × 20 pixels is determined to be the upsampling module. Then, the first feature map with a resolution of 20 pixels × 20 pixels can be input into the upsampling module for upsampling to obtain the target first feature map with a resolution of 40 pixels × 40 pixels corresponding to the resolution scale of 20 pixels × 20 pixels.
[0113] Next, the features in the first feature map with a resolution of 40 pixels × 40 pixels, the features in the target first feature map with a resolution of 40 pixels × 40 pixels corresponding to the resolution scale of 80 pixels × 80 pixels, and the features in the target first feature map with a resolution of 40 pixels × 40 pixels corresponding to the resolution scale of 20 pixels × 20 pixels are fused to obtain the second feature map with a resolution of 40 pixels × 40 pixels.
[0114] Similarly, the implementation manner of fusing the first feature maps of multiple scales based on the target sampling module to obtain the second feature map with a resolution of 80 pixels × 80 pixels or the second feature map with a resolution of 20 pixels × 20 pixels is similar to the implementation manner of fusing the first feature maps of multiple scales based on the target sampling module to obtain the second feature map with a resolution of 40 pixels × 40 pixels introduced above. For details, reference can be made to the above introduction and will not be elaborated here.
[0115] By determining the target sampling module for any other scale of the first feature map, the target sampling module can be used to perform upsampling or downsampling on the first feature map of the other scale, and then the resolution of the first feature map of the other scale can be adjusted to the resolution of the first feature map of this scale, so as to realize fusing each first feature map of the other scale into the first feature map of this scale to obtain the second feature map of this scale.
[0116] Figure 2 It is a schematic flowchart of another method for detecting transmission towers provided by this application. As Figure 2 shown, based on the Figure 1 embodiment, the method for detecting transmission towers is described in detail. The method includes:
[0117] S201. Obtain an image frame to be detected, where the image frame to be detected is an image including a transmission tower.
[0118] Exemplarily, the execution subject of this embodiment can be any device among a terminal device, other electronic devices / computer devices, a server, a distributed system, and other devices that can implement the solution of this application, and no limitation is imposed thereon.
[0119] This embodiment will be described with the execution entity being a terminal device.
[0120] Exemplarily, the implementation manners of step S201 and step S101 are similar. For details, reference can be made to the description in step S101, which will not be elaborated here.
[0121] S202: Input the image frame to be detected into the backbone network unit in the target detection model for feature extraction, and respectively obtain the first feature maps output by each of the multiple feature extraction layers cascaded in the backbone network unit, so as to obtain the first feature maps of multiple scales.
[0122] Exemplarily, the implementation manners of step S202 and step S102 are similar. For details, reference can be made to the description in step S102, which will not be elaborated here.
[0123] Specifically, it can be combined with Figure 3 to understand the target detection model. Figure 3 is a schematic diagram of the network structure of the target detection model provided in this application. As Figure 3 shown, the target detection model includes a backbone network unit 32, a neck network unit 33 connected to the backbone network unit 32, and a head network unit 34 connected to the neck network unit 33. The image frame 31 can be understood as the image frame to be detected, for example.
[0124] Among them, in the backbone network unit 32, there are 3 cascaded feature extraction layers. In the first cascaded feature extraction layer, there is an initial feature extraction module 3201, a depth feature extraction module 3202 connected to the initial feature extraction module 3201 (the depth feature extraction module is used for deep feature extraction, using state space model technology to capture features at different levels and merge and process time series information, etc.), a visual cue fusion module 3203 connected to the depth feature extraction module 3202 (the visual cue fusion module is used for feature fusion and optimization), and a depth feature extraction module 3204 connected to the visual cue fusion module 3203. The first cascaded feature extraction layer is used to obtain the output first feature map with high resolution.
[0125] In the second cascaded feature extraction layer, there is a visual cue fusion module 3205 connected to the depth feature extraction module 3204 and a depth feature extraction module 3206 connected to the visual cue fusion module 3205. The second cascaded feature extraction layer is used to obtain the output first feature map with medium resolution.
[0126] The third cascaded feature extraction layer includes a visual cue fusion module 3207 connected to the depth feature extraction module 3206, a depth feature extraction module 3208 connected to the visual cue fusion module 3207, and a multi-scale fusion module 3209 (e.g., a Fast Spatial Pyramid Pooling (SPPF) module) connected to the depth feature extraction module 3208. The third cascaded feature extraction layer is used to obtain the first feature map with a low resolution as the output.
[0127] In one example, as Figure 3 shown, after the terminal device obtains the image frame 31, it inputs the image frame 31 into Figure 3 the first cascaded feature extraction layer in the backbone network unit 32 of the object detection model in
[0128] for feature extraction to obtain the first feature map P3 with a high resolution (assuming the first feature map with a resolution of 80 pixels × 80 pixels).
[0129] Then, it inputs the first feature map P3 with a high resolution into the second cascaded feature extraction layer in the backbone network unit 32 for feature extraction to obtain the first feature map P4 with a medium resolution (assuming the first feature map with a resolution of 40 pixels × 40 pixels).
[0130] Among them, the meanings of the first feature map, the first feature map with a high resolution, the first feature map with a medium resolution, and the first feature map with a low resolution have been introduced in step S102. For details, refer to the description in step S102 and will not be elaborated here.
[0131] S203. Input the first feature maps of each scale into the neck network unit of the object detection model, and the neck network unit fuses the first feature map of each scale with the first feature maps of the other scales to obtain the second feature maps of multiple scales.
[0132] In one example, the neck network unit includes convolutional streams corresponding to multiple scales respectively; among them, each convolutional stream includes at least one cascaded convolutional layer; and the resolution of each first feature map remains unchanged after the convolutional operation of any convolutional layer.
[0133] In some embodiments, the first feature maps of each scale are input into the neck network unit in the object detection model. For each scale, the target resolution of the first feature map corresponding to this scale, the first target convolutional layer in the convolutional stream corresponding to this scale, and the transitional first feature map corresponding to this convolutional stream to be input into the first target convolutional layer are determined, and the second target convolutional layer is determined in other convolutional streams; wherein, the transitional first feature map is the output of any convolutional layer in the convolutional stream.
[0134] The output of the second target convolutional module in other convolutional streams is processed by the target sampling module and then input into the first target convolutional module, and the transitional first feature map corresponding to this convolutional stream is input into the first target convolutional module, and the first target convolutional module outputs the fused intermediate first feature map corresponding to this scale.
[0135] In one example, the neck network unit includes deep feature extraction modules corresponding to multiple scales respectively.
[0136] In some embodiments, step S203 includes:
[0137] In the order from the largest scale to the smallest scale, the second feature maps of each scale are determined step by step based on the following first operation:
[0138] Obtain the first convolution result obtained by convolving the first feature map of this scale by the convolutional stream of this scale, the second convolution results obtained by convolving the first feature maps of the scales smaller than this scale by the convolutional streams of their respective scales, and the candidate second feature maps output by the deep feature extraction modules of other scales larger than this scale;
[0139] Input the first convolution result, each second convolution result, and the candidate second feature maps into the deep feature extraction module of this scale, and the deep feature extraction module fuses the first convolution result, each second convolution result, and each candidate second feature map to obtain the second feature map of this scale;
[0140] Take the next scale smaller than the current scale as the updated current scale, and repeat the first operation.
[0141] Exemplarily, in order to enhance the detailed features in the feature maps of each scale, for example, the feature maps of this scale can be first enhanced through one or more convolutional layers, and then the features in the feature maps of other scales are fused in the feature maps of this scale.
[0142] Since the features become gradually more abstract and semantic from the features in the high-resolution feature maps to the features in the low-resolution feature maps, for each scale of feature maps, the features in the feature maps of the remaining scales can be gradually fused layer by layer into the feature maps of this scale in the order from large to small. Repeatedly and gradually fusing the features of the remaining scales in this way can enable the object detection model to gradually construct more complex feature combinations. Each fusion can be regarded as further optimizing the feature representation on the basis of the previous one. As the number of fusions increases, the correlation between different-scale features is gradually strengthened, and finally a feature map including more comprehensive and detailed features is output.
[0143] In one example, it can be combined with Figure 3 , and the example in step S202 to understand the neck network unit in the object detection model. As shown in Figure 3 , in the neck network unit 33 of the object detection model in Figure 3 , it includes a fusion stream corresponding to the high resolution, a fusion stream corresponding to the medium resolution, and a fusion stream corresponding to the low resolution.
[0144] Among them, the fusion stream corresponding to the high resolution includes a convolutional stream corresponding to the high resolution and a deep feature extraction module 3305 connected to the convolutional layer 3304, the convolutional layer 3308, and the convolutional layer 3311 respectively. The convolutional stream corresponding to the high resolution includes the convolutional layer 3301, the convolutional layer 3302 connected to the convolutional layer 3301, the convolutional layer 3303 connected to the convolutional layer 3302 and the convolutional layer 3306 respectively, and the convolutional layer 3304 connected to the convolutional layer 3303, the convolutional layer 3307, and the convolutional layer 3310 respectively. The fusion stream corresponding to the high resolution is used to fuse the features in the first feature map P4 of the medium resolution and the features in the first feature map P5 of the low resolution into the first feature map P3 of the high resolution to obtain the second feature map P3' of the high resolution. Among them, the resolutions of the second feature map P3' of the high resolution and the first feature map P3 of the high resolution are both 80 pixels × 80 pixels.
[0145] The fusion stream corresponding to the medium resolution includes the convolutional stream corresponding to the medium resolution, and the deep feature extraction module 3309 connected to the convolutional layer 3308, the convolutional layer 3311, and the deep feature extraction module 3305 respectively. The convolutional stream corresponding to the medium resolution includes the convolutional layer 3306, the convolutional layer 3307 connected to the convolutional layer 3302 and the convolutional layer 3306 respectively, and the convolutional layer 3308 connected to the convolutional layer 3303, the convolutional layer 3307, and the convolutional layer 3310 respectively. Among them, the fusion stream corresponding to the medium resolution is used to fuse the features in the high-resolution first feature map P3 and the features in the low-resolution first feature map P5 into the medium-resolution first feature map P4 to obtain the medium-resolution second feature map P4'. Among them, the resolutions of the medium-resolution second feature map P4' and the medium-resolution first feature map P4 are both 40 pixels × 40 pixels.
[0146] The fusion stream corresponding to the low resolution includes the convolutional stream corresponding to the low resolution, and the deep feature extraction module 3312 connected to the convolutional layer 3311, the deep feature extraction module 3305, and the deep feature extraction module 3309 respectively. The convolutional stream corresponding to the low resolution includes the convolutional layer 3310, and the convolutional layer 3311 connected to the convolutional layer 3303, the convolutional layer 3307, and the convolutional layer 3310 respectively. Among them, the fusion stream corresponding to the low resolution is used to fuse the features in the high-resolution first feature map P3 and the features in the medium-resolution first feature map P4 into the low-resolution first feature map P5 to obtain the low-resolution second feature map P5'. Among them, the resolutions of the low-resolution second feature map P5' and the low-resolution first feature map P5 are both 20 pixels × 20 pixels.
[0147] Among them, the convolutional layer 3301, the convolutional layer 3306, and the convolutional layer 3310 can be, for example, 1×1 convolutional layers. The convolutional layer is used to perform a convolutional operation on the input data through a series of filters to extract the features in the input data.
[0148] In one example, as Figure 3 shown, after obtaining the high-resolution first feature map P3, the medium-resolution first feature map P4, and the low-resolution first feature map P5, the first feature maps of the above three scales can be input into Figure 3 the neck network unit 33 in the object detection model in
[0149] for cross-scale feature fusion. Specifically, for the high-resolution scale, the resolution 80×80 of the high-resolution first feature map P3 can be determined as the target resolution corresponding to the high-resolution scale. For example, it can also be understood that the feature maps output in the fusion stream corresponding to the high resolution are all of the resolution 80 pixels × 80 pixels. For example, it can be determined that the first target convolutional layer corresponding to the high-resolution scale includes the convolutional layer 3303 and the convolutional layer 3304.
[0150] For the first target convolutional layer 3303 corresponding to the high-resolution scale, the feature map output by the convolutional layer 3302 can be determined as the transitional first feature map of the convolutional stream corresponding to the high resolution of the first target convolutional layer 3303 In the convolutional stream corresponding to the medium resolution, the second target convolutional layer corresponding to the first target convolutional layer 3303 can be determined to be the convolutional layer 3306.
[0151] For the first target convolutional layer 3304 corresponding to the high-resolution scale, the feature map output by the convolutional layer 3303 can be determined as the transitional first feature map of the convolutional stream corresponding to the high resolution of the first target convolutional layer 3304 In the convolutional stream corresponding to the medium resolution, the second target convolutional layer corresponding to the first target convolutional layer 3304 can be determined to be the convolutional layer 3307; in the convolutional stream corresponding to the low resolution, the second target convolutional layer corresponding to the first target convolutional layer 3304 can be determined to be the convolutional layer 3310.
[0152] For the medium-resolution scale, the resolution 40×40 of the first feature map P4 at medium resolution can be determined as the target resolution corresponding to the medium-resolution scale. For example, it can also be understood that the feature maps output in the fusion stream corresponding to the medium resolution are all at a resolution of 40 pixels × 40 pixels; for example, it can be determined that the first target convolutional layers corresponding to the medium-resolution scale include the convolutional layer 3307 and the convolutional layer 3308.
[0153] For the first target convolutional layer 3307 corresponding to the medium-resolution scale, the feature map output by the convolutional layer 3306 can be determined as the transitional first feature map of the convolutional stream corresponding to the medium resolution of the first target convolutional layer 3307 In the convolutional stream corresponding to the high resolution, the second target convolutional layer corresponding to the first target convolutional layer 3307 can be determined to be the convolutional layer 3302.
[0154] For the first target convolutional layer 3308 corresponding to the medium-resolution scale, the feature map output by the convolutional layer 3307 can be determined as the transitional first feature map of the convolutional stream corresponding to the medium resolution of the first target convolutional layer 3308 In the convolutional stream corresponding to the high resolution, the second target convolutional layer corresponding to the first target convolutional layer 3308 can be determined to be the convolutional layer 3303; in the convolutional stream corresponding to the low resolution, the second target convolutional layer corresponding to the first target convolutional layer 3308 can be determined to be the convolutional layer 3310.
[0155] For the low-resolution scale, the resolution of 20×20 of the first feature map P5 with low resolution can be determined as the target resolution corresponding to the low-resolution scale. For example, it can also be understood that the feature maps output in the fusion stream corresponding to the low resolution are all at a resolution of 20 pixels × 20 pixels. For example, it can be determined that the first target convolutional layer of the low-resolution scale includes convolutional layer 3311.
[0156] For the first target convolutional layer 3311 corresponding to the low-resolution scale, the feature map output by convolutional layer 3310 can be determined as the transitional first feature map of the convolutional stream corresponding to the low resolution of the first target convolutional layer 3311 In the convolutional stream corresponding to the high resolution, the second target convolutional layer corresponding to the first target convolutional layer 3311 can be determined as convolutional layer 3303; in the convolutional stream corresponding to the medium resolution, the second target convolutional layer corresponding to the first target convolutional layer 3311 can be determined as convolutional layer 3307.
[0157] After determining the target resolution, the first target convolutional layer, the second target convolutional layer corresponding to the first target convolutional layer, and the transitional first feature map corresponding to the first target convolutional layer for each scale, for example, the output after processing the transitional first feature map corresponding to the first target convolutional layer and the output of the second target convolutional layer corresponding to the first target convolutional layer by the target sampling module can be input into the first target convolutional layer for feature fusion to obtain the intermediate first feature map of this scale. Among them, the purpose of the target sampling module is to adjust the resolution of the feature map output by the second target convolutional layer to the target resolution of this scale. The implementation method of the target sampling module is similar to Figure 1 the implementation method of the target sampling module introduced in the embodiment. For details, reference can be made to Figure 1 the description in the embodiment, which will not be elaborated here.
[0158] For example, it can be combined with Figure 3 to understand the process of obtaining the intermediate first feature map. As Figure 3 shown, after obtaining the first feature map P3 with high resolution (assuming the resolution is 80 pixels × 80 pixels), based on Figure 3 the object detection model in, the first feature map P3 with high resolution is input into convolutional layer 3301 for feature enhancement processing to obtain a transitional first feature map with a resolution of 80 pixels × 80 pixels
[0159] The transitional first feature map is input into convolutional layer 3302 for further feature enhancement processing to obtain a transitional first feature map with a resolution of 80 pixels × 80 pixels
[0160] After obtaining the first feature map P4 with medium resolution (assuming the resolution is 40 pixels × 40 pixels), based on Figure 3 the object detection model, the first feature map P4 with medium resolution is input into the convolutional layer 3306 for feature enhancement processing to obtain a transitional first feature map with a resolution of 40 pixels × 40 pixels
[0161] In the convolutional stream corresponding to a resolution of 80 pixels × 80 pixels, the transitional first feature map with a resolution of 40 pixels × 40 pixels output by the second target convolutional layer 3306 corresponding to the first target convolutional layer 3303 in the convolutional stream corresponding to a resolution of 80 pixels × 80 pixels is input into an upsampling module ( Figure 3 not shown in the figure) to upsample the transitional first feature map with a resolution of 40 pixels × 40 pixels to obtain a first transitional first feature map with a resolution of 80 pixels × 80 pixels.
[0162] The features in the first transitional first feature map with a resolution of 80 pixels × 80 pixels and the features in the transitional first feature map are fused to obtain a first fused feature map with a resolution of 80 pixels × 80 pixels.
[0163] The first fused feature map with a resolution of 80 pixels × 80 pixels is input into the first target convolutional layer 3303 in the convolutional stream corresponding to a resolution of 80 pixels × 80 pixels for further feature enhancement and fusion processing to obtain an intermediate first feature map with a resolution of 80 pixels × 80 pixels
[0164] In the convolutional stream corresponding to a resolution of 40 pixels × 40 pixels, the transitional first feature map with a resolution of 80 pixels × 80 pixels output by the second target convolutional layer 3302 corresponding to the first target convolutional layer 3307 in the convolutional stream corresponding to a resolution of 40 pixels × 40 pixels is input into a downsampling module ( Figure 3 not shown in the figure) to downsample the transitional first feature map to obtain a first transitional first feature map with a resolution of 40 pixels × 40 pixels.
[0165] The features in the first transitional first feature map with a resolution of 40 pixels × 40 pixels and the features in the transitional first feature map with a resolution of 40 pixels × 40 pixels are fused to obtain a first fused feature map with a resolution of 40 pixels × 40 pixels.
[0166] The first fused feature map with a resolution of 40 pixels × 40 pixels is input into the first target convolutional layer 3307 corresponding to the convolutional stream with a resolution of 40 pixels × 40 pixels for further feature enhancement and fusion processing, obtaining an intermediate first feature map with a resolution of 40 pixels × 40 pixels
[0167] The low-resolution first feature map P5 is input into the convolutional layer 3310 for feature enhancement processing, obtaining a transitional first feature map with a resolution of 20 pixels × 20 pixels
[0168] In the convolutional stream corresponding to a resolution of 80 pixels × 80 pixels, the transitional first feature map with a resolution of 20 pixels × 20 pixels output by the second target convolutional layer 3310 corresponding to the first target convolutional layer 3304 corresponding to the convolutional stream with a resolution of 80 pixels × 80 pixels is input into the upsampling module ( Figure 3 not shown in the figure) to upsample the transitional first feature map and obtain a second transitional first feature map with a resolution of 80 pixels × 80 pixels
[0169] In the convolutional stream corresponding to a resolution of 80 pixels × 80 pixels, the intermediate first feature map with a resolution of 40 pixels × 40 pixels output by the second target convolutional layer 3307 corresponding to the first target convolutional layer 3304 corresponding to the convolutional stream with a resolution of 80 pixels × 80 pixels is input into the upsampling module ( Figure 3 not shown in the figure) to upsample the intermediate first feature map with a resolution of 40 pixels × 40 pixels and obtain a third transitional first feature map with a resolution of 80 pixels × 80 pixels
[0170] The features in the third transitional first feature map with a resolution of 80 pixels × 80 pixels, the features in the second transitional first feature map with a resolution of 80 pixels × 80 pixels, and the features in the intermediate first feature map with a resolution of 80 pixels × 80 pixels are fused to obtain a fourth transitional first feature map with a resolution of 80 pixels × 80 pixels
[0171] The fourth transitional first feature map with a resolution of 80 pixels × 80 pixels is input into the first target convolutional layer 3304 corresponding to the convolutional stream with a resolution of 80 pixels × 80 pixels for further feature enhancement and fusion processing, obtaining an intermediate first feature map with a resolution of 80 pixels × 80 pixels
[0172] In the convolution stream corresponding to a resolution of 40 pixels × 40 pixels, the intermediate first feature map with a resolution of 80 pixels × 80 pixels output by the second target convolutional layer 3303 corresponding to the first target convolutional layer 3308 in the convolution stream corresponding to a resolution of 40 pixels × 40 pixels is input into a downsampling module ( Figure 3 not shown in the figure), and the intermediate first feature map with a resolution of 80 pixels × 80 pixels is downsampled to obtain a second transitional first feature map with a resolution of 40 pixels × 40 pixels.
[0173] In the convolution stream corresponding to a resolution of 40 pixels × 40 pixels, the transitional first feature map with a resolution of 20 pixels × 20 pixels output by the second target convolutional layer 3310 corresponding to the first target convolutional layer 3308 in the convolution stream corresponding to a resolution of 40 pixels × 40 pixels is input into an upsampling module ( Figure 3 not shown in the figure), and the transitional first feature map with a resolution of 20 pixels × 20 pixels is upsampled to obtain a third transitional first feature map with a resolution of 40 pixels × 40 pixels.
[0174] The features in the second transitional first feature map with a resolution of 40 pixels × 40 pixels, the features in the third transitional first feature map with a resolution of 40 pixels × 40 pixels, and the features in the intermediate first feature map with a resolution of 40 pixels × 40 pixels are fused to obtain a fourth transitional first feature map with a resolution of 40 pixels × 40 pixels.
[0175] The fourth transitional first feature map with a resolution of 40 pixels × 40 pixels is input into the first target convolutional layer 3308 in the convolution stream corresponding to a resolution of 40 pixels × 40 pixels for further feature enhancement and fusion processing to obtain the intermediate first feature map
[0176] In the convolution stream corresponding to a resolution of 20 pixels × 20 pixels, the intermediate first feature map with a resolution of 80 pixels × 80 pixels output by the second target convolutional layer 3303 corresponding to the first target convolutional layer 3311 in the convolution stream corresponding to a resolution of 20 pixels × 20 pixels is input into a downsampling module ( Figure 3 not shown in the figure), and the intermediate first feature map with a resolution of 80 pixels × 80 pixels is downsampled to obtain a first transitional first feature map with a resolution of 20 pixels × 20 pixels.
[0177] In the convolution stream corresponding to a resolution of 20 pixels × 20 pixels, the intermediate first feature map with a resolution of 40 pixels × 40 pixels output by the second target convolutional layer 3307 corresponding to the first target convolutional layer 3311 in the convolution stream with a resolution of 20 pixels × 20 pixels is input into a downsampling module ( Figure 3 not shown in the figure), and the intermediate first feature map with a resolution of 40 pixels × 40 pixels is downsampled to obtain a second transitional first feature map with a resolution of 20 pixels × 20 pixels.
[0178] The features in the first transitional first feature map with a resolution of 20 pixels × 20 pixels, the features in the second transitional first feature map with a resolution of 20 pixels × 20 pixels, and the features in the transitional first feature map with a resolution of 20 pixels × 20 pixels are fused to obtain a third transitional first feature map with a resolution of 20 pixels × 20 pixels.
[0179] The third transitional first feature map with a resolution of 20 pixels × 20 pixels is input into the first target convolutional layer 3311 corresponding to the convolution stream with a resolution of 20 pixels × 20 pixels for further feature enhancement processing, to obtain an intermediate first feature map with a resolution of 20 pixels × 20 pixels
[0180] In some embodiments, for example, the intermediate first feature map with a resolution of 80 pixels × 80 pixels the intermediate first feature map with a resolution of 40 pixels × 40 pixels the intermediate first feature map with a resolution of 20 pixels × 20 pixels are respectively used as the second feature map P3′ with a resolution of 80 pixels × 80 pixels, the second feature map P4′ with a resolution of 40 pixels × 40 pixels, and the second feature map P5′ with a resolution of 20 pixels × 20 pixels.
[0181] In other embodiments, for example, through the deep feature extraction modules 3305, 3309, and 3312 in the neck network unit 33 Figure 3 the intermediate first feature map with a resolution of 80 pixels × 80 pixels the intermediate first feature map with a resolution of 40 pixels × 40 pixels the intermediate first feature map with a resolution of 20 pixels × 20 pixels continue to perform multi-scale feature fusion to obtain a multi-scale second feature map.
[0182] Specifically, for example, first, for the scale of 80 pixels × 80 pixels resolution, obtain the first convolution result of 80 pixels × 80 pixels resolution obtained by convolving the first feature map P3 of 80 pixels × 80 pixels resolution with the convolution stream corresponding to 80 pixels × 80 pixels resolution (that is, the first intermediate feature map of 80 pixels × 80 pixels resolution ), convolve the first feature map P4 of 40 pixels × 40 pixels resolution with the convolution stream corresponding to 40 pixels × 40 pixels resolution, which is smaller than the resolution of 80 pixels × 80 pixels, to obtain the second convolution result of 40 pixels × 40 pixels resolution (that is, the first intermediate feature map of 40 pixels × 40 pixels resolution ), convolve the first feature map P5 of 20 pixels × 20 pixels resolution with the convolution stream corresponding to 20 pixels × 20 pixels resolution, which is smaller than the resolution of 80 pixels × 80 pixels, to obtain the second convolution result of 20 pixels × 20 pixels resolution (that is, the first intermediate feature map of 20 pixels × 20 pixels resolution ).
[0183] Next, input the second convolution result of 40 pixels × 40 pixels resolution into the upsampling module ( Figure 3 not shown in the figure) to upsample the second convolution result R 3 2 of 40 pixels × 40 pixels resolution to obtain the target first feature map of 80 pixels × 80 pixels resolution corresponding to the convolution stream of 40 pixels × 40 pixels resolution.
[0184] Input the second convolution result of 20 pixels × 20 pixels resolution into the upsampling module ( Figure 3 not shown in the figure) to upsample the second convolution result of 20 pixels × 20 pixels resolution to obtain the target first feature map of 80 pixels × 80 pixels resolution corresponding to the convolution stream of 20 pixels × 20 pixels resolution.
[0185] Fuse the features in the target first feature map of 80 pixels × 80 pixels resolution corresponding to the convolution stream of 40 pixels × 40 pixels resolution, the features in the target first feature map of 80 pixels × 80 pixels resolution corresponding to the convolution stream of 20 pixels × 20 pixels resolution, and the features in the first convolution result of 80 pixels × 80 pixels resolution to obtain the high-resolution fusion feature map of 80 pixels × 80 pixels resolution.
[0186] The high-resolution fused feature map with a resolution of 80 pixels × 80 pixels is input into the deep feature extraction module 3305 for further feature enhancement and fusion processing to obtain a high-resolution second feature map P3'. Among them, the resolution of the high-resolution second feature map P3' is 80 pixels × 80 pixels.
[0187] Since in the fusion stream corresponding to the high resolution, each feature map is of high resolution, and in the high-resolution second feature map, the features in the medium-resolution feature map and the features in the low-resolution feature map are fused multiple times. Therefore, in the high-resolution second feature map, both high resolution is maintained and the fusion of cross-resolution information interaction features is achieved, and thus the features in the high-resolution second feature map are more accurate.
[0188] After obtaining the high-resolution second feature map P3', for example, for the scale of the resolution of 40 pixels × 40 pixels, the first convolution result with a resolution of 40 pixels × 40 pixels obtained by convolving the first feature map P4 with a resolution of 40 pixels × 40 pixels using the convolution stream corresponding to the resolution of 40 pixels × 40 pixels (that is, the intermediate first feature map with a resolution of 40 pixels × 40 pixels ), the second convolution result with a resolution of 20 pixels × 20 pixels obtained by convolving the first feature map P5 with a resolution of 20 pixels × 20 pixels using the convolution stream corresponding to the resolution of 20 pixels × 20 pixels with a size smaller than the resolution of 40 pixels × 40 pixels (that is, the intermediate first feature map with a resolution of 20 pixels × 20 pixels ), and the candidate second feature map P3' (that is, the high-resolution second feature map P3') output by the deep feature extraction module 3305 in the fusion stream corresponding to the resolution of 80 pixels × 80 pixels with a size larger than the resolution of 40 pixels × 40 pixels.
[0189] Then, the candidate second feature map P3' is input into a downsampling module ( Figure 3 not shown in the figure) for downsampling the candidate second feature map P3' to obtain a target first feature map with a resolution of 40 pixels × 40 pixels corresponding to the convolution stream with a resolution of 80 pixels × 80 pixels.
[0190] The second convolution result with a resolution of 20 pixels × 20 pixels is input into an upsampling module ( Figure 3 not shown in the figure) for upsampling the second convolution result with a resolution of 20 pixels × 20 pixels Perform upsampling to obtain a target first feature map with a resolution of 40 pixels × 40 pixels corresponding to the convolution stream with a resolution of 20 pixels × 20 pixels.
[0191] Fuse the features in the target first feature map with a resolution of 40 pixels × 40 pixels corresponding to the convolution stream with a resolution of 80 pixels × 80 pixels, the features in the target first feature map with a resolution of 40 pixels × 40 pixels corresponding to the convolution stream with a resolution of 20 pixels × 20 pixels, and the features in the first convolution result with a resolution of 40 pixels × 40 pixels to obtain a medium-resolution fused feature map with a resolution of 40 pixels × 40 pixels.
[0192] Input the medium-resolution fused feature map with a resolution of 40 pixels × 40 pixels into the deep feature extraction module 3309 for further feature enhancement and fusion processing to obtain a second feature map P4' with medium resolution. Among them, the resolution of the second feature map P4' with medium resolution is 40 pixels × 40 pixels.
[0193] Since in the fusion stream corresponding to the medium resolution, each feature map is of medium resolution, and in the second feature map with medium resolution, the features in the high-resolution feature map and the features in the low-resolution feature map are fused multiple times. Therefore, in the second feature map with medium resolution, the medium resolution is maintained, and at the same time, the fusion of cross-resolution information interaction features is achieved, and thus the features in the second feature map with medium resolution are more accurate.
[0194] After obtaining the second feature map P4' with medium resolution, for example, for the scale with a resolution of 20 pixels × 20 pixels, obtain the first convolution result with a resolution of 20 pixels × 20 pixels obtained by convolving the first feature map P5 with a resolution of 20 pixels × 20 pixels by the convolution stream corresponding to the resolution of 20 pixels × 20 pixels (that is, the intermediate first feature map with a resolution of 20 pixels × 20 pixels ), the candidate second feature map P3' output by the deep feature extraction module 3305 in the fusion stream with a resolution of 80 pixels × 80 pixels larger than the resolution of 20 pixels × 20 pixels (that is, the second feature map P3' with high resolution), and the candidate second feature map P4' output by the deep feature extraction module 3309 in the fusion stream with a resolution of 40 pixels × 40 pixels larger than the resolution of 20 pixels × 20 pixels (that is, the second feature map P4' with medium resolution).
[0195] Next, input the candidate second feature map P3' into the downsampling module ( Figure 3In (not shown in the figure), the candidate second feature map P3' is downsampled to obtain a target first feature map with a resolution of 20 pixels × 20 pixels corresponding to the convolution stream with a resolution of 80 pixels × 80 pixels.
[0196] The candidate second feature map P4' is input into the downsampling module ( Figure 3 not shown in the figure), and the candidate second feature map P4' is downsampled to obtain a target first feature map with a resolution of 20 pixels × 20 pixels corresponding to the convolution stream with a resolution of 40 pixels × 40 pixels.
[0197] The features in the target first feature map with a resolution of 20 pixels × 20 pixels corresponding to the convolution stream with a resolution of 80 pixels × 80 pixels, the features in the target first feature map with a resolution of 20 pixels × 20 pixels corresponding to the convolution stream with a resolution of 40 pixels × 40 pixels, and the features in the first convolution result R 3 in 3 are fused to obtain a low-resolution fused feature map with a resolution of 20 pixels × 20 pixels.
[0198] The low-resolution fused feature map with a resolution of 20 pixels × 20 pixels is input into the deep feature extraction module 3312 for further feature enhancement and fusion processing to obtain a low-resolution second feature map P5'. Among them, the resolution of the low-resolution second feature map P5' is 20 pixels × 20 pixels.
[0199] Since in the fusion stream corresponding to the low resolution, each feature map is of low resolution, and in the low-resolution second feature map, the features in the high-resolution feature map and the features in the medium-resolution feature map are fused multiple times. Therefore, in the low-resolution second feature map, both the low resolution is maintained and the fusion of cross-resolution information interaction features is achieved at the same time, and thus the features in the low-resolution second feature map are more accurate.
[0200] S204. Input each scale of the second feature map into the head network unit in the target detection model, and the head network unit outputs the detection result.
[0201] In one example, the head network unit includes detection modules corresponding to multiple scales respectively, and each scale of the detection module is used to perform object detection on the second feature map of that scale.
[0202] Exemplarily, it can be combined with Figure 3 to understand the head network unit in the target detection model, such as Figure 3As shown in the figure, the head network unit 34 includes a detection module 3401 connected to the deep feature extraction module 3305, a detection module 3402 connected to the deep feature extraction module 3309, and a detection module 3403 connected to the deep feature extraction module 3312.
[0203] Among them, the detection module 3401 is used to detect small-scale transmission towers in the image frame 31 based on the second feature map P3' with high resolution. The detection module 3402 is used to detect medium-scale transmission towers in the image frame 31 based on the second feature map P4' with medium resolution. The detection module 3403 is used to detect large-scale transmission towers in the image frame 31 based on the second feature map P5' with low resolution.
[0204] In one example, as Figure 3 shown, still referring to the example in step S203, after obtaining the second feature map P3' with high resolution, based on Figure 3 the object detection model in, the second feature map P3' with high resolution is input into the detection module 3401 to detect the transmission towers, and the categories of small-scale transmission towers in the image frame 31 are obtained. Among them, the categories of transmission towers have been introduced in step S104. For details, please refer to the description in step S104 and will not be elaborated here.
[0205] Since the second feature map with high resolution has high resolution, it contains more small-scale detailed features in the image frame, and more global semantic features and context information in more image frames at medium and low resolutions are fused in the second feature map with high resolution. Therefore, the second feature map with high resolution has both more small-scale detailed features in the image frame and more global semantic features and context information in more image frames. Furthermore, the accuracy of identifying the categories of small-scale transmission towers based on the second feature map with high resolution is higher.
[0206] After obtaining the second feature map P4' with medium resolution, based on Figure 3 the object detection model in, the second feature map P4' with medium resolution is input into the detection module 3402 to identify the transmission towers, and the categories of medium-scale transmission towers in the image frame 31 are obtained.
[0207] Since the second feature map with medium resolution has medium resolution, it contains more medium-scale detailed features in the image frame, and more detailed features and semantic features of other scales in the image frame at high and low resolutions are fused. Therefore, the second feature map with medium resolution has both medium-scale detailed features in the image frame and more detailed features and semantic features of other scales. Furthermore, the accuracy of identifying the categories of medium-scale transmission towers based on the second feature map with medium resolution is higher.
[0208] After obtaining the second feature map P5' with low resolution, based on the Figure 3 object detection model in, the second feature map P5' with low resolution is input into the detection module 3403 to identify transmission towers, and the categories of large-scale transmission towers in the image frame 31 are obtained.
[0209] Since the second feature map with low resolution has low resolution, it contains more large-scale semantic features in the image frames, and integrates more image detail features and semantic features of the image frames at high resolution and medium resolution. Therefore, the second feature map with low resolution has both more large-scale semantic features in the image frames and more detail features and semantic features of other scales. Furthermore, the accuracy of identifying the categories of large-scale transmission towers based on the second feature map with low resolution is higher.
[0210] S205. Compare the number and categories of transmission towers detected in the two adjacent image frames respectively; in response to the comparison result meeting the preset conditions, output the categories of each transmission tower in the latter image frame.
[0211] In one example, the image frame to be detected includes two continuously captured image frames.
[0212] Exemplarily, when inputting second feature maps of multiple scales into the object detection model to detect transmission towers and obtaining the categories of transmission towers in the image frame to be detected, the target boxes of the transmission towers in the image frame to be detected can also be obtained; wherein, the target box represents the position range of the transmission tower in the image frame, and each target box is associated with the category of the transmission tower in the target box.
[0213] In the image frame to be detected, for example, the target boxes associated with the categories can be sorted according to a preset rule to obtain the order of the categories.
[0214] Among them, the preset rule can be to calculate the coordinates of the center position point of each target box in the image frame to be detected, and then sort them in ascending or descending order according to the distance between the center position points of the target boxes and the upper left corner of the image frame to be detected; or it can also be to calculate the area of each target box in the image frame to be detected and sort them in ascending or descending order according to the area. This embodiment does not limit the preset rule, as long as it can sort the target boxes in the image frame to be detected and obtain the order of the categories associated with the target boxes.
[0215] In one example, assume that the image frame T1 to be detected includes 3 transmission towers, namely transmission tower A, transmission tower B, and transmission tower C. The target box of transmission tower A detected based on the target detection model is A1, the category is L1, and the distance from the center position point of the target box A1 to the upper left corner of the image frame T1 to be detected is A2; the target box of transmission tower B is B1, the category is L2, and the distance from the center position point of the target box B1 to the upper left corner of the image frame T1 to be detected is B2; the target box of transmission tower C is C1, the category is L3, and the distance from the center position point of the target box C1 to the upper left corner of the image frame T1 to be detected is C2, and A2 < B2 < C2.
[0216] Assume that the preset rule is to sort the categories associated with the target boxes in ascending order according to the distance from the center position points of the target boxes to the upper left corner of the image frame to be detected. Then, in the image frame T1 to be detected, the order of category L1 is 1, the order of category L2 is 2, and the order of category L3 is 3.
[0217] In one example, the preset condition can be that the number of transmission towers identified in the previous and subsequent image frames is equal, and the categories of the transmission towers identified in the previous and subsequent image frames are the same in order.
[0218] Assume that the image frame previous to the image frame T1 to be detected is the image frame T2, and assume that the image frame T2 includes 3 transmission towers, namely transmission tower D, transmission tower E, and transmission tower F. The target box of transmission tower D detected based on the target detection model is D1, the category is L1, and the distance from the center position point of the target box D1 to the upper left corner of the image frame T2 is D2; the target box of transmission tower E is E1, the category is L2, and the distance from the center position point of the target box E1 to the upper left corner of the image frame T2 is E2; the target box of transmission tower F is F1, the category is L3, and the distance from the center position point of the target box F1 to the upper left corner of the image frame T2 is F2, and D2 < E2 < F2. Then, in the image frame T2, the order of category L1 is 1, the order of category L2 is 2, and the order of category L3 is 3.
[0219] Therefore, in the image frame T1 to be detected and the image frame T2, the number of target boxes (i.e., the number of transmission towers) is 3, and the categories of each order are also the same. Then, output the categories in the image frame T1 to be detected and the target boxes associated with the categories. Otherwise, do not output the categories in the image frame T1 to be detected and the target boxes associated with the categories.
[0220] In two consecutive image frames, if the number of transmission towers detected based on the object detection model is the same, and the categories in each order are also the same, then the categories in the latter image frame and the object bounding boxes associated with the categories are output, which can avoid the problem of inaccurate recognition of the categories and object bounding boxes of transmission towers in a single image frame by the object detection model, and improve the accuracy of recognizing the categories and object bounding boxes of transmission towers.
[0221] Exemplarily, before step S202, an initial object detection model can also be constructed first according to the Figure 3 network structure therein. Then, as Figure 3 shown, the image frame to be trained (for example, image frame 31) and the calibrated categories and calibrated object bounding boxes corresponding to the transmission towers in this image frame to be trained are input into the backbone network unit 32 in the initial object detection model to perform feature extraction on this image frame to be trained, and first feature maps of multiple scales are obtained.
[0222] Among them, the implementation of obtaining the image frame to be trained is similar to the implementation of obtaining the image frame to be detected introduced in step S101. For details, reference can be made to the description in step S101, which will not be elaborated here. The calibrated categories and calibrated object bounding boxes can be understood as correct categories and correct object bounding boxes, and can be obtained through manual annotation or from the recognition results of the categories and object bounding boxes of historical transmission towers. The present embodiment does not limit the acquisition method of the calibrated categories and calibrated object bounding boxes.
[0223] Secondly, based on the initial object detection model, the first feature maps of each scale are input into the neck network unit 33 in the initial object detection model, and the neck network unit 33 performs feature fusion on the first feature map of each scale with the first feature maps of the other scales to obtain second feature maps of multiple scales.
[0224] Then, the second feature map of each scale is input into the head network unit 34 in the initial object detection model, and the head network unit outputs the categories and object bounding boxes of the transmission towers in the image frame to be trained.
[0225] Finally, the categories, object bounding boxes, and the calibrated categories and calibrated object bounding boxes corresponding to the transmission towers are input into the loss function ( Figure 3 not shown in the figure) in the initial object detection model to calculate the loss value, and backpropagation is performed in the initial object detection model according to the loss value to update the parameters in the initial object detection model and minimize the loss value.
[0226] Repeatedly input multiple batches of training image frames, the calibrated categories and calibrated target bounding boxes corresponding to the transmission towers in the training image frames, into the initial object detection model, and perform iterative training on the initial object detection model until the loss value tends to be stable and no longer decreases significantly (i.e., the initial object detection model reaches convergence) or reaches a preset number of iterations, and use the initial object detection model as the object detection model. Among them, the object detection model is used to detect the categories and target bounding boxes of the transmission towers in the image frames to be detected in steps S202 - S205.
[0227] Since when the initial object detection model is being trained, it simultaneously captures the multi-scale information of the training image frames in the second feature maps of multiple scales, enhances the semantic information of the image frames, and thus improves the accuracy of detecting the categories and target bounding boxes of the transmission towers in the training image frames. Furthermore, based on the more accurate categories and target bounding boxes, the parameters of the initial object detection model can be adjusted, and the parameters of the initial object detection model can be updated more accurately.
[0228] Figure 4 It is a schematic structural diagram of the detection device for transmission towers provided by this application, as Figure 4 shown. The detection device 40 for transmission towers provided in this embodiment includes:
[0229] An acquisition module 401, configured to acquire an image frame to be detected, and the image frame to be detected is an image including a transmission tower;
[0230] An extraction module 402, configured to extract first feature maps of multiple scales of the image frame to be detected;
[0231] A fusion module 403, configured to, for each scale of the first feature maps, perform feature fusion on the first feature maps of the remaining scales and the first feature maps of this scale to obtain second feature maps of this scale;
[0232] A detection module 404, configured to detect each scale of the second feature maps to obtain a detection result, and the detection result includes the category of the transmission tower.
[0233] In a possible implementation manner, the extraction module 402 is specifically configured to:
[0234] Input the image frame to be detected into the backbone network unit in the object detection model for feature extraction, and respectively obtain the first feature maps output by each of the multiple feature extraction layers cascaded by the backbone network unit to obtain first feature maps of multiple scales.
[0235] In a possible implementation manner, the fusion module 403 is specifically configured to:
[0236] For any first feature map of the remaining scales, determine the target sampling module corresponding to the remaining scale; use the target sampling module to process the first feature map of the remaining scale to obtain the target first feature map of the remaining scale; the resolution of the target first feature map is the same as that of the first feature map of this scale.
[0237] Fuse the target first feature map corresponding to each remaining scale and the first feature map of this scale to obtain the second feature map of this scale.
[0238] In a possible implementation, the target sampling module includes an upsampling module for upsampling or a downsampling module for downsampling.
[0239] In a possible implementation, the fusion module 403 is specifically configured to:
[0240] Input the first feature maps of each scale into the neck network unit in the target detection model, and the neck network unit performs feature fusion on the first feature map of each scale and the first feature maps of the remaining scales to obtain the second feature maps of multiple scales.
[0241] In a possible implementation, the neck network unit includes convolution streams corresponding to multiple scales respectively; wherein, each convolution stream includes at least one cascaded convolution layer; wherein, the resolution of each first feature map remains unchanged after the convolution operation of any convolution layer.
[0242] And "the neck network unit performs feature fusion on the first feature map of each scale and the first feature maps of the remaining scales" in the fusion module 403 is specifically configured to:
[0243] For each scale, determine the target resolution of the first feature map corresponding to this scale, the first target convolution layer in the convolution stream corresponding to this scale, and the transitional first feature map corresponding to the convolution stream to be input to the first target convolution layer, and determine the second target convolution layer in other convolution streams; wherein, the transitional first feature map is the output of any convolution layer in the convolution stream.
[0244] Input the output of the second target convolution layer in other convolution streams into the first target convolution layer after being processed by the target sampling module, and input the transitional first feature map corresponding to the convolution stream into the first target convolution layer, and the first target convolution layer outputs the fused intermediate first feature map corresponding to this scale.
[0245] In a possible implementation, the neck network unit includes deep feature extraction modules corresponding to multiple scales respectively.
[0246] In a possible implementation, the fusion module 403 is specifically configured to:
[0247] Determine the second feature map of each scale step by step based on the following first operation in descending order of scale:
[0248] Obtain the first convolution result obtained by convolving the convolution stream of the current scale with the first feature map of this scale, the second convolution results obtained by convolving the convolution streams of each scale smaller than the current scale with their respective first feature maps, and the candidate second feature maps output by the deep feature extraction modules of other scales larger than the current scale;
[0249] Input the first convolution result, each second convolution result, and the candidate second feature maps into the deep feature extraction module of this scale, and the deep feature extraction module fuses the first convolution result, each second convolution result, and each candidate second feature map to obtain the second feature map of the current scale;
[0250] Take the next scale smaller than the current scale as the updated current scale, and repeat the first operation.
[0251] In a possible implementation manner, the detection module 404 is specifically configured to:
[0252] Input the second feature map of each scale into the head network unit in the object detection model, and the head network unit outputs the detection result.
[0253] In a possible implementation manner, the head network unit includes detection modules corresponding to multiple scales respectively, and the detection module of each scale is used to perform object detection on the second feature map of this scale.
[0254] In a possible implementation manner, the image frame to be detected includes two continuously captured image frames; the device 40 further includes:
[0255] Compare the number and category of transmission towers detected in the front and rear two image frames respectively;
[0256] In response to the comparison result meeting the preset condition, output the category of each transmission tower in the latter image frame.
[0257] The detection device of the transmission tower provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effect are similar, and will not be elaborated here in this embodiment.
[0258] Figure 5 It is a schematic structural diagram of the electronic device provided in this application. As Figure 5 shown, the electronic device 50 provided in this embodiment includes: at least one processor 501 and a memory 502. Optionally, the device 50 further includes a communication component 503. Among them, the processor 501, the memory 502, and the communication component 503 are connected through a bus 504.
[0259] In a specific implementation process, at least one processor 501 executes computer-executable instructions stored in a memory 502, so that at least one processor 501 executes the above-mentioned method.
[0260] For the specific implementation process of the processor 501, reference may be made to the above method embodiments. Their implementation principles and technical effects are similar, and will not be elaborated here in this embodiment.
[0261] In the above embodiments, it should be understood that the processor may be a central processing unit (Central Processing Unit, CPU for short), or may also be other general-purpose processors, digital signal processors (Digital Signal Processor, DSP for short), application specific integrated circuits (Application Specific Integrated Circuit, ASIC for short), etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor.
[0262] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.
[0263] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.
[0264] This application also provides a computer program product, including a computer program, which implements the above-mentioned method when executed by a processor.
[0265] This application also provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the processor executes the computer-executable instructions, the above-mentioned method is implemented.
[0266] The above-readable storage medium may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium may be any available medium accessible by a general-purpose or special-purpose computer.
[0267] An exemplary readable storage medium is coupled to the processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium may also be an integral part of the processor. The processor and the readable storage medium may be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium may also exist as discrete components in a device.
[0268] The division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other may be indirect couplings or communication connections through some interfaces, devices or units, and may be in electrical, mechanical or other forms.
[0269] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0270] In addition, the functional units in various embodiments of the present invention may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit.
[0271] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs and other various media that can store program codes.
[0272] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When this program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: ROMs, RAMs, magnetic disks, or optical discs and other various media that can store program codes.
[0273] Finally, it should be noted that: After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily think of other implementation manners of the present invention. The present invention is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field not disclosed in the present invention. It is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.
Claims
1. A method for detecting a transmission tower, characterized in that: include: Acquire an image frame to be detected, wherein the image frame to be detected is an image including a transmission tower; Extracting first feature maps of multiple scales of the image frame to be detected; For the first feature map of each scale, the first feature maps of other scales are fused with the first feature map of the scale to obtain the second feature map of the scale; The second feature map of each scale is detected to obtain a detection result, wherein the detection result includes a category of a transmission tower.
2. The method according to claim 1, characterized in that The step of extracting first feature maps of multiple scales of the image frame to be detected includes: The image frame to be detected is input into a backbone network unit in a target detection model for feature extraction, and first feature maps of respective outputs are respectively obtained from a plurality of feature extraction layers cascaded from the backbone network unit to obtain first feature maps of the plurality of scales.
3. The method according to claim 1, characterized in that For the first feature map of each scale, the first feature maps of other scales are subjected to feature fusion with the first feature map of the scale to obtain the second feature map of the scale, including: For any first feature map of other scales, determine a target sampling module corresponding to the other scales; use the target sampling module to process the first feature map of the other scales to obtain a target first feature map of the other scales; the resolution of the target first feature map is the same as the resolution of the first feature map of the scale; The target first feature map corresponding to each of the remaining scales and the first feature map of the scale are fused to obtain a second feature map of the scale.
4. The method according to claim 3, characterized in that The target sampling module includes an up-sampling module for up-sampling or a down-sampling module for down-sampling.
5. The method according to claim 1, characterized in that: For the first feature map of each scale, the first feature maps of other scales are subjected to feature fusion with the first feature map of the scale to obtain the second feature map of the scale, including: The first feature map of each scale is input into the neck network unit in the target detection model, and the neck network unit performs feature fusion on the first feature map of each scale with the first feature maps of other scales to obtain second feature maps of multiple scales.
6. The method according to claim 5, characterized in that The neck network unit includes convolution flows corresponding to a plurality of scales respectively; wherein each convolution flow includes at least one convolution layer in cascade; wherein the resolution of each first feature map remains unchanged after the convolution operation of any convolution layer; And the neck network unit performs feature fusion on the first feature map of each scale with the first feature maps of other scales, including: For each scale, determine the target resolution of the first feature map corresponding to the scale, the first target convolution layer in the convolution flow corresponding to the scale, and the transition first feature map corresponding to the convolution flow to be input to the first target convolution layer, and determine the second target convolution layer in other convolution flows; wherein the transition first feature map is the output of any convolution layer in the convolution flow; The output of the second target convolution layer in other convolution flows is processed by the target sampling module and then input into the first target convolution layer, and the transition first feature map corresponding to the convolution flow is input into the first target convolution layer, and the first target convolution layer outputs the fused intermediate first feature map corresponding to the scale.
7. The method according to claim 6, characterized in that The neck network unit includes deep feature extraction modules corresponding to multiple scales respectively; For the first feature map of each scale, the first feature maps of other scales are subjected to feature fusion with the first feature map of the scale to obtain the second feature map of the scale, including: In order from large to small scales, the second feature map of each scale is determined step by step based on the following first operation: Obtain a first convolution result obtained by convolving a convolution flow of a current scale with a first feature map of the scale, a second convolution result obtained by convolving convolution flows of scales smaller than the current scale with respective first feature maps, and a candidate second feature map output by a deep feature extraction module of other scales larger than the current scale; Inputting the first convolution result, each of the second convolution results, and the candidate second feature map into a deep feature extraction module of the scale, and the deep feature extraction module fuses the first convolution result, each of the second convolution results, and each of the candidate second feature maps to obtain a second feature map of the current scale; The next scale whose scale is smaller than the current scale is used as the updated current scale, and the first operation is repeatedly performed.
8. The method according to any one of claims 1 to 7, characterized in that The detecting of the second feature map of each scale to obtain the detection result includes: The second feature map of each scale is input into the head network unit in the target detection model, and the head network unit outputs the detection result.
9. The method according to claim 8, characterized in that The head network unit includes detection modules corresponding to multiple scales, and the detection module of each scale is used to perform target detection on the second feature map of the scale.
10. The method according to any one of claims 1 to 7, characterized in that The image frames to be detected include two image frames taken continuously; the method further includes: Compare the number and type of transmission towers detected in the two image frames before and after; In response to the comparison result satisfying a preset condition, the category of each transmission tower in the next image frame is output.
11. A detection device for a transmission tower, characterized in that: include: An acquisition module, used for acquiring an image frame to be detected, wherein the image frame to be detected is an image including a transmission tower; An extraction module, used for extracting first feature maps of multiple scales of the image frame to be detected; A fusion module is used to fuse the first feature maps of other scales with the first feature map of the scale for each scale, so as to obtain a second feature map of the scale; The detection module is used to detect the second feature map of each scale to obtain a detection result, wherein the detection result includes the category of the transmission tower.
12. An electronic device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 10 when executed by a processor.
14. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 10 when being executed by a processor.