Antenna identification method based on unmanned aerial vehicle side cloud cooperation mode
By applying an antenna recognition edge model on the drone, filtering and keyframe extraction of video data, the problem of low video data transmission efficiency of drone is solved, and efficient data transmission and storage is achieved.
Patent Information
- Application Number
- CN202510076006.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, drones directly transmit video frames to the server, resulting in large data transmission volume, low transmission efficiency, and occupying a large amount of data storage space of cloud servers.
An antenna recognition method based on the drone edge cloud collaboration mode is adopted, and video data is detected through the antenna recognition edge model to obtain the first video frame, and the video frame is filtered according to the pixel threshold and the number of occurrence thresholds, video keyframes and positioning information are obtained, and only uploaded to the cloud server.
It reduces the amount of data transmission, improves data transmission efficiency, saves data storage space in cloud servers, and improves the accuracy of small target detection.
Smart Images

Figure CN120071194A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of antenna detection, and particularly relates to an antenna recognition method based on the edge-cloud collaboration mode of an unmanned aerial vehicle (UAV). Background Art
[0002] In urban construction, if residents install antennas without permission, it will cause spectrum congestion in the communication frequency band in the same geographical area, damaging the communication stability of the legally built base station communication system. In the prior art, the method of using UAV patrol and shooting is adopted to find out whether there are antennas installed by residents without permission in the geographical area. However, directly transmitting the collected video frames by the UAV to the server has problems of large data transmission volume and low transmission efficiency, and it will also occupy a large amount of data storage space of the cloud server. Summary of the Invention
[0003] The purpose of this application is to overcome the deficiencies in the prior art and provide an antenna recognition method based on the edge-cloud collaboration mode of an unmanned aerial vehicle, which can reduce the data transmission volume, improve the data transmission efficiency, and save the data storage space occupied by the cloud server.
[0004] The first aspect of the embodiments of this application provides an antenna recognition method based on the edge-cloud collaboration mode of an unmanned aerial vehicle, which is applied to the UAV and includes:
[0005] Input the video data captured by the UAV into the antenna recognition edge model to obtain a plurality of first video frames; the first video frames include a number of recognized antenna targets;
[0006] Filter the first video frames according to a preset pixel threshold and the pixels of each antenna target to obtain a plurality of second video frames;
[0007] Obtain video key frames according to a preset number threshold and the appearance times of each antenna target in the plurality of second video frames; the appearance times of the antenna targets in the video key frames are greater than the number threshold;
[0008] Upload the video key frames including a number of antenna targets and the positioning information of the video key frames to the cloud server;
[0009] Wherein, the antenna recognition edge model includes a backbone network, an enhancement network, a neck network, and a head network;
[0010] The step of inputting the video data captured by the UAV into the antenna recognition edge model to obtain a plurality of first video frames includes:
[0011] Input the video data into the backbone network for feature extraction processing to obtain first features of each video frame;
[0012] Input the first features of each video frame into the enhancement network for enhanced convolutional processing to obtain the second features of each video frame;
[0013] Input the second features of each video frame into the neck network for bidirectional feature convolutional processing to obtain the third features of each video frame;
[0014] Input the third features of each video frame into the head network for feature prediction processing to obtain multiple first video frames.
[0015] Compared with the prior art, the present application detects video data through an antenna recognition edge model to obtain multiple first video frames. Among them, the antenna recognition edge model includes a backbone network, an enhancement network, a neck network, and a head network. The second features of video frames can be quickly obtained through the backbone network and the enhancement network, and then the second features are subjected to bidirectional feature convolutional processing through the neck network to obtain the third features of each video frame for the head network to perform feature prediction processing, so as to obtain the first video frames including several recognized antenna targets, and the antenna targets can be efficiently and accurately detected from the video frames of the video data; then filter the first video frames according to the pixel threshold to obtain multiple second video frames, and then obtain the video key frames whose occurrence times are greater than the preset number threshold, so as to upload the video key frames and the positioning information of the video key frames to the cloud server, which can improve the small target detection accuracy, reduce the data transmission volume, improve the data transmission efficiency, and efficiently and quickly upload the video key frames and the positioning information of the video key frames to the cloud server, saving the data storage space occupied by the cloud server.
[0016] In order to understand the present application more clearly, the specific implementation manners of the present application will be described below in conjunction with the accompanying drawings. Brief Description of the Drawings
[0017] Figure 1 It is a flowchart of an antenna recognition method based on the drone edge-cloud collaboration mode according to an embodiment of the present application.
[0018] Figure 2 It is a schematic diagram of an antenna recognition edge model of an antenna recognition method based on the drone edge-cloud collaboration mode according to an embodiment of the present application.
[0019] Figure 3 It is a flowchart of step S1 of an antenna recognition method based on the drone edge-cloud collaboration mode according to an embodiment of the present application.
[0020] Figure 4 It is a schematic diagram of the backbone network of an antenna recognition edge model of an antenna recognition method based on the drone edge-cloud collaboration mode according to an embodiment of the present application.
[0021] Figure 5 Schematic diagram of the target enhancement layer of the antenna recognition edge model of the antenna recognition method based on the drone edge-cloud collaboration mode according to an embodiment of the present application.
[0022] Figure 6 Schematic diagram of the heterogeneous convolution module of the antenna recognition edge model of the antenna recognition method based on the drone edge-cloud collaboration mode according to an embodiment of the present application. Detailed implementation manners
[0023] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0024] It should be clear that the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the embodiments of the present application.
[0025] When the following description involves the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. In the description of the present application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects, and do not have to be used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances. The singular forms of "a", "the" and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. The word "if" / "when" used herein can be interpreted as "when" or "while" or "in response to determining".
[0026] In addition, in the description of the present application, unless otherwise specified, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0027] Please refer to Figure 1 , which is a flowchart of the antenna recognition method based on the drone edge-cloud collaboration mode according to an embodiment of the present application, and is applied to a drone. A drone refers to a remotely controlled flying device with a shooting module, a communication module and a video frame data processing module.
[0028] The method includes:
[0029] S1: Input the video data captured by the drone into the antenna recognition edge model to obtain multiple first video frames; the first video frames include several recognized antenna targets.
[0030] Among them, the antenna recognition edge model can be an algorithm model running on the drone.
[0031] Please refer to Figure 2 , the antenna recognition edge model includes a backbone network, an enhancement network, a neck network, and a head network;
[0032] Please refer to Figure 3 , the steps of S1 include:
[0033] S11: Input the video data into the backbone network for feature extraction processing to obtain the first features of each video frame;
[0034] S12: Input the first features of each video frame into the enhancement network for enhanced convolution processing to obtain the second features of each video frame.
[0035] S13: Input the second features of each video frame into the neck network for bidirectional feature convolution processing to obtain the third features of each video frame.
[0036] S14: Input the third features of each video frame into the head network for feature prediction processing to obtain multiple first video frames.
[0037] S2: Filter the first video frames according to a preset pixel threshold and the pixels of each antenna target to obtain multiple second video frames.
[0038] In a feasible embodiment, step S2 includes:
[0039] S21: Compare the pixels of each antenna target in the first video frame with the pixel threshold respectively.
[0040] Among them, the first video frame will frame the recognized antenna through a detection box, and the image content within the detection box is the antenna target. The size of the detection box is different, and the corresponding pixel size of the antenna target is also different. All antenna targets in the first video frame need to be compared with the pixel threshold respectively.
[0041] S22: If the pixels of at least one antenna target are less than or equal to the pixel threshold, determine the first video frame as the second video frame.
[0042] Among them, since the pixel size of the antenna in the video frames captured by the drone is limited, through a pixel threshold, the first video frame including only the antenna target with overly large pixels misdetected by the antenna recognition edge model can be filtered to obtain a second video frame including antenna targets with pixels less than or equal to the pixel threshold, preventing the misdetected targets from interfering with the key frames.
[0043] S3: Obtain video key frames according to a preset number threshold and the occurrence times of each of the antenna targets in the multiple second video frames; the occurrence times of the antenna targets in the video key frames are greater than the number threshold.
[0044] In a feasible embodiment, step S3 includes:
[0045] S311: Accumulate the occurrence times of the continuously appearing antenna targets according to the order of the second video frames.
[0046] Among them, antenna identifiers can be added to each antenna target for distinction, and the antenna identifiers and the occurrence times of the corresponding antenna targets are stored in an associated manner.
[0047] S312: Compare the occurrence times of the antenna targets in the current second video frame with the number threshold.
[0048] Among them, the number threshold can be set by the user.
[0049] S313: If the occurrence times are greater than the number threshold, determine the current second video frame as a video key frame.
[0050] Among them, since the video data of one second in length already includes multiple video frames, the accurately recognized antenna targets will appear repeatedly in multiple video frames. Therefore, the second video frames can be screened through the number threshold to obtain the second video frames where the antenna targets appear multiple times as video key frames.
[0051] In this embodiment, according to the number threshold, screening to obtain video key frames where the occurrence times of the antenna targets are greater than the number threshold can reduce the number of video frames that need to be uploaded to the cloud server, thereby reducing the data transmission volume and improving the data transmission efficiency.
[0052] S4: Upload the video key frames including several antenna targets and the positioning information of the video key frames to the cloud server.
[0053] Among them, the video key frames and the positioning information can be uploaded to the cloud server through a wireless communication module for storage in the cloud server, facilitating the user to call and view the video key frames stored in the cloud server.
[0054] Compared with the prior art, in the present application, the antenna recognition edge model is used to detect video data to obtain a plurality of first video frames. Among them, the antenna recognition edge model includes a backbone network, an enhancement network, a neck network, and a head network. The second feature of the video frame can be quickly obtained through the backbone network and the enhancement network, and then the second feature is subjected to bidirectional feature convolution processing through the neck network to obtain the third feature of each video frame for the head network to perform feature prediction processing, so as to obtain the first video frame including several recognized antenna targets, and the antenna targets can be efficiently and accurately detected from the video frames of the video data; then the first video frames are filtered according to the pixel threshold to obtain a plurality of second video frames, and then the video key frames with the number of occurrences greater than the preset number threshold are obtained, so as to upload the video key frames and the positioning information of the video key frames to the cloud server, which can improve the detection accuracy of small targets, reduce the data transmission volume, improve the data transmission efficiency, and upload the video key frames and the positioning information of the video key frames to the cloud server efficiently and quickly, saving the data storage space occupied by the cloud server.
[0055] In a feasible embodiment, the step S313: if the number of occurrences is greater than the number threshold and determine the current second video frame as the video key frame includes:
[0056] If the number of occurrences is greater than the number threshold, and the antenna target of the current second video frame is different from the key target of the already determined video key frame, determine the current second video frame as the video key frame, and determine the antenna target of the key frame as the key target.
[0057] In this embodiment, in combination with the number threshold and whether the antenna target of the current second video frame is the same as the key target of the already determined video key frame to determine whether the current second video frame is determined as the video key frame, it is possible to prevent multiple second video frames of the same antenna target that appear multiple times from being used as video key frames, so as to reduce the number of video key frames with repeated content, and further reduce the data transmission volume and improve the data transmission efficiency.
[0058] In another feasible embodiment, the step S3: obtaining the video key frame according to the preset number threshold and the number of occurrences of each antenna target of the plurality of second video frames includes:
[0059] S321: Obtain the number of occurrences of each antenna target in a plurality of consecutive second video frames.
[0060] S322: Compare the number of occurrences of each antenna target with the number threshold.
[0061] S323: If the number of occurrences is greater than the number threshold, determine the corresponding antenna target as the key target.
[0062] S324: Determine a second video frame including the key target as the video key frame.
[0063] Among them, a second video frame including the key target can be any second video frame in which the antenna target appears. That is, as long as the number of appearances of the antenna target is greater than the number threshold, one second video frame can be obtained from the consecutive second video frames in which the antenna target appears as the video key frame. For example, starting from the first second video frame in which the antenna target appears to the last second video frame in which the antenna target appears, and among the intermediate consecutive second video frames in which the antenna target appears, any one second video frame can be used as the video key frame.
[0064] In this embodiment, determining a second video frame including the key target as the video key frame can prevent repeated acquisition of video key frames with the same content, so as to reduce the data transmission volume.
[0065] In a feasible embodiment, the backbone network includes a feature extraction layer and a plurality of sequentially cascaded combined convolutional layers; among them, the combined convolutional layer includes a first convolutional module and a second convolutional module;
[0066] S11: The step of inputting the video data into the backbone network for feature extraction processing to obtain the first features of each video frame includes:
[0067] S111: Input the video data into the feature extraction layer for feature extraction processing to obtain the fourth features of each video frame.
[0068] S112: Input the fourth features of each video frame into the first convolutional module of the first combined convolutional layer for convolutional processing to obtain the fifth features of each video frame.
[0069] S113: Input the fifth features of each video frame into the second convolutional module of the same combined convolutional layer for convolutional processing to obtain the sixth features of each video frame.
[0070] Please refer to Figure 4 , among which, the second convolutional module includes a plurality of sequentially cascaded phantom convolutional units and convolutional fusion units;
[0071] The step of S113: Inputting the fifth features of each video frame into the second convolutional module of the same combined convolutional layer for convolutional processing to obtain the sixth features of each video frame includes:
[0072] S1131: Input the fifth feature into a plurality of sequentially cascaded phantom convolutional units for convolutional processing to obtain the seventh features output by each phantom convolutional unit.
[0073] S1132: Input the fifth feature and the seventh feature output by each phantom convolution unit into the convolution fusion unit for convolution fusion to obtain the sixth feature.
[0074] Among them, combining S1131 and S1132 of step S113, the sixth feature can be obtained through the following formula:
[0075]
[0076] Y is the sixth feature, x is the fifth feature, cat(·) is the data concatenation function, and y i represents the seventh feature output after i cascaded phantom convolution units, and n is the number of cascaded phantom convolution units.
[0077] S114: Input the sixth feature of each video frame into the next-level combined convolution layer for convolution processing, and determine the sixth feature of each video frame output by the last-level combined convolution layer as the first feature.
[0078] In this embodiment, by using the first convolution module and the second convolution module including a plurality of sequentially cascaded phantom convolution units and convolution fusion units to perform feature extraction processing on each video frame, the original details of the video frame can be maximally retained.
[0079] Please continue to refer to Figure 2 , in a feasible embodiment, the backbone network includes a plurality of sequentially cascaded combined convolution layers; the enhancement network includes a first convolution layer, a second convolution layer, and a target enhancement layer;
[0080] The step S12: Input the first feature of each video frame into the enhancement network for enhanced convolution processing to obtain the second feature of each video frame includes:
[0081] S121: Input the first features output by the penultimate and antepenultimate combined convolution layers of the backbone network into the first convolution layer and the second convolution layer respectively for convolution processing to obtain a first enhanced feature and a second enhanced feature.
[0082] S122: Input the first feature output by the last combined convolution layer of the backbone network into the target enhancement layer for enhanced convolution processing to obtain a third enhanced feature.
[0083] Please refer to Figure 5 , the network structure of the target enhancement layer is as Figure 5As shown, the target enhancement layer first smoothly extracts a concise feature map through 1×1 convolution and 3×3 convolution, and then performs dilated residual attention operation on the feature map to filter branch features through convolutions with different dilation rates. Each branch in the dilated residual attention operation has a unique receptive field, which can form a comprehensive feature representation. Among them, due to the small characteristics of the antenna interference source target, in this embodiment, the dilation rate 5 used in the convolution is replaced with dilation rate 2 to minimize the redundancy in the receptive field. Finally, we use convolution to compress the channels to perform convolution compression on the output results of multiple branches in the dilated residual attention operation to obtain a comprehensive third enhanced feature.
[0084] S123: Obtain the second feature according to the first enhanced feature, the second enhanced feature, and the third enhanced feature.
[0085] In this embodiment, through the first convolutional layer, the second convolutional layer, and the target enhancement layer, the first enhanced feature, the second enhanced feature, and the third enhanced feature obtained by convolutional enhancement can obtain a more comprehensive second feature.
[0086] Please continue to refer to Figure 2 , in a feasible embodiment, the second feature includes the first enhanced feature, the second enhanced feature, and the third enhanced feature; the neck network includes a first bidirectional feature extraction layer, a second bidirectional feature extraction layer, a third bidirectional feature extraction layer, and a fourth bidirectional feature extraction layer;
[0087] The step S13: Input the second feature of each video frame into the neck network for bidirectional feature convolution processing to obtain the third feature of each video frame includes:
[0088] S131: Input the third enhanced feature into the first bidirectional feature extraction layer for upsampling processing, and then perform bidirectional feature convolution processing with the second enhanced feature to obtain a first bidirectional convolution feature.
[0089] S132: Input the first bidirectional convolution feature into the second bidirectional feature extraction layer for upsampling processing, and then perform bidirectional feature convolution processing with the first enhanced feature to obtain a second bidirectional convolution feature.
[0090] S133: Input the second bidirectional convolution feature into the third bidirectional feature extraction layer for convolution processing, and then perform bidirectional feature convolution processing with the second enhanced feature to obtain a third bidirectional convolution feature.
[0091] S134: Input the third bidirectional convolution feature into the fourth bidirectional feature extraction layer for convolution processing, and then perform bidirectional feature convolution processing with the third enhanced feature to obtain a fourth bidirectional convolution feature.
[0092] S135: Obtain the third feature according to the second bidirectional convolution feature, the third bidirectional convolution feature, and the fourth bidirectional convolution feature.
[0093] Please continue to refer to Figure 2 , the first bidirectional feature extraction layer and the second bidirectional feature extraction layer each include an upsampling module, a bidirectional feature extraction module, and a heterogeneous convolution module (HetBlock); the third bidirectional feature extraction layer and the fourth bidirectional feature extraction layer each include a convolution module, a bidirectional feature extraction module, and a heterogeneous convolution module.
[0094] Please refer to Figure 6 , the network structure of the heterogeneous convolution module (Hetblock) is as Figure 6 shown. The heterogeneous convolution module includes a first convolution unit, a splitting unit (SPlit), a first bottleneck feature extraction unit (HetBottleneck[1]), a second bottleneck feature extraction unit (HetBottleneck[3]), a feature fusion unit, and a second convolution unit; wherein, the input end of the first convolution unit is the input end of the heterogeneous convolution module, the output end of the first convolution unit is connected to the input end of the splitting unit, the output end of the splitting unit is respectively connected to the input end of the first bottleneck feature extraction unit and the input end of the feature fusion unit, the output end of the first bottleneck feature extraction unit is respectively connected to the input end of the second bottleneck feature extraction unit and the input end of the feature fusion unit, the output end of the second bottleneck feature extraction unit is connected to the input end of the feature fusion unit, the output end of the feature fusion unit is connected to the input end of the second convolution unit, and the output end of the second convolution unit is the output end of the heterogeneous convolution module.
[0095] Among them, both the first bottleneck feature extraction unit and the second bottleneck feature extraction unit adopt the network structure of heterogeneous kernel convolution (HetConv). Specifically, in HetConv, half of the convolution kernels are replaced with 1×1 convolution kernels and are alternately arranged with the remaining half of 3×3 convolution kernels to form a filter. The computational cost of these 3×3 convolution kernels is:
[0096]
[0097] And the computational cost of the remaining half of 1×1 convolution kernels is:
[0098]
[0099] Therefore, the total computational cost of a single HetConv operation is:
[0100]
[0101] In this embodiment, by using HetBlock, the number of parameters and computational complexity in the bottleneck can be minimized. At the same time, HetConv still retains a quarter of the alternating convolutional kernels to ensure that the filter can capture the spatial correlation of specific channels.
[0102] The device embodiments described above are merely illustrative. The components described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application. A person of ordinary skill in the art can understand and implement it without creative work.
[0103] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0104] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the selected functions in Figure 1 one or more of the processes Figure 1 or multiple processes and / or blocks Figure 1 one or more of the blocks Figure 1 or multiple blocks.
[0105] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the selected functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 steps of the selected functions in one block or multiple blocks.
[0106] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0107] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0108] Computer-readable media includes permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0109] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.
[0110] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. An antenna identification method based on UAV edge-cloud collaborative mode, characterized in that: Applications in drones include: Inputting the video data captured by the drone into the antenna recognition edge model to obtain a plurality of first video frames; the first video frames include a plurality of recognized antenna targets; Filtering the first video frame according to a preset pixel threshold and pixels of each of the antenna targets to obtain a plurality of second video frames; Obtaining a video key frame according to a preset number threshold and the number of occurrences of each of the antenna targets in the plurality of second video frames; the number of occurrences of the antenna target in the video key frame is greater than the number threshold; Uploading the video key frames including the plurality of antenna targets and the positioning information of the video key frames to a cloud server; Wherein, the antenna recognition edge model includes a backbone network, an enhancement network, a neck network and a head network; The step of inputting the video data shot by the drone into the antenna recognition edge model to obtain a plurality of first video frames includes: Inputting the video data into the backbone network for feature extraction processing to obtain the first feature of each video frame; Inputting the first feature of each video frame into the enhancement network for enhanced convolution processing to obtain the second feature of each video frame; Inputting the second feature of each video frame into the neck network for bidirectional feature convolution processing to obtain the third feature of each video frame; The third feature of each video frame is input into the head network for feature prediction processing to obtain multiple first video frames.
2. The antenna identification method based on the UAV edge-cloud collaborative mode according to claim 1 is characterized in that: The step of filtering the first video frame according to a preset pixel threshold and pixels of each of the antenna targets to obtain a plurality of second video frames comprises: Comparing the pixels of each of the antenna targets in the first video frame with the pixel threshold respectively; If a pixel of at least one of the antenna targets is less than or equal to the pixel threshold, the first video frame is determined as a second video frame.
3. The antenna identification method based on the UAV edge-cloud collaborative mode according to claim 1 is characterized in that: The step of obtaining a video key frame according to a preset number threshold and the number of occurrences of each of the antenna targets in the plurality of second video frames comprises: According to the sequence of the second video frames, accumulating the number of occurrences of the antenna target that appears continuously; Compare the number of occurrences of the antenna target in the current second video frame with the number threshold; If the number of occurrences is greater than the number threshold, the current second video frame is determined as a video key frame.
4. The antenna identification method based on the UAV edge-cloud collaborative mode according to claim 3 is characterized in that: The step of determining the current second video frame as a video key frame if the number of occurrences is greater than the number threshold comprises: If the number of occurrences is greater than the number threshold, and the antenna target of the current second video frame is different from the key target of the determined video key frame, the current second video frame is determined as the video key frame, and the antenna target of the key frame is determined as the key target.
5. The antenna identification method based on the UAV edge-cloud collaborative mode according to claim 1 is characterized in that: The step of obtaining a video key frame according to a preset number threshold and the number of occurrences of each of the antenna targets in the plurality of second video frames comprises: Obtaining the number of occurrences of each of the antenna targets in a plurality of consecutive second video frames; Comparing the number of occurrences of each of the antenna targets with the number threshold; If the number of occurrences is greater than the number threshold, determining the corresponding antenna target as a key target; A second video frame including the key object is determined as the video key frame.
6. The antenna identification method based on the UAV edge-cloud collaborative mode according to claim 1 is characterized in that: The backbone network includes a feature extraction layer and a plurality of sequentially cascaded combined convolutional layers; wherein the combined convolutional layer includes a first convolutional module and a second convolutional module; The step of inputting the video data into the backbone network for feature extraction processing to obtain the first feature of each video frame includes: Inputting the video data into the feature extraction layer for feature extraction processing to obtain a fourth feature of each video frame; Inputting the fourth feature of each video frame into the first convolution module of the first combined convolution layer for convolution processing to obtain the fifth feature of each video frame; Inputting the fifth feature of each video frame into the second convolution module of the same combined convolution layer for convolution processing to obtain the sixth feature of each video frame; The sixth feature of each video frame is input into the next-level combined convolution layer for convolution processing, and the sixth feature of each video frame output by the last-level combined convolution layer is determined as the first feature.
7. The antenna identification method based on the UAV edge-cloud collaborative mode according to claim 6 is characterized by: The second convolution module includes a plurality of phantom convolution units and convolution fusion units cascaded in sequence; The step of inputting the fifth feature of each video frame into the second convolution module of the same combined convolution layer for convolution processing to obtain the sixth feature of each video frame includes: Inputting the fifth feature into a plurality of phantom convolution units cascaded in sequence for convolution processing, to obtain a seventh feature output by each phantom convolution unit; The fifth feature and the seventh feature output by each phantom convolution unit are input into the convolution fusion unit for convolution fusion to obtain the sixth feature.
8. The antenna identification method based on the UAV edge-cloud collaborative mode according to claim 6 is characterized in that: The step of inputting the fifth feature of each video frame into the second convolution module of the same combined convolution layer for convolution processing to obtain the sixth feature of each video frame includes: The sixth characteristic is obtained by the following formula: Y is the sixth feature, x is the fifth feature, cat(·) is the data concatenation function, y i represents the seventh feature output by i cascaded phantom convolution units, and n is the number of cascaded phantom convolution units.
9. The antenna identification method based on the UAV edge-cloud collaborative mode according to claim 1 is characterized in that: The backbone network includes a plurality of sequentially cascaded combined convolutional layers; the enhanced network includes a first convolutional layer, a second convolutional layer and a target enhancement layer; The step of inputting the first feature of each video frame into the enhanced network for enhanced convolution processing to obtain the second feature of each video frame includes: Inputting the first feature output by the combined convolutional layer of the third-to-last level and the second-to-last level of the backbone network into the first convolutional layer and the second convolutional layer for convolution processing respectively, to obtain a first enhanced feature and a second enhanced feature; Inputting the first feature output by the last level combined convolution layer of the backbone network into the target enhancement layer for enhanced convolution processing to obtain a third enhanced feature; The second feature is obtained according to the first enhanced feature, the second enhanced feature and the third enhanced feature.
10. The antenna identification method based on the UAV edge-cloud collaborative mode according to claim 1 is characterized in that: The second feature includes a first enhancement feature, a second enhancement feature and a third enhancement feature; the neck network includes a first bidirectional feature extraction layer, a second bidirectional feature extraction layer, a third bidirectional feature extraction layer and a fourth bidirectional feature extraction layer; The step of inputting the second feature of each video frame into the neck network for bidirectional feature convolution processing to obtain the third feature of each video frame includes: After inputting the third enhanced feature into the first bidirectional feature extraction layer for upsampling processing, the third enhanced feature is subjected to bidirectional feature convolution processing with the second enhanced feature to obtain a first bidirectional convolution feature; After the first bidirectional convolution feature is input into the second bidirectional feature extraction layer for upsampling processing, the bidirectional feature convolution processing is performed with the first enhanced feature to obtain a second bidirectional convolution feature; After the second bidirectional convolution feature is input into the third bidirectional feature extraction layer for convolution processing, the bidirectional feature convolution processing is performed with the second enhanced feature to obtain a third bidirectional convolution feature; After the third bidirectional convolution feature is input into the fourth bidirectional feature extraction layer for convolution processing, the third bidirectional convolution feature is subjected to bidirectional feature convolution processing with the third enhanced feature to obtain a fourth bidirectional convolution feature; The third feature is obtained according to the second bidirectional convolution feature, the third bidirectional convolution feature and the fourth bidirectional convolution feature.
Citation Information
Patent Citations
Method and device for obtaining video key frames, storage medium and electronic equipment
CN112667851A