Unmanned aerial vehicle online target counting method based on multi-modal dynamic neural network
By applying the multimodal dynamic neural network DyNN framework on drones, combining CGAN, multimodal gating and modal conversion network, the drone's target counting problem in complex environments is solved, and the multimodal target counting effect with high precision and low latency is achieved.
Patent Information
- Application Number
- CN202510113505.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The existing drone aerial image target counting methods are difficult to accurately identify effective targets in images in complex environments such as severe weather and insufficient lighting, and multimodal models are difficult to deploy on drones with limited resources.
A framework based on multimodal dynamic neural network DyNN is proposed, which combines multimodal feature fusion through conditional generation adversarial network (CGAN), and combines decoder-assisted multimodal gating and density map loss modal conversion network to achieve lightweight multimodal feature alignment and dynamic environment adaptation.
High-precision target counting is achieved in complex environments, low inference delay is maintained, and efficient deployment of multimodal models is achieved on resource-constrained drones.
Smart Images

Figure CN120032280A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of machine vision, and in particular is an online target counting method for unmanned aerial vehicles based on a multimodal dynamic neural network. Background Art
[0002] Most existing aerial image object counting methods assume that the UAV is deployed in an environment with high visibility and clear field of view, and can capture rich details about the appearance and texture of the target. [1] However, in reality, drones often encounter bad weather, overexposure, underexposure, or shadows, which can cause the quality of visible light images to deteriorate significantly. [2] , as shown in Figure 1(a). In this case, it is difficult for the counting method that relies solely on visible light to accurately identify the valid targets in the image. [3] The advantage of thermal infrared images is that they can capture surface thermal information of objects that cannot be directly perceived by the human eye. Thermal infrared cameras can be integrated into aerial drones. The collected thermal infrared images, as shown in Figure 1(b), contain the location of the target and the spatial background, which can make up for the target information missing in the visible light image. [4] This multimodal target counting method for visible light / thermal infrared (RGB / Thermal, RGB-T) images is more robust than single-modality methods. [5] .
[0003] Researchers have explored cross-modal feature fusion and modality perception. Fang et al. [6] proposed a cross-modal feature fusion strategy based on self-attention mechanism, which improved the performance of multispectral target detection in aerial images by integrating rich information from different modalities. Sun et al. [7] Develop an uncertainty-aware cross-modal vehicle detection framework that fuses visible light and thermal infrared image data and reduces the detection error caused by uncertain targets through illumination quantization, thereby achieving robust target detection in complex environments.
[0004] Although many multimodal object detection methods have been proposed by researchers, few consider online multimodal object counting methods for aerial images. A practical problem is that complex multimodal models are difficult to directly deploy on drones with limited capabilities, especially consumer electronics drones. The Dynamic Neural Network (DyNN) architecture is a potential solution. Typical solutions include early exit [8] , Skip Layer [9] and many experts
[10] etc. Drone online counting has a natural dependence on multimodality and DyNN. The combination of the two forms a multimodal DyNN method that can adapt to the diversity of detection environments and instances, and balance the accuracy and inference cost of online counting.
[0005] Despite its many advantages, designing a low-cost and highly available multimodal dynamic neural network counting model faces unique challenges due to the cluttered background, large object scale variations and uneven distribution in drone aerial images, as well as limited computing resources:
[0006] 1) Lightweight multimodal feature alignment. Since visible light and thermal infrared images are inconsistent in spatial distribution, it is not advisable to blindly design a feature extraction network, which will make the extracted features uninterpretable. For multimodal feature learning, a common method is to use a two-stream network structure.
[11] , that is, using two parallel network structures to extract features separately. However, the generated model has a large number of parameters and is difficult to deploy on drones with limited resources. In addition, the two modal information is independent of each other. The dual-stream structure does not fully exploit the complementarity of the modalities and ignores the parameters shared between the two modalities.
[0007] 2) Environmental adaptation in dynamic deployment. During the mission execution, the importance of different modal data for drones changes dynamically over time and space. In other words, the demand for different modal data for drone online counting changes dynamically. For example, when there is sufficient light during the day, the model decodes two modal information, visible light images and thermal infrared images. Although the counting accuracy is optimized, the inference cost of the model is doubled. For resource-constrained drones, it is not worth it to exchange a higher inference cost for a slight improvement in accuracy. Pruning is a common method for lightweight models. However, although static pruning helps to reduce the inference cost, the loss of detection accuracy is inevitable. Considering the diversity of the environment during the drone's mission execution, the multimodal model must have the ability to self-regulate to achieve the best trade-off between counting performance and inference cost.
[0008] 3) Robustness optimization of multimodal models. During dynamic deployment, the multimodal data collected by airborne sensors will inevitably degrade in quality or even be lost. In the image pair generation framework based on the CycleGAN
[12] In this case, the visible light data is converted into thermal infrared data by a reversible neural network. These data and the visible light data are used as inputs to the model for object counting. Since the loss function does not constrain local details, it is difficult for CycleGAN to focus on local patterns and texture features. Summary of the invention
[0009] In order to support online target counting of UAV aerial images in complex scenarios such as insufficient lighting, rainy, foggy and haze, and night operations, the present invention proposes a framework based on a multimodal dynamic neural network DyNN for target counting in UAV aerial images, taking into account the adaptability, robustness and real-time performance of the model.
[0010] An online target counting method for unmanned aerial vehicles based on a multimodal dynamic neural network comprises the following steps: firstly, the unmanned aerial vehicle uses an airborne visible light camera and a thermal infrared camera to respectively take aerial photos of a visible light image and a thermal infrared image; then, the images are segmented based on scale estimation; the segmented images are used as inputs of a target counting model to perform target counting;
[0011] The target counting model is a target counting model based on a multimodal dynamic neural network DyNN. The architecture of the target counting model includes a feature extraction network, a decoder, a multimodal gating network, and a modality conversion network. In the target counting model:
[0012] S1. When a thermal infrared image cannot be obtained, the visible light image is converted into a thermal infrared image by a modal conversion network, and the generated thermal infrared image is provided to the thermal infrared modal branch of the target counting model; (the visible light image and the generated thermal infrared image are used together for counting in the target counting model)
[0013] S2, multi-modal gating analysis of the scene illumination of the image; if the illumination is good, only the visible light modality branch of the target counting model is selected; if the illumination is insufficient, the visible light modality branch and the thermal infrared modality branch are selected, and the thermal infrared image is used as auxiliary data;
[0014] S3, the feature extraction network extracts the modal features of the selected modal branch and fuses the features of the multimodal reasoning flow;
[0015] S4. The features obtained in step S3 are passed through the decoder to obtain the output of the target counting model.
[0016] The main contributions of the present invention include:
[0017] Lightweight multimodal feature fusion. The modal information is separated by an additional modality extractor and concatenated to the shared features. Between the feature representations of different modalities, the Conditional Generative Adversarial Network (CGAN) is used to establish the complementarity of the two modal features, promote model learning and capture the consistency information between the modalities.
[0018] Decoder-assisted multimodal gating. The decoder provides training data for multimodal gating. Multimodal gating is used to guide the opening and closing of the thermal infrared branch network. The decoder upsamples the features to restore the resolution and generates high-quality density maps from different modalities. The training of multimodal gating is based on the comparison results of density maps.
[0019] Density-Map Modal Transformation (DMMT) with Density Map Loss. In adversarial learning, thermal infrared images and real thermal infrared images tend to be similar in spatial distribution based on density map counts. The model gradually acquires the ability to generate high-quality thermal infrared data using visible light data to make up for the lack of (thermal infrared) modal data.
[0020] Experiments on multiple authoritative datasets (including DroneRGBT and RGBT-CC) show that the proposed model is adaptive. Even when the modal data input is unstable, the proposed method can still achieve good counting accuracy and maintain low inference latency. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1(a) and Figure 1(b) respectively show aerial images in dim scenes, where:
[0022] Figure 1(a) is a visible light image.
[0023] Figure 1(b) is a thermal infrared image;
[0024] Figure 2 Represents a framework / model for object counting in UAV images based on multimodal dynamic neural networks;
[0025] Figure 3(a) and Figure 3(b) respectively show the decision of multi-modal gating under different lighting environments, where:
[0026] Figure 3(a) shows the case with good lighting (PRRGB=97, PRfinal=106).
[0027] Figure 3(b) shows the case of insufficient illumination (PRRGB=4, PRfinal=16);
[0028] Figure 4 Represents the modal conversion framework based on DMMT. Explanation of terms in the figure:
[0029] RGB image-visible light image, Fake RGB image-(generated) fake visible light image,
[0030] Infrared image-infrared image, Fake Infrared image-(generated) fake infrared image,
[0031] Predicted density-Map–predicted density map, DM loss-density map loss,
[0032] Pre-trained Counter (Shared and Frozen)-pre-trained target counting model,
[0033] Cycle loss-cycle consistency loss, GAN loss-adversarial loss,
[0034] G R抗性 -The process of generating thermal infrared images (Thermal Infrared) from visible light images (RGB),
[0035] G T过程 -The process of generating visible light images from thermal infrared images,
[0036] D R -Visible light image discrimination process, D T -Thermal infrared image recognition process;
[0037] Figure 5 Represents the mean absolute error MAE and mean square error MSE for different k values;
[0038] Figure 6 Represents MAE and MSE for different Q values (large, medium, and small scales);
[0039] Figure 7 shows the visual analysis of three samples of the DroneRGBT dataset (compared with the unimodal model), where:
[0040] Figure 7(a) is a visible light image.
[0041] Figure 7(b) is a thermal infrared image.
[0042] Figure 7(c) is the true value density map.
[0043] Figure 7(d) is the predicted density map of comparison method-4.
[0044] Figure 7(e) is the predicted density map of the proposed method-1;
[0045] Figure 8 shows the visual analysis of three samples of the DroneRGBT dataset (compared with the multimodal model), where:
[0046] Figure 8(a) is a visible light image.
[0047] Figure 8(b) is a thermal infrared image.
[0048] Figure 8(c) is the true value density map.
[0049] Figure 8(d) is the predicted density map of the proposed method-2.
[0050] Figure 8(e) is the predicted density map of comparison method-8.
[0051] Figure 8(f) is the predicted density map of the proposed method-1;
[0052] Figure 9(a) to Figure 9(d) Represents the visual analysis of three samples of the RGBT-CC dataset, where:
[0053] Figure 9(a) is a visible light image.
[0054] Figure 9(b) is a thermal infrared image.
[0055] Figure 9(c) is the true value density map.
[0056] Figure 9(d) is the predicted density map of the proposed method-1;
[0057] Figure 10(a) to Figure 10(c) is the modal conversion effect diagram, where:
[0058] Figure 10(a) is a visible light image.
[0059] Figure 10(b) is a real thermal infrared image.
[0060] Figure 10(c) is the generated thermal infrared image. DETAILED DESCRIPTION
[0061] 1 Overview
[0062] Taking into account the insufficient counting ability of UAVs in complex environments (e.g., weak light, bad weather), the present invention designs an online target counting method for UAVs based on a multimodal dynamic neural network.
[0063] Modal information is fused through conditional adversarial generation to enhance multimodal feature expression and reduce redundant computation of feature extraction.
[0064] When scheduling modal branches, the accuracy of each modal density map is taken into account to optimize the judgment of multi-modal gating. Modal conversion is a cycle generative adversarial network CycleGAN that introduces density map loss to solve the problem of missing modal data.
[0065] Experiments on two authoritative datasets show that the method achieves 8.6%–19.5% lower mean absolute error and 79% lower average inference latency compared to CMCRL.
[0066] 2. Design
[0067] like Figure 2As shown, the counting model based on multimodal DyNN proposed in the present invention mainly includes the following three networks:
[0068] 1) The feature extraction network of modal fusion is used to extract and fuse multimodal information. The discriminator is used to supervise the generator to learn modal consistency information to achieve the purpose of modal feature fusion.
[0069] 2) The decoder-assisted multimodal gating network is responsible for scheduling modal information in complex scenes. When the lighting conditions in the input image scene are good, the thermal infrared branch is turned off and only the visible light modality is used; when the image is relatively dim, the thermal infrared branch is activated to assist in counting.
[0070] 3) A modal conversion network based on density map loss is used to make up for the lack of modal information. If thermal infrared data cannot be obtained, thermal infrared data is generated from visible light data to enhance the counting performance under dark light conditions.
[0071] The main workflows in the proposed framework include:
[0072] a. When thermal infrared images are not available, the modal conversion network based on density map loss provides the thermal infrared branch with the function of generating thermal infrared images.
[0073] b. Decoder-assisted multimodal gating analyzes scene illumination and selects necessary modal branches.
[0074] c. The feature extraction network for modal fusion extracts the modal features of the selected branch and fuses the features of the multimodal reasoning stream.
[0075] In this example, the image input to the object counting model is a segmented image based on scale estimation. Here, the scale estimation is based on the flight altitude and elevation angle of the drone.
[0076] 2.1 Feature Extraction Network for Modal Fusion
[0077] like Figure 2 As shown in the figure, multimodal feature fusion relies on a conditional generative adversarial network (Conditional GAN, CGAN), which contains a pair of generators and discriminators. The former generates modality-related features, which enter the discriminator after a scale-aware convolution operation. The latter determines the modality to which the features generated by the generator belong. In the adversarial process, the discriminator can align the features of the optical image with the features of the thermal infrared image to the same domain, so that the generator has the ability to express modality consistency. The generator and discriminator are introduced below.
[0078] A. Generator
[0079] Since the visible light image and the thermal infrared image are not consistent in spatial distribution, the present invention must reconstruct the feature extraction method.
[0080] Unlike the two-stream network with a large number of parameters, the present invention uses a universal feature extractor to learn shared parameters. Mboilenet v3 is used to build the backbone network. Two low-parameter modal extractors are added to the backbone network to extract modal features of visible light and thermal infrared data respectively. The role of the modal extractor is to enhance the features extracted from the two modalities so that the mined shared features do not lose model information. After the backbone network extracts the modal consistency features, they are added pixel by pixel with the features of the modal extractor. Taking a thermal infrared data as an example, its features are calculated as
[0081]
[0082] Among them, F g and They represent the general feature extractor and modal feature extractor for thermal infrared data respectively. and represents the parameters of the relevant extractor. After the convolution function is introduced, formula (1) is simplified to
[0083]
[0084] Among them, W g represents the parameters of the general feature extractor, represents the parameters of the modal feature extractor, and the convolution operation is denoted by ×. The new matrix generated, namely M T , representing the parameters related to the mode.
[0085] B. Discriminator
[0086] The object counting based on adversarial learning mainly focuses on the game between the predicted count density map and the real density map. In contrast, the present invention constructs an adversarial learning framework for the feature representation of visible light and thermal infrared modalities, so that the features of these two modalities can complement each other and achieve alignment between the modalities.
[0087] The features of the visible light mode and thermal infrared mode output by the generator are represented as F RGB and F T After passing through the 1×1 convolution layer, they are transformed into the corresponding feature map S through the Sigmoid function. RGB and S T Subsequently, these feature maps are combined with the RGB image (visible light image) into a 4-channel feature map and passed to the discriminator to determine the original modality type of the input sample. The generator can learn the consistent expression between multimodal features. The discriminator uses a SalGAN
[13] network structure.
[0088] C. Adversarial Training
[0089] k-root optimized Euclidean loss and adversarial loss are used for model training. The features F of the two modalities involved in adversarial learning are RGB and F T are mapped to the same feature space, and the discriminator's discriminant result is obtained through the convolution layer and the Sigmoid function. When training the discriminator, the generator is locked. When training the generator, the parameters of the generator are updated, and at the same time, the loss of the discriminator is back-propagated by the gradient. It is worth noting that the discriminator here is only used in the training phase to support adversarial learning between modal features; that is, only the generator and decoder are used in the inference phase. The loss function of the discriminator is expressed as
[0090] L D =L(D(data RGB ,S RGB ),1)+L(D(data RGB ,S T ),0) (3)
[0091] Where data RGB represents the true label used as auxiliary information in CGAN, i.e., the original visible light image. 1 and 0 represent the corresponding classification of visible light and thermal infrared modalities, respectively. The overall loss of the final density map obtained by training the model is expressed as
[0092]
[0093] 2.2 Decoder-Assisted Multimodal Gating Network
[0094] The key to high-precision and low-redundancy counting is to adaptively schedule modal branches according to input instances. To this end, the present invention proposes a decoder-assisted multimodal gating network. For instances captured in dim light, dust, etc., the method achieves the best match of multiple modal information by dynamically calling modal information. The method is jointly constructed by a decoder and multimodal gating. Multimodal gating is a binary classification neural network used to decide whether to open or close branches. The output result of the decoder is used as the basis for multimodal gating training. The decoder and multimodal gating are introduced below.
[0095] A.Decoder
[0096] The input original image is downsampled by multiple convolutions. Although deep semantic information can be extracted, the target details will be lost. In response to this, the features are upsampled and the resolution is restored to generate a high-quality density map. The decoder first upsamples the features three times to restore the image to its initial size. Then, the density map branches of the two modalities are separated. These two branches can output density maps of visible light and thermal infrared modes respectively. Finally, a convolution operation is performed on the two spliced density maps to obtain a density map with multimodal (two modal) information fusion.
[0097] Three density maps (density map PR generated by visible light image RGB , density map PR generated by thermal infrared image T The density map PR after the fusion of the two modal image information final ) determines the label value of the multimodal gating during the training process (detailed introduction later).
[0098] When only visible light modality data is used, PR RGB As the final output of the network;
[0099] When using multimodal data, PR RGB and PR T Further fusion generates a more accurate density map.
[0100] The decoder provides flexible options for density map output to adapt to the decision results of the previous multimodal gating.
[0101] B. Multimodal Gating
[0102] The input of multimodal gating is the segmented visible light image. The features extracted by the shallow convolution are fed into the fully connected layer, which outputs the confidence of the modality selection and guides the opening and closing of the thermal infrared branch. The convolution uses the first block of Resnet50 to analyze the light intensity in the image. The label during training is set to
[0103]
[0104] Among them, PR RGB Represents the density map generated by the visible light image, PR final Density map representing the fusion of visible light and thermal infrared information. Represents the threshold value of the label.
[0105] When the visible light and thermal infrared images pass through the entire network, PR RGB , PR T and PR final According to formula (5), label p = 1 means that the thermal infrared modal branch can detect more targets that are not identified by the visible light branch.RGB There are large shadow areas in the image, and the thermal infrared branch is used to enhance the visible light branch. The label p=0 means "turn off the thermal infrared branch".
[0106] When the UAV is in a well-lit environment (see Figure 3(a)), a relatively accurate counting result can be obtained using only visible light data. Therefore, the UAV tends to choose visible light unimodal data. When the UAV is in an environment with insufficient lighting (see Figure 3(b)), thermal infrared information is added to the reasoning process as auxiliary data.
[0107] 2.3 Density-Map Modal Transformation (DMMT) based on density map loss
[0108] Although thermal infrared images are essential in dim environments, drones may not be able to collect them in some environments. In this case, visible light data is used to generate thermal infrared data (i.e., "pseudo thermal infrared images"), and the visible light data and the generated thermal infrared data are input into the target counting model for counting.
[0109] CycleGAN is often used to solve the problem of modality conversion, but considering that the loss function cannot constrain local details (local patterns and texture features), this paper proposes a density map loss-based modality conversion network (DMMT) to generate thermal infrared images. The purpose is to shorten the spatial distance between the generated (pseudo) thermal infrared image and the real thermal infrared image for target counting as much as possible.
[0110] In the modality conversion process, the pre-trained object counting model is regarded as a spatial mapping converter. The pseudo thermal infrared images and real thermal infrared images generated by CycleGAN are converted into density map space respectively (density map space refers to the density distribution of certain features (such as objects, people, materials, etc.) in space represented by density maps). The density map mean square loss is used in the additionally introduced density map space to force the generated image to be close to the real image. This additional loss forces CycleGAN to focus on the objects in the image.
[0111] The DMMT framework mainly consists of a CycleGAN and a pre-trained target counter model, such as Figure 4 shown.
[0112] The generator and discriminator of CycleGAN are denoted as G and D respectively. R and T represent the visible light and thermal infrared domains respectively, and r and t represent the data in the visible light domain and the data in the thermal infrared domain. GAN and cycle consistency loss L cycle Expressed as
[14]
[0113]
[0114] and
[0115]
[0116] The mean square loss of the density map is expressed as
[0117]
[0118] where ψ represents the spatial mapping performed by the pre-trained count model and M(·,·) represents the mean squared error between the generated image and the real image in the density map space.
[0119] Based on equations (6), (7) and (8), the loss function of DMG (density management graphic) is expressed as
[0120]
[0121] Among them, λ and μ are the weights of cycle consistency loss and density map loss, respectively.
[0122] 3 Experimental design and result analysis
[0123] In order to improve the reference value of the experimental results, two authoritative datasets containing complex scenes were selected.
[15] It is a multimodal crowd counting dataset collected from the perspective of drones. A total of 3,600 images were collected, covering a variety of complex scenes from different altitudes, lighting conditions to crowd density. It contains images taken at three flight altitudes, which can be used to verify the effectiveness of the proposed multi-scale receptive field strategy. In terms of lighting conditions, the image pairs cover a variety of conditions such as dark, dusk, and bright, which can simulate various complex environments encountered by drones in dynamic deployment. In terms of density, the density of the dataset ranges widely, ranging from 1 to 401. DroneRGBT images are collected by drones (DJIPhantom 4) and cover complex backgrounds such as parks, roads, streets, squares, subways, and shopping malls. RGBT
[16] It is a large-scale crowd counting dataset, which contains 2030 pairs of 640×480 resolution RGB thermal images taken in various scenes. The dataset covers a variety of scenes, such as shopping malls, streets, playgrounds, train stations and subway stations. Among them, 1013 pairs of images were collected under bright lighting conditions, while the other 1017 pairs were collected under dark lighting conditions. Compared with some datasets from the Internet, the crowd density distribution of RGBT-CC is closer to real urban scenes.
[0124] The Adam optimizer is used for density map generation training, where the learning rate is set to 1×10 -5 The adversarial loss is used to alternately optimize the alignment model and the decision model to align the visible light and thermal infrared features. The learning rate of the Adam optimizer in the adversarial training phase is set to 1×10 -5 These settings help ensure the stability and convergence speed of the model, taking into account the accuracy and robustness of crowd counting. λ and μ in formula (7) are set to 0.1 and 0.1 respectively.
[0125] The comparative experiments on the DroneRGBT dataset aim to determine the value of k in the distribution-aware block loss. Figure 5 As shown in Figure 2, since MAE and MSE have the best results when k = 3, this value is used in subsequent experiments. I obtained the best settings for Q at various scales through experiments on DroneRGBT, which are summarized in Figure 6 When MAE and MSE reach their minimum values, Q = (1, 2, 3), (3, 4, 5), (5, 6, 7).
[0126] The parameter settings of the discriminator are shown in Table 1, which mainly include 6 layers of 3×3 convolution operations Conv1~Conv6, using ReLU activation function after each convolution layer, and using a maximum pooling operation after every two convolution operations. Fc7~Fc9 represent 3 layers of fully connected layers, and the activation functions used are Tanh and Sigmoid respectively.
[0127] Table 1: Composition of the discriminator
[0128] Table 1: Composition of discriminator
[0129]
[0130] This experiment divides the proposed methods into the following six categories to observe the counting performance under different inference flow modes:
[0131] Proposed method-1: Dynamic reasoning using data from visible light and real thermal infrared modalities;
[0132] Proposed method-2: using only visible light modality data for inference;
[0133] Proposed method-3: using only data from the real thermal infrared modality for reasoning;
[0134] Proposed method-4: using both visible light and real thermal infrared modality data for reasoning;
[0135] Proposed method-5: Dynamic reasoning using data from visible light and thermal infrared modalities;
[0136] Proposed method-6: Two-stream structure, using data from visible light and real thermal infrared modalities for dynamic reasoning.
[0137] In order to verify the environmental adaptability and counting performance of the proposed method, this experiment selected some existing mainstream single-modal and multi-modal models as benchmark methods:
[0138] Comparison method-1 (derived from MCNN
[17] : A lightweight single-modal multi-column model that uses multiple sets of convolutional kernels to cover a wide range of receptive fields;
[0139] ·Comparison Method-2 (derived from SANET
[18] ): Introducing a scale aggregation module in the encoder to improve the representation ability and scale diversity of features;
[0140] ·Comparison method-3 (from CSRNET
[19] ): Encoding deeper features in crowded scenes using dilated convolutions;
[0141] ·Comparison method-4 (derived from CANNET
[20] ): A robust model that uses detail information in features as attention to guide scale fusion;
[0142] ·Comparison method-5 (derived from BL
[21] ): Use the loss function Bayesian Loss to supervise the expected value of the count of each point annotation;
[0143] ·Comparison Method-6 (from AAVCNET
[22] ): Optimizing counting accuracy and inference latency of drone images through spatial information and multi-scale receptive fields;
[0144] ·Comparison Method-7 (from RMMCC
[23] ): Add counting constraints in the multimodal fusion process and use a multi-scale token transformer to perceive objects of different scales;
[0145] ·Comparison Method-8 (derived from CMCRL
[24] ): The framework integrates cross-modal collaborative representation learning and learns complementary information between different modalities through an information aggregation and distribution module.
[0146] 3.1 Comparative Experiments with Unimodal Model
[0147] Table 2: Comparison of counting performance with unimodal models on the DroneRGBT dataset
[0148] Table 2: Counting performance comparison with mono-modality models onDroneRGBT
[0149]
[0150]
[0151] In order to prove that the multimodal model proposed in this chapter is better than a single modality, the experiment tested the effect of the existing counting model on single modality data. These models were trained on single modality data and tested on the corresponding test sets. The experimental results are shown in Table 2. When using single modality data of visible light or thermal infrared for counting tasks, the proposed method-2 and the proposed method-3 are significantly better than the comparison method-1 in counting accuracy, and slightly inferior to the comparison method-4. The comparison results with the comparison method-6 show that after fusing multiple modality information, the proposed method still maintains similar single modality information understanding capabilities. When using thermal infrared data to assist visible light for target counting, the counting accuracy of the proposed method-1 surpasses all single modality models, which shows that the method designed in this chapter effectively understands and utilizes the information of targets in multiple modalities. By Figure 7(a) to Figure 7(e) From the density map visualization results in , we can see that the advantage of Method-4 is that the visible light part is more accurate, but it still cannot effectively identify objects in the dark light area. The proposed Method-1 achieves accurate counting for aerial images under various lighting conditions.
[0152] 3.2 Comparative Experiments with Multimodal Models
[0153] This section selects a representative multimodal counting model for comparison with the proposed method. The proposed method-6 is a two-stream network. Both of its branches use the first four blocks of Mobilenetnet V3 to extract features related to each modality. Then the feature map output by each branch is connected according to the channel, and then reduced in dimension by a 1×1 convolution kernel. Finally, the fused features are input into the regression module to obtain the final density map.
[0154] Table 3: Comparison of counting performance with multimodal models on DroneRGBT and RGBT-CC datasets
[0155] Table 3: Counting comparison with multi-modality models on DroneRGBTand RGBT-CC
[0156]
[0157] As shown in Table 3, in the DroneRGBT dataset, the MAE of the proposed method-1 is 7.26, which is close to 96.1% of the comparison method-7. However, it is lower than the proposed method-6 because the two-stream network improves the model understanding ability by increasing the number of parameters. The proposed method-1 achieves acceptable counting performance with a much smaller number of parameters than the proposed method-6. When running the multimodal model, thermal infrared information helps to distinguish potential targets from cluttered backgrounds, while visible light information helps to eliminate negative thermal objects in thermal infrared images and eliminate interference from similar objects. It can be seen that the model fully analyzes the key features of the target in the image and the modal information is efficiently utilized. In addition, in the proposed method, the processing of the two modal information is decoupled from each other, that is, it can work even when there is only visible light image. When the whole scene is well lit, the model can switch to single-modal reasoning mode to achieve high-precision counting at a lower cost.
[0158] The proportion of dark light areas is highly correlated with the inference latency of the model. Table 3 summarizes the average inference latency on the DroneRGBT dataset. It is worth noting that these average inference latencies are not linearly related to the number of modalities. The proposed multimodal gating can dynamically activate and mobilize the thermal infrared branch for dark light environments, thereby reducing the inference cost. Figure 8(a) to Figure 8(f) and Figure 9(a) to Figure 9(d) The visualization inference results of the proposed method on the DroneRGBT dataset and RGBT-CC dataset are shown in Figure 2. Since the features of visible light images are not obvious in dark environments, there are many false detections and missed detections in the detection results of the proposed method-2. The baseline method-referenced both thermal infrared and visible light images and captured most targets in dark environments. In contrast, the proposed method-1 can distinguish targets in various lighting environments and maintain low false detection and missed detection rates.
[0159] 3.2 Ablation Experiment
[0160] In order to verify the effectiveness of each network in the proposed framework, this section designs ablation experiments on the DroneRGBT dataset by module.
[0161] This experiment first verifies whether the feature extraction network of modal fusion helps improve counting accuracy. As shown in Table 4, CGAN is not used to extract features, and the data of the two modalities pass through the same backbone network. According to the results, it is easy to find that the MAE without using CGAN is significantly lower than that of the proposed method-1, because the interference between the features of different modal data hinders the correct expression of the features. This result indirectly confirms the unique and critical role of CGAN in understanding and extracting the common expression of each modality.
[0162] Table 4: Comparison of ablation results on the DroneRGBT dataset
[0163] Table 4: Comparison of results from ablation experiments on DroneRGBT
[0164]
[0165] Next, we examine the impact of multimodal gating on counting performance and inference latency. As shown in Table 4, when multimodal gating is running (proposed method-1), the MAE of the model is better than the global single-modal model (proposed method-2 and proposed method-3); when multimodal gating is removed (proposed method-4), the model always uses two modalities for counting. In this case, its MAE is only reduced by 2.1%, but it brings a lot of unnecessary calculations. It can be seen that multimodal gating plays an important role in maintaining the stability of counting and reducing the cost of counting inference.
[0166] Finally, the effectiveness of the mode conversion network is verified. Figure 10(a) to Figure 10(c) Compared with the real thermal infrared images, the generated thermal infrared images show the contour and brightness of the object. Table 4 shows the experimental results on the DroneRGBT dataset. When the real thermal infrared images in the dataset are all replaced by thermal infrared images generated by the modality conversion network (proposed method-5), the MAE is only improved by 9.4%, indicating that the modality conversion network can meet the model's demand for thermal infrared data when the modality is missing.
[0167] 4 Summary
[0168] This paper proposes a method for online target counting of drones based on multimodal dynamic neural network. This method integrates multimodal information through CGAN to improve the counting accuracy of drone multimodal images while maintaining a low computational load. Multimodal gating dynamically schedules the reasoning of modal branches to balance counting accuracy and reasoning latency. In addition, when only visible light data is available, the model can perform modal conversion through DMMT and generate thermal infrared images using visible light images.
[0169] Experimental results show that the proposed model can perform efficient counting in multiple modal input scenarios. The proposed method can cope with harsh natural environments and adapt to changes in the flight altitude, pitch angle and ambient light of the UAV itself. The proposed multimodal DyNN framework is universal and customizable, and can be further applied to aerial target detection, target segmentation and other scenarios.
[0170] References
[0171] [1]Iftikhar S,Asim M,Zhang Z,et al.Target detection and recognitionfor traffic congestion in smart cities using deep learning-enabled UAVs:Areview and analysis[J].Applied Sciences,2023,13(6):3995.
[0172] [2]Sun Y,Cao B,Zhu P,et al.Detfusion:A detection-driven infrared andvisible image fusion network[C] / / Proceedings of the 30th ACM internationalconference on multimedia.2022:4003-4011.
[0173] [3]Zhang,X.,Ye,P.,Peng,S.,Liu,J.,Xiao,G.:Dsiammft:An rgb-t fusiontracking method via dynamic siamese networks using multi-layer featurefusion.Signal Processing:Image Communication(2020)115756.
[0174] [4]Li X,Chen H,Li Y,et al.MAFusion:Multiscale attention network forinfrared and visible image fusion[J].IEEE Transactions on Instrumentation andMeasurement,2022,71:1-16.
[0175] [5]Zhao W,Xie S,Zhao F,et al.Metafusion:Infrared and visible imagefusion via meta-feature embedding from object detection[C] / / Proceedings ofthe IEEE / CVF Conference on Computer Vision and Pattern Recognition.2023:13955-13965.
[0176] [6]Qingyun F,Dapeng H,Zhaokui W.Cross-Modality Fusion Transformer forMultispectral Object Detection[J].2021.
[0177] [7]Sun Y,Cao B,Zhu P,et al.Drone-based RGB-infrared cross-modalityvehicle detection via uncertainty-aware learning[J].IEEE Transactions onCircuits and Systems for Video Technology,2022,32(10):6700-6713.
[0178] [8]G.Huang,D.Chen,T.Li,F.Wu,L.Van Der Maaten,and K.Q.Weinberger,Multi-scale dense networks for resource efficient image classification,arXivpreprint,2017,DOI 10.48550 / arXiv.1703.09844.
[0179] [9]A.Veit and S.Belongie,Convolutional networks with adaptiveinference graphs,Proc.of the European Conference on Computer Vision(ECCV),Munich,Germany),2018,pp.3–18.
[0180]
[10] D.Eigen,M.Ranzato,and I.Sutskever,Learning factoredrepresentations in a deep mixture of experts,arXiv preprint,2013,DOI10.48550 / arXiv.1312.4314.
[0181]
[11] Li G,Liu Z,Ling H.ICNet:Information conversion network for RGB-Dbased salient object detection.IEEE Trans.on Image Processing,2020,29:4873-4884.
[0182]
[12] Li F,Zha Y,Zhang L,et al.Information lossless multi-modal imagegeneration for RGB-T tracking[C] / / Chinese Conference on Pattern Recognitionand Computer Vision(PRCV).Cham:Springer Nature Switzerland,2022:671-683.
[0183]
[13] Pan J,Sayrol E,Nieto X G,et al.Salgan:Visual saliency predictionwith adversarial networks[C] / / CVPR scene understanding workshop(SUNw).2017.
[0184]
[14] Goodfellow I,Pouget-Abadie J,Mirza M,et al.Generative adversarialnetworks[J].Communications of the ACM,2020,63(11):139-144.
[0185]
[15] Peng T,Li Q,Zhu P.Rgb-t crowd counting from drone:A benchmark andmmccn network[C] / / Proceedings of the Asian conference on computervision.2020.
[0186]
[16] Liu Z,Wu W,Tan Y,et al.RGB-T Multi-Modal Crowd Counting Based onTransformer[J].arXiv preprint arXiv:2301.03033,2023.Shen H,Liu X,Shi N,etal.Blockchain-enabled solution for secure and scalable V2V video contentdissemination[J].Peer-to-Peer Networking and Applications,2022:1-17.
[0187]
[17] Yingying Zhang,Desen Zhou,Siqin Chen,Shenghua Gao,and Yi Ma,Single-image crowd counting via multi-column convolutional neural network,Proc.of the IEEE conference on computer vision and pattern recognition(2016),589–597.
[0188]
[18] Cao X,Wang Z,Zhao Y,et al.Scale aggregation network for accurateand efficient crowd counting[C].In Proceedings of the European Conference onComputer Vision(ECCV),2018:734–750.
[0189]
[19] Yuhong Li,Xiaofan Zhang,and Deming Chen,CSRNet:Dilatedconvolutional neural networks for understanding the highly congested scenes,Proc.of the IEEE conference on computer vision and pattern recognition(2018),1091–1100.
[0190]
[20] Liu W,Salzmann M,Fua P.Context-aware crowd counting[C].InProceedings of the IEEE Conference on Computer Vision and PatternRecognition,2019:5099–5108.
[0191]
[21] Ma Z,Wei X,Hong X,et al.Bayesian loss for crowd count estimationwith point supervision [C].In Proceedings of the IEEE InternationalConference on Computer Vision,2019:6142–6151.
[0192]
[22] Liu Y,Shen H,Wang T,et al.Vehicle counting in drone images:Anadaptive method with spatial attention and multiscale receptive fields[J].ETRI Journal,2024.
[0193]
[23] Liu Z,Wu W,Tan Y,et al.RGB-T Multi-Modal Crowd Counting Based onTransformer[J].arXiv preprint arXiv:2301.03033,2023.
[0194]
[24] Liu L,Chen J,Wu H,et al.Cross-Modal Collaborative RepresentationLearning and a Large-Scale RGBT Benchmark for Crowd Counting[J].2020.DOI:10.48550 / arXiv.2012.04529.
Claims
1. A method for online target counting of unmanned aerial vehicles based on a multimodal dynamic neural network, comprising: First, the drone uses the onboard visible light camera and thermal infrared camera to take aerial photos of visible light images and thermal infrared images respectively; Then, these images are segmented based on scale estimation; The segmented image is used as the input of the target counting model to perform target counting; Its characteristics are that the target counting model is a target counting model based on a multimodal dynamic neural network DyNN; the architecture of the target counting model includes a feature extraction network, a decoder, a multimodal gating network and a modality conversion network; in the target counting model: S1. When a thermal infrared image cannot be obtained, the visible light image is converted into a thermal infrared image by a modality conversion network, and the generated thermal infrared image is provided to the thermal infrared modality branch of the target counting model; S2, multi-modal gating analysis of the scene illumination of the image; if the illumination is good, only the visible light modality branch of the target counting model is selected; if the illumination is insufficient, the visible light modality branch and the thermal infrared modality branch are selected, and the thermal infrared image is used as auxiliary data; S3, the feature extraction network extracts the modal features of the selected modal branch and fuses the features of the multimodal reasoning flow; S4, the features obtained in step S3 are passed through a decoder to obtain the output of a target counting model; In step S1, the modality transfer network is a density map loss-based modality transfer network DMMT; the DMMT includes a cycle generative adversarial network CycleGAN and a pre-trained target counter model; in the modality transfer process, the pre-trained target count model is regarded as a space mapping converter; The thermal infrared images generated by CycleGAN and the real thermal infrared images are converted into density map space respectively; the density map mean square loss is used in the additionally introduced density map space to force the generated images to be close to the real images; The generator and discriminator of CycleGAN are represented as G and D respectively, R and T represent the visible light domain and thermal infrared domain respectively, r and t represent the data in the visible light domain and the data in the thermal infrared domain; Adversarial loss L GAN and cycle consistency loss L cycle Respectively expressed as: and The mean square loss of the density map is expressed as: Where ψ represents the spatial mapping performed by the pre-trained count model, and M(·,·) represents the mean square error between the generated image and the real image in the density map space; The loss function of density map generation DMG is expressed as: L final (G R→T ,G T→R ,D R ,D T ,R,T) =L GAN (G R→T ,D T ,R,T)+L GAN (G T→R ,D R ,T,R) +λL cycle (G R→T ,G T→R ,R,T)+μL DM (G R→T ,G T→R ,R,T) Among them, λ and μ are the weights of cycle consistency loss and density map loss, respectively.
2. The method for online target counting of unmanned aerial vehicles based on a multimodal dynamic neural network according to claim 1 is characterized in that in step S2, the multimodal gating network is a decoder-assisted multimodal gating network, wherein: The multimodal gating network is a binary classification neural network used to decide whether to open or close the modal branch. The output of the decoder is used as the basis for training the multimodal gating network. S2.
1. In the decoder, first, the features obtained in S3 are upsampled to restore the image to its initial size; then, the density map branches of the visible light modality and the thermal infrared modality are separated to obtain the density maps of the visible light modality and the thermal infrared modality; finally, the two density maps are convolved to obtain the density map of the fusion of the two modal information; The three density maps are represented as follows: Density map PR generated by visible light image RGB , density map PR generated by thermal infrared image T The density map PR after the fusion of the two modal image information final PR RGB , PR T and PR final The comparison results determine the label values of the multimodal gating network during the training process; When only visible light modality data is used, PR RGB As output; when using two modal data, PR final As output; S2.2, in the multimodal gating network, its input is the visible light image; the features extracted by the shallow convolution enter the fully connected layer, and the confidence of the modality selection is output to guide the opening and closing of the thermal infrared branch; the convolution uses the first block of Resnet50 to analyze the light intensity in the image; Labels of the multimodal gating network during training in, represents the threshold of the label; The label p = 1 means that the thermal infrared modality branch can detect enough targets that the visible light branch has not identified; If label p=0, the thermal infrared mode branch is closed.
3. The method for online target counting of unmanned aerial vehicles based on multimodal dynamic neural network according to claim 1 is characterized in that In step S3, the feature extraction network is a feature extraction network of modal fusion, and multimodal feature fusion relies on the generative adversarial network GAN, which includes a generator and a discriminator; The generator generates modality-related features, which enter the discriminator after scale-aware convolution operations; The discriminator determines the modality to which the features generated by the generator belong; During adversarial training, the discriminator can align the features of the light image and the thermal infrared image to the same domain, so that the generator has the ability to express modality consistency; The discriminator is only used in the training phase to support adversarial learning between modality features; The inference phase only uses the generator and decoder.
4. The method for online target counting of unmanned aerial vehicles based on multimodal dynamic neural network according to claim 3 is characterized by feature extraction In the network, GAN is a conditional generative adversarial network (CGAN).
5. The method for online target counting of unmanned aerial vehicles based on multimodal dynamic neural network according to claim 4, Its characteristics are in CGAN: S3.
1. Generator Mboilenet v3 is used to construct the backbone network; two low-parameter modal extractors are added to the backbone network to extract modal features of visible light image data and thermal infrared image data respectively; After the backbone network extracts the modality consistency features, it is added pixel by pixel with the features of the modality extractor; For any thermal infrared image data, its thermal infrared modal feature is calculated as Among them, F g and They represent the general feature extractor and modal feature extractor of thermal infrared image data respectively; θ g and They represent the parameters of the general feature extractor and the modal feature extractor respectively; After introducing the convolution function, the above formula is simplified to Among them, W g represents the parameters of the general feature extractor, Represents the parameters of the modal feature extractor, the convolution operation is denoted by ×, and the generated new matrix M T Represents parameters related to the mode; For any visible light image data, the same method as the feature extraction of thermal infrared image data is used to obtain the visible light modal feature F RGB ; S3.
1. Discriminator First, F RGB and F T After passing through the 1×1 convolution layer, they are transformed into the corresponding feature map S through the Sigmoid function. RGB and S T ; Then, S. RGB and S T The RGB image is merged into a 4-channel feature map and passed to the discriminator to determine the original modality type of the input sample; The discriminator adopts a network structure similar to SalGAN; S3.3 Adversarial Training k-root optimized Euclidean loss and adversarial loss were used for training; F participating in adversarial learning RGB and F T It is mapped to the same feature space and passes through the convolution layer and the Sigmoid function to obtain the discriminant result; When training the discriminator, the generator is locked; when training the generator, the parameters of the generator are updated, and at the same time, the loss of the discriminator is back-propagated by the gradient; The loss function of the discriminator is expressed as: L D =L(D(data RGB ,S RGB ),1)+L(D(data RGB ,S T ),0) Where data RGB represents the true label used as auxiliary information in CGAN, i.e., the original visible light image; 1 and 0 represent the corresponding classification of the visible light modality and thermal infrared modality, respectively; the overall loss of the final density map obtained by training is expressed as L=L(Θ)+L D 。
Citation Information
Patent Citations
Method for converting thermal infrared image into visible-light-like color image based on CycleGAN
CN114119356A
Multi-scene crowd counting method and device based on multi-modal fusion
CN116245845A
Control method and system of automobile electronic outside rear-view mirror and storage medium
CN117104140A
Object-level infrared-and-visible-light image fusion method based on fully convolutional neural network
WO2024174488A1
Cited By
Multi-source target detection method based on illumination perception and hybrid expert model
CN120088465A