Multi-source Object Detection Method Based on Light Perception and Mixture of Experts Model
Through the multi-source object detection method of light perception and hybrid expert model, the problem of insufficient RGB infrared multi-source object detection accuracy from the perspective of the drone is solved, and high-precision and robust detection effects are achieved, suitable for rapidly changing environments.
Patent Information
- Application Number
- CN202510542977.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The existing multi-source object detection methods have insufficient accuracy in complex application scenarios, especially when the light, morphology, shooting perspective and background changes rapidly from the perspective of the drone, the detection accuracy and robustness of RGB infrared multi-source object detection are insufficient.
Using a multi-source object detection method based on light perception and hybrid expert model, feature extraction is performed through a dual-flow backbone network, combined with the light adaptive hybrid expert dynamic feature decomposition combination module and the spatial pyramid pooling module, the light perception technology is used to accurately predict the lighting scene and make comprehensive decisions, improving detection accuracy and robustness.
It effectively improves the accuracy and scene robustness of RGB infrared multi-source target detection, can adapt to the rapid changes in lighting, morphology, shooting perspective and background from the perspective of the drone, and meets the real-time detection needs of drone tasks.
Smart Images

Figure CN120088465B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the technical field of artificial intelligence and deep learning, and particularly to a multi-source object detection method based on light perception and a mixture of experts model. Background Art
[0002] Object detection is a core task in computer vision technology, and its purpose is to identify and locate objects of interest in images or video frames. Object detection methods based on deep learning have currently been widely applied in fields such as security monitoring, intelligent transportation, intelligent retail, and industrial manufacturing. RGB images (visible light images) can provide rich texture information and detail information under good lighting conditions, but their imaging effects are poor at night or in bad weather. To improve the accuracy and robustness of object detection, research teams at home and abroad have focused on RGB-infrared multi-source object detection technology. RGB-infrared multi-source object detection technology makes full use of the complementary characteristics of two different modal images and can adapt to different environmental conditions and object characteristics.
[0003] To achieve the complementary fusion of multi-source features, a large number of attention mechanism modules have been designed in related research to improve object detection performance. Zhang L et al. proposed the CIAN network (Cross-modality interactive attention network), which uses a cross-modal interactive fusion module to fuse RGB features and infrared features and uses a context enhancement module to further enhance the network's ability to extract and transmit context information. Zhang H et al. designed an intra-modal attention module and an inter-modal attention module to enhance single-modal features and fuse complementary multi-modal features respectively. Yuan et al. proposed a network structure called Calibration and Complementary Transformer based on the powerful correlation modeling ability of Transformer, and calibrated modal features by calculating the cross-attention relationship between the RGB modality and the infrared modality.
[0004] However, most of the multi-source object detection methods proposed currently adopt a static network structure, which has the problem of insufficient accuracy in complex application scenarios. For example, in the RGB-infrared multi-source object detection task from the perspective of a drone, the objects captured by the drone change rapidly in terms of lighting, shape, shooting angle, and background. The feature performances of RGB samples and infrared samples are different at different times and in different scenarios. A single and fixed feature fusion strategy and a static network structure cannot adapt to the differences brought about by external environmental changes, which results in serious insufficient detection accuracy of RGB-infrared multi-source object detection. Summary of the Invention
[0005] In view of this, an embodiment of the present application proposes a multi-source object detection method based on light perception and a mixture of experts model. By determining the current light scene through light perception technology and then making comprehensive decisions using the corresponding mixture of experts model, the accuracy and scene robustness of RGB-infrared multi-source object detection are effectively improved.
[0006] To achieve the above object, an embodiment of the present application proposes a multi-source object detection method based on light perception and a mixture of experts model, and the method includes the following steps:
[0007] Use a two-stream backbone network to perform downsampling and multi-scale feature extraction on the input paired RGB images and infrared images, so as to obtain RGB features at different scales and infrared features ; where , corresponding to downsampling resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 respectively;
[0008] Input and into the first light-adaptive mixture of experts dynamic feature decomposition and combination module, and through light perception and multi-expert decision-making, obtain the recombined features and , input and into the second light-adaptive mixture of experts dynamic feature decomposition and combination module, and through light perception and multi-expert decision-making, obtain the recombined features and , perform channel splicing on and to obtain the spliced feature , input into the spatial pyramid pooling module to obtain the fused feature ;
[0009] Input , and into the RGB neck network for multi-path feature aggregation to generate the RGB aggregated feature , and , input , and into the infrared neck network for multi-path feature aggregation to generate the infrared aggregated feature , and ;
[0010] Input , and Input it into the RGB detection network for detection. After performing non-maximum suppression processing, the RGB detection result is obtained. Then, 、 and are input into the infrared detection network for detection. After performing non-maximum suppression processing, the infrared detection result is obtained.
[0011] To achieve the above object, an embodiment of the present application further provides an electronic device, where the electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, instructions executable by the at least one processor are stored in the memory, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute a multi-source target detection method based on light perception and a mixture of experts model as described above.
[0012] To achieve the above object, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which when executed by a processor, can implement a multi-source target detection method based on light perception and a mixture of experts model as described above.
[0013] The multi-source target detection method based on light perception and a mixture of experts model proposed by the embodiment of the present application uses a dual-stream backbone network, a first light-adaptive mixture of experts dynamic feature decomposition and combination module, a second light-adaptive mixture of experts dynamic feature decomposition and combination module, a spatial pyramid pooling module, an RGB neck network, an infrared neck network, an RGB detection network, and an infrared detection network to work together to achieve multi-source target detection. The light-adaptive mixture of experts dynamic feature decomposition and combination module integrates light perception technology and configures a mixture of experts model. Through the light perception technology, the currently belonging light scene can be accurately predicted, and then the mixture of experts model corresponding to the currently belonging light scene is used for comprehensive decision-making, effectively improving the accuracy and scene robustness of RGB-infrared multi-source target detection. The method proposed in the present application is particularly applicable to the RGB-infrared multi-source target detection task from the perspective of an unmanned aerial vehicle with rapid changes in aspects such as light, morphology, shooting angle, and background, and can well adapt to the differences brought about by external environmental changes, thereby providing a reliable basic support for unmanned aerial vehicle tasks.
[0014] Optionally, the dual-stream backbone network is composed of an RGB feature extraction branch and an infrared feature extraction branch;
[0015] The RGB feature extraction branch consists of five sequentially connected visible light convolution modules. The input of the first visible light convolution module is the RGB image, and the input of the next visible light convolution module is the output of the previous visible light convolution module. The five visible light convolution modules are used to downsample and extract features of their own inputs, and obtain RGB features with resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the RGB image, respectively. The five RGB features are recorded in the order of the visible light convolution modules as , , , and ;
[0016] The infrared feature extraction branch consists of five sequentially connected thermal infrared convolution modules. The input of the first thermal infrared convolution module is the infrared image, and the input of the next thermal infrared convolution module is the output of the previous thermal infrared convolution module. The five thermal infrared convolution modules are used to downsample and extract features of their own inputs, and obtain infrared features with resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the infrared image, respectively. The five infrared features are recorded in the order of the thermal infrared convolution modules as , , , and .
[0017] Optionally, the network structure of the first illumination adaptive hybrid expert dynamic feature decomposition combination module and the second illumination adaptive hybrid expert dynamic feature decomposition combination module is the same, both of which are composed of an illumination perception network and three sparse gated hybrid expert models, wherein the three sparse gated hybrid expert models are daytime sparse gated hybrid expert models, , Night sparse gated hybrid expert model and Dark Night Sparse Gated Hybrid Expert Model ;
[0018] The illumination perception network of the first illumination adaptive hybrid expert dynamic feature decomposition combination module receives Afterwards, based on Predict image illumination values , and get the maximum illumination value index ,based on Determine the current lighting scene, and start the corresponding sparse gated hybrid expert model based on the current lighting scene;
[0019] The sparse gated hybrid expert model activated by the first illumination adaptive hybrid expert dynamic feature decomposition combination module will and Perform channel splicing to obtain the spliced features , the gating weights are then calculated using a noisy Top-K sparse gating network, and the corresponding Top-K expert networks are activated for comprehensive decision-making. Finally, the gating weights and the outputs of the expert networks involved in the calculation are weighted to obtain the recombined features and ;
[0020] After the light perception network of the second light-adaptive hybrid expert dynamic feature decomposition and combination module receives , based on , the image light value is predicted, and at the same time, the maximum light value index is obtained. Based on , the current belonging light scene is judged, and the corresponding sparse gating hybrid expert model is started based on the current belonging light scene; among them, the current belonging light scene is daytime, night or dark night;
[0021] The started sparse gating hybrid expert model of the second light-adaptive hybrid expert dynamic feature decomposition and combination module will and are concatenated in channels to obtain the concatenated features , the gating weights are then calculated using a noisy Top-K sparse gating network, and the corresponding Top-K expert networks are activated for comprehensive decision-making. Finally, the gating weights and the outputs of the expert networks involved in the calculation are weighted to obtain the recombined features and .
[0022] Optionally, the light perception network consists of a first convolutional module, a max pooling module, a second convolutional module, an adaptive average pooling module, a first linear module, a second linear module, a linear layer, and a ReLU activation function layer. The network structures of the first convolutional module and the second convolutional module are the same, both consisting of a 1×1 convolutional layer with a stride of 1, a batch normalization layer, and a ReLU activation function layer. The network structures of the first linear module and the second linear module are the same, both consisting of a linear layer, a ReLU activation function layer, and Droupout;
[0023] The input of the light perception network is successively subjected to channel dimensionality reduction and spatial downsampling through the first convolutional module, the max pooling module, and the second convolutional module, then subjected to adaptive average pooling by the adaptive average pooling module, and then the light is predicted by the first linear module, the second linear module, and the linear layer. Finally, the predicted image light value and the maximum light value index are output by the ReLU activation function layer; among them, .
[0024] Optionally, each expert network in the sparse gating mixture-of-experts model adopts an independent feature decomposition and combination module, which is composed of a decomposition unit, four feature extractors with the same network structure, and a multi-source feature recombination unit. The four feature extractors are a visible light basic feature extractor, a visible light detail feature extractor, a thermal infrared basic feature extractor, and a thermal infrared detail feature extractor;
[0025] Input to the expert network is first decomposed by the decomposition unit into and , and the visible light basic feature extractor and the visible light detail feature extractor respectively extract features from to obtain the RGB basic feature and the RGB detail feature . The thermal infrared basic feature extractor and the thermal infrared detail feature extractor respectively extract features from to obtain the infrared basic feature and the infrared detail feature ; where ;
[0026] The multi-source feature recombination unit performs multi-source feature recombination based on , and to obtain , and at the same time performs multi-source feature recombination based on , and to obtain .
[0027] Optionally, the feature extractor is composed of a first 1×1 convolutional unit, a channel splitting unit, a first 3×3 convolutional unit, a second 3×3 convolutional unit, a channel concatenation unit, and a second 1×1 convolutional unit. The first 1×1 convolutional unit and the second 1×1 convolutional unit have the same network structure, both consisting of a 1×1 convolutional layer with a stride of 1, a batch normalization layer, and a SiLU activation function layer. The first 3×3 convolutional unit and the second 3×3 convolutional unit have the same network structure, both consisting of a 3×3 convolutional layer with a stride of 1, a batch normalization layer, and a SiLU activation function layer;
[0028] Let the input feature of the feature extractor be , After being processed by the first 1×1 convolutional unit, it is split by the channel splitting unit along the channel dimension into and , After being processed by the first 3×3 convolutional unit and the second 3×3 convolutional unit, it is concatenated by the channel concatenation unit with and Concatenate along the channel dimension to obtain the concatenated feature , Reduce the dimension after being processed by the second 1×1 convolution unit to obtain the output feature .
[0029] Optionally, the network structures of the RGB neck network and the infrared neck network are both top-down and bottom-up multi-path aggregation networks;
[0030] In the RGB neck network, the top-down path is composed of an ordinary PAN module, a first visible light PAN module, and a second visible light PAN module connected in sequence. The bottom-up path is composed of a third visible light PAN module and a fourth visible light PAN module connected in sequence. The output of the ordinary PAN module is also used as the input of the fourth visible light PAN module, and the output of the first visible light PAN module is also used as the input of the third visible light PAN module. , and are respectively input into the ordinary PAN module, the first visible light PAN module, and the second visible light PAN module. The second visible light PAN module, the third visible light PAN module, and the fourth visible light PAN module respectively output three different scales of RGB aggregation features 、 and ;
[0031] In the infrared neck network, the top-down path is composed of an ordinary PAN module, a first thermal infrared PAN module, and a second thermal infrared PAN module connected in sequence. The bottom-up path is composed of a third thermal infrared PAN module and a fourth thermal infrared PAN module connected in sequence. The output of the ordinary PAN module is also used as the input of the fourth thermal infrared PAN module, and the output of the first thermal infrared PAN module is also used as the input of the third thermal infrared PAN module. , and are respectively input into the ordinary PAN module, the first thermal infrared PAN module, and the second thermal infrared PAN module. The second thermal infrared PAN module, the third thermal infrared PAN module, and the fourth thermal infrared PAN module respectively output three different scales of infrared aggregation features 、 and .
[0032] Optionally, the dual-stream backbone network, the first light-adaptive mixture-of-experts dynamic feature decomposition combination module, the second light-adaptive mixture-of-experts dynamic feature decomposition combination module, the spatial pyramid pooling module, the RGB neck network, the infrared neck network, the RGB detection network, and the infrared detection network together constitute a multi-source object detection model. The overall loss function used during the training of the multi-source object detection model is expressed by the formula:
[0033] ;
[0034] where, represents the overall loss function used during the training of the multi-source object detection model, represents the classification loss, represents the th category, represents the regression loss, represents the distribution focus loss, represents the multi-source feature decomposition loss, which is used to increase the correlation between basic features and reduce the correlation between detailed features, represents the light perception loss, which is calculated based on and obtained, represents the importance loss, which is used to measure the dispersion of the batch weights of each expert network to ensure that each expert network is equally important, represents the load balancing loss, which is used to avoid the imbalance in the number of samples assigned to each expert network, , , , are respectively the loss weights of the classification loss, the regression loss, the distribution focus loss, and the multi-source feature decomposition loss, is the binary cross entropy loss, is the CIoU loss, is the cross entropy loss. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the related art, the following will briefly introduce the drawings required for the description of the embodiments of the present application or the related art. The following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. The drawings described here are only used to explain the present application and are not used to limit the present application.
[0036] Figure 1 is provided in an embodiment of the present application, which is a flowchart of a multi-source object detection method based on light perception and a mixture-of-experts model;
[0037] Figure 2 is a schematic structural diagram of a multi-source target detection model provided in an embodiment of the present application;
[0038] Figure 3 is a schematic structural diagram of a light-adaptive mixture-of-experts dynamic feature decomposition and combination module provided in an embodiment of the present application;
[0039] Figure 4 is a schematic structural diagram of a light perception network provided in an embodiment of the present application;
[0040] Figure 5 is a schematic structural diagram of a sparse gating mixture-of-experts model provided in an embodiment of the present application;
[0041] Figure 6 is a schematic structural diagram of a feature extractor provided in an embodiment of the present application;
[0042] Figure 7 is a schematic diagram of the experimental results of a comparative experiment conducted on the Drone Vehicle dataset provided in an embodiment of the present application;
[0043] Figure 8 is a schematic structural diagram of an electronic device provided in another embodiment of the present application. Detailed implementation manners
[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below with reference to the accompanying drawings. Those of ordinary skill in the art can understand that in the embodiments of the present application, many technical details are provided to help readers better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented. The following division of the embodiments is only for convenience of description and should not constitute any limitation on the specific implementation manners of the present application. The embodiments can be combined with and cited from each other on the premise of not being contradictory.
[0045] To solve the problem that the existing multi-source target detection methods cannot adapt to the differences brought by changes in the external environment, an embodiment of the present application proposes a multi-source target detection method based on light perception and a mixture-of-experts model, which is applied to an electronic device. The electronic device can be a terminal or a server. In this embodiment and the following embodiments, the electronic device is described by taking the server as an example. The implementation details of the multi-source target detection method based on light perception and a mixture-of-experts model proposed in this embodiment are specifically described below. The following content is only the implementation details provided for convenient understanding and is not necessary for implementing this solution.
[0046] The specific process of a multi-source target detection method based on light perception and a mixture of experts model proposed in this embodiment can be as follows Figure 1 shown, including:
[0047] Step 101, use a two-stream backbone network to downsample and perform multi-scale feature extraction on the input paired RGB images and infrared images, so as to obtain RGB features and infrared features .
[0048] In a specific implementation, when multi-source target detection is required, obtain the paired RGB images and infrared images that need to be subjected to multi-source target detection, and use a two-stream backbone network to downsample and perform multi-scale feature extraction on the input paired RGB images and infrared images, so as to obtain RGB features and infrared features , , corresponding to downsampling resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 respectively.
[0049] In an example, the two-stream backbone network, the first light-adaptive mixture of experts dynamic feature decomposition and combination module, the second light-adaptive mixture of experts dynamic feature decomposition and combination module, the spatial pyramid pooling module, the RGB neck network, the infrared neck network, the RGB detection network, and the infrared detection network together constitute a multi-source target detection model. The multi-source target detection task is actually implemented by the multi-source target detection model. The structure of the multi-source target detection model can be as follows Figure 2 shown.
[0050] In an example, as shown Figure 2 below, the two-stream backbone network consists of an RGB feature extraction branch and an infrared feature extraction branch. The network structures of the RGB feature extraction branch and the infrared feature extraction branch are basically the same.
[0051] The RGB feature extraction branch consists of five sequentially connected visible light convolution modules. The input of the first visible light convolution module is the RGB image, and the input of the subsequent visible light convolution module is the output of the previous visible light convolution module. All five visible light convolution modules are used to downsample and extract features from their own inputs, and RGB features with resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the RGB image are obtained in sequence. The five RGB features are sequentially denoted as , , , and .
[0052] The infrared feature extraction branch consists of five sequentially connected thermal infrared convolution modules. The input of the first thermal infrared convolution module is the infrared image, and the input of the next thermal infrared convolution module is the output of the previous thermal infrared convolution module. The five thermal infrared convolution modules are used to downsample and extract features of their own inputs, and obtain infrared features with resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the infrared image, respectively. The five infrared features are recorded in the order of the thermal infrared convolution modules as , , , and .
[0053] It is important to note that , , , It will not be input into the subsequent network of the multi-source target detection model and will not participate in the calculation of the subsequent network.
[0054] In one example, the network structure of the visible light convolution module and the thermal infrared convolution module is the same, both of which are composed of a basic convolution unit and a cross-stage local convolution unit in series. The basic convolution unit is composed of a 3×3 convolution, a batch normalization layer, and a SiLU activation function layer in series. The cross-stage local convolution unit first uses a basic convolution unit to increase the dimension of its own input features, and then splits them into two feature maps according to the channel dimension. and , For residual connection, The features are extracted through multiple basic convolutional units in turn, and finally the two parts of the feature map are channel-joined and output with dimensionality reduction using basic convolutional units.
[0055] Step 102: and Input to the first illumination adaptive hybrid expert dynamic feature decomposition combination module, and obtain the reorganized features through illumination perception and multi-expert decision making and ,Will and The input is sent to the second illumination adaptive hybrid expert dynamic feature decomposition combination module, and the reorganized features are obtained through illumination perception and multi-expert decision making. and ,right and Perform channel splicing to obtain the spliced features ,Will Input to the spatial pyramid pooling module to obtain fusion features .
[0056] In a specific implementation, , , , , , will participate in the calculations of the subsequent network, and are input into the first light-adaptive hybrid expert dynamic feature decomposition and combination module, and through light perception and multi-expert decision-making, the recombined features and are obtained. Similarly, and are input into the second light-adaptive hybrid expert dynamic feature decomposition and combination module, and through light perception and multi-expert decision-making, the recombined features and are obtained. and will perform channel concatenation to obtain the concatenated feature , is input into the spatial pyramid pooling module to obtain the fused feature .
[0057] In an example, the network structures of the first light-adaptive hybrid expert dynamic feature decomposition and combination module and the second light-adaptive hybrid expert dynamic feature decomposition and combination module are the same, and both are composed of a light perception network and three sparse gating hybrid expert models. The three sparse gating hybrid expert models are the daytime sparse gating hybrid expert model , the nighttime sparse gating hybrid expert model and the dark night sparse gating hybrid expert model . The specific structure of the light-adaptive hybrid expert dynamic feature decomposition and combination module can be as shown in Figure 3 .
[0058] After the light perception network of the first light-adaptive hybrid expert dynamic feature decomposition and combination module receives , based on it predicts the image light value , and at the same time obtains the maximum light value index . Based on it determines the current belonging light scene, and based on the current belonging light scene, it activates the corresponding sparse gating hybrid expert model; among them, the current belonging light scene is daytime, nighttime or dark night.
[0059] The activated sparse gating hybrid expert model of the first light-adaptive hybrid expert dynamic feature decomposition and combination module will and perform channel concatenation to obtain the concatenated feature , then use the noise Top-K sparse gating network to calculate the gating weights, activate the corresponding Top-K expert networks for comprehensive decision-making, and finally weight the gating weights and the outputs of the expert networks participating in the calculation to obtain the recombined features and .
[0060] After the light perception network of the second light-adaptive hybrid expert dynamic feature decomposition and combination module receives , based on predict the image light value , and at the same time obtain the maximum light value index , based on judge the current light scene, and start the corresponding sparse gating hybrid expert model based on the current light scene; among them, the current light scene is day, night or dark night.
[0061] The started sparse gating hybrid expert model of the second light-adaptive hybrid expert dynamic feature decomposition and combination module will and perform channel splicing to obtain the spliced features , then use the noise Top-K sparse gating network to calculate the gating weights, activate the corresponding Top-K expert networks for comprehensive decision-making, and finally weight the gating weights and the outputs of the expert networks participating in the calculation to obtain the recombined features and .
[0062] In an example, the structure of the light perception network is as Figure 4 shown. The light perception network consists of a first convolutional module, a max pooling module, a second convolutional module, an adaptive average pooling module, a first linear module, a second linear module, a linear layer, and a ReLU activation function layer. The network structures of the first convolutional module and the second convolutional module are the same, and both consist of a 1×1 convolutional layer with a stride of 1, a batch normalization layer, and a ReLU activation function layer. The network structures of the first linear module and the second linear module are the same, and both consist of a linear layer, a ReLU activation function layer, and Droupout.
[0063] The input of the light perception network sequentially passes through the first convolutional module, the max pooling module, and the second convolutional module for channel dimensionality reduction and spatial downsampling, and then is adaptively average pooled by the adaptive average pooling module, and the size becomes . Then, the first linear module, the second linear module, and the linear layer are used for light prediction, and finally the predicted image light value with a length of 3 and the maximum light value index are output by the ReLU activation function layer. Used to calculate the illumination loss together with the illumination ground truth during the training phase , supervise the illumination perception network to perform effective illumination prediction on the input features, which is then used to determine the current illumination scene. Among them, , , , respectively represent height, width, and number of channels.
[0064] In one example, the working principle of the sparse gated mixture-of-experts model can be expressed by the formula as:
[0065] ;
[0066] ;
[0067] ;
[0068] ;
[0069] Among them, represents the channel concatenation operation, represents the th expert network, represents the gating weight corresponding to the th expert network.
[0070] Taking the input sample as an example, the calculation process of the noise Top-K sparse gating network used in this embodiment is shown in the following formula:
[0071] ;
[0072] ;
[0073] ;
[0074] First, use the trainable weight to calculate an initial weight for the input sample . Then, use the trainable weight to introduce noise, and use the Softplus activation function and the StandardNormal standard normal distribution to generate tunable Gaussian noise, the purpose of which is to achieve load balancing. After adding, the expert weight with noise is finally obtained. After that, use the method to retain the top values and set the remaining values to (The corresponding gating value after being processed by the Softmax function is 0). Finally, the Softmax function is used to generate sparse gating weights.
[0075] In one example, the specific structure of the sparse gating mixture-of-experts model can be as Figure 5 shown. In the sparse gating mixture-of-experts model, each expert network adopts an independent feature decomposition and combination module, which consists of a decomposition unit, four feature extractors with the same network structure, and a multi-source feature recombination unit. The four feature extractors with the same network structure are a visible light basic feature extractor, a visible light detail feature extractor, a thermal infrared basic feature extractor, and a thermal infrared detail feature extractor respectively.
[0076] The input to the expert network is first decomposed by the decomposition unit into and . The visible light basic feature extractor and the visible light detail feature extractor respectively perform feature extraction on to obtain an RGB basic feature and an RGB detail feature . The thermal infrared basic feature extractor and the thermal infrared detail feature extractor respectively perform feature extraction on to obtain an infrared basic feature and an infrared detail feature . Among them, .
[0077] The multi-source feature recombination unit performs multi-source feature recombination based on , and to obtain , and at the same time performs multi-source feature recombination based on , and to obtain .
[0078] and The calculation processes of are expressed by formulas as:
[0079] ;
[0080] .
[0081] In one example, the specific structure of the feature extractor is as Figure 6As shown, the feature extractor is specifically composed of a first 1×1 convolution unit, a channel splitting unit, a first 3×3 convolution unit, a second 3×3 convolution unit, a channel concatenation unit, and a second 1×1 convolution unit. The network structures of the first 1×1 convolution unit and the second 1×1 convolution unit are the same, and both are composed of a 1×1 convolution layer with a stride of 1, a batch normalization layer, and a SiLU activation function layer. The network structures of the first 3×3 convolution unit and the second 3×3 convolution unit are the same, and both are composed of a 3×3 convolution layer with a stride of 1, a batch normalization layer, and a SiLU activation function layer.
[0082] Let the input feature of the feature extractor be , After being processed by the first 1×1 convolution unit, it is split by the channel splitting unit along the channel dimension into and , After being processed by the first 3×3 convolution unit and the second 3×3 convolution unit, it is then concatenated by the channel concatenation unit with and along the channel dimension to obtain the concatenated feature . Finally, after being processed by the second 1×1 convolution unit to achieve dimensionality reduction, the output feature is obtained.
[0083] Step 103: Input , and into the RGB neck network for multi-path feature aggregation to generate the RGB aggregation feature , and . Input , and into the infrared neck network for multi-path feature aggregation to generate the infrared aggregation feature , and .
[0084] In a specific implementation, after the work of the first light-adaptive mixture-of-experts dynamic feature decomposition and combination module, the second light-adaptive mixture-of-experts dynamic feature decomposition and combination module, and the spatial pyramid pooling module is completed, , and will be input into the RGB neck network for multi-path feature aggregation, and the RGB neck network will generate the RGB aggregation feature , and , , and It will be input into the infrared neck network for multi-path feature aggregation, and the infrared neck network will generate infrared aggregation features , and .
[0085] In one example, the network structures of the RGB neck network and the infrared neck network are as Figure 2 shown, both are top-down and bottom-up multi-path aggregation networks (Path Aggregation Network, PAN).
[0086] In the RGB neck network, the top-down path is composed of a general PAN module, a first visible light PAN module, and a second visible light PAN module connected in sequence. The bottom-up path is composed of a third visible light PAN module and a fourth visible light PAN module connected in sequence. The output of the general PAN module is also used as the input of the fourth visible light PAN module, and the output of the first visible light PAN module is also used as the input of the third visible light PAN module, , and are respectively input into the general PAN module, the first visible light PAN module, and the second visible light PAN module. The second visible light PAN module, the third visible light PAN module, and the fourth visible light PAN module respectively output three different scales of RGB aggregation features 、 and .
[0087] In the infrared neck network, the top-down path is composed of a general PAN module, a first thermal infrared PAN module, and a second thermal infrared PAN module connected in sequence. The bottom-up path is composed of a third thermal infrared PAN module and a fourth thermal infrared PAN module connected in sequence. The output of the general PAN module is also used as the input of the fourth thermal infrared PAN module, and the output of the first thermal infrared PAN module is also used as the input of the third thermal infrared PAN module, , and are respectively input into the general PAN module, the first thermal infrared PAN module, and the second thermal infrared PAN module. The second thermal infrared PAN module, the third thermal infrared PAN module, and the fourth thermal infrared PAN module respectively output three different scales of infrared aggregation features 、 and .
[0088] Step 104, take , and It is input into the RGB detection network for detection. After using non-maximum suppression processing, the RGB detection result is obtained. 、 and are input into the infrared detection network for detection. After using non-maximum suppression processing, the infrared detection result is obtained.
[0089] In a specific implementation, after the RGB neck network and the infrared neck network are calculated, 、 and will be input into the RGB detection network for detection, and at the same time, non-maximum suppression processing is used to obtain the RGB detection result. 、 and will be input into the infrared detection network for detection, and after using non-maximum suppression processing at the same time, the infrared detection result can be obtained.
[0090] In an example, the overall loss function used when training the multi-source target detection model can be expressed by the formula:
[0091] ;
[0092] Among them, represents the overall loss function used when training the multi-source target detection model, represents the classification loss, represents the th category, represents the regression loss, represents the distribution focus loss, represents the multi-source feature decomposition loss, which is used to increase the correlation between basic features and reduce the correlation between detailed features. represents the light perception loss, which is calculated based on and obtained. represents the importance loss, which is used to measure the dispersion degree of the batch weights of each expert network to ensure that each expert network is equally important. represents the load balancing loss, which is used to avoid the imbalance of the number of samples assigned to each expert network. 、 、 、 are respectively the loss weights of the classification loss, the loss weights of the regression loss, the loss weights of the distribution focus loss, and the loss weights of the multi-source feature decomposition loss. is the binary cross loss, is the CIoU loss, is the cross-entropy loss.
[0093] In one example, , , , are generally set to 0.5, 7.5, 1.5, 20.
[0094] In one example, the multi-source feature decomposition loss is calculated as follows:
[0095] ;
[0096] where is the Pearson correlation coefficient operator, , used to ensure that the calculation result is positive.
[0097] In one example, the purpose of setting the importance loss and the load balancing loss is to avoid the phenomenon of "self-reinforcement" in the training process of the gating network. "Self-reinforcement" means that large weights are always assigned to a fixed number of expert networks, while another part of the expert networks are always in an inactive state. This phenomenon will lead to the loss of dynamics of the gating network.
[0098] The importance mentioned in this embodiment refers to the batch sum of the gating values of each expert network for the input sample , and the calculation formula is: .
[0099] The importance loss calculates the coefficient of variation (CV) of the importance set to measure the dispersion of the batch weights of each expert, aiming to ensure that all expert networks are equally important and avoid the problem of unbalanced weights of expert networks. The calculation formula is expressed as: .
[0100] To avoid the problem of unbalanced sample numbers assigned to each expert network, the smoothing estimator is used to smooth the number of samples assigned to each expert network, and the load balancing loss is calculated based on this. The calculation formula is expressed as: .
[0101] A multi-source object detection method based on light perception and a mixture of experts model proposed in this embodiment uses a two-stream backbone network, a first light-adaptive mixture of experts dynamic feature decomposition and combination module, a second light-adaptive mixture of experts dynamic feature decomposition and combination module, a spatial pyramid pooling module, an RGB neck network, an infrared neck network, an RGB detection network, and an infrared detection network to work together to achieve multi-source object detection. The light-adaptive mixture of experts dynamic feature decomposition and combination module integrates light perception technology and configures a mixture of experts model. Through the light perception technology, the current belonging light scene can be accurately predicted, and then the mixture of experts model corresponding to the current belonging light scene is used for comprehensive decision-making, effectively improving the accuracy and scene robustness of RGB-infrared multi-source object detection. The method proposed in this embodiment is particularly suitable for RGB-infrared multi-source object detection tasks from the perspective of drones with rapid changes in aspects such as light, morphology, shooting angle, and background, and can well adapt to the differences brought about by external environmental changes, thus providing a scientific and reliable basic support for drone tasks.
[0102] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step, or some steps can be split into multiple steps. As long as the same logical relationship is included, it is within the protection scope of this application. Minor modifications added to the algorithm or process or insignificant designs introduced, but without changing the core design of its algorithm and process, are within the protection scope of this application.
[0103] In one embodiment, the multi-source object detection model is obtained by transforming the lightweight single-source one-stage YOLOv8 network, which will be hereinafter referred to as IMoE-DMDet. To verify the effectiveness of the method proposed in this application, we iteratively trained and tested IMoE-DMDet using the publicly available RGB-infrared multi-source object detection dataset DroneVehicle from the perspective of drones and compared it with other mainstream methods. Among them, the DroneVehicle training set contains 17,990 pairs of RGB-infrared images, and the test set contains 8,980 pairs of RGB-infrared images, including 5 types of aerial detection targets such as cars, vans, buses, trucks, and box trucks.
[0104] The algorithm evaluation indicators involve multiple aspects such as accuracy, speed, and number of parameters. We use the mean average precision (mAP@0.5 and mAP@0.5:0.95) to evaluate the algorithm accuracy, use frames per second (FPS) to evaluate the algorithm running time, and use Parms and GFLOPs to evaluate the model parameters and complexity. The relevant experimental results are as Figure 7As shown, in IMoE-DMDet, the configuration is E4K2, that is, there are 4 expert networks in the light-adaptive mixture-of-experts dynamic feature decomposition combination module, and the TOP-2 networks are selected to participate in the decision-making.
[0105] Compared with the mainstream method I 2 MDet, IMoE-DMDet improves by 3.73% and 13.16% respectively in terms of the mAP@0.5 and mAP@0.5:0.95 accuracy metrics. The number of model parameters is less than 1 / 6 of that of I 2 MDet. On a single NVIDIA GeForce RTX3090 graphics card, the detection speed reaches an astonishing 25.63 frames per second, basically meeting the requirements of the RGB infrared multi-source real-time target detection task from the perspective of drones.
[0106] Another embodiment of the present application proposes an electronic device, the structure of which is as Figure 8 shown, including: at least one processor 201; and a memory 202 communicatively connected to the at least one processor 201; wherein, the memory 202 stores instructions executable by the at least one processor 201, and the instructions are executed by the at least one processor 201 to enable the at least one processor 201 to execute a multi-source target detection method based on light perception and mixture-of-experts model as described in the above method embodiment.
[0107] Among them, the memory and the processor are connected in a bus manner. The bus can include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and the memory together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be further described in this application. The bus interface provides an interface between the bus and the transceiver. The transceiver can be one element or multiple elements, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor is transmitted over the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor.
[0108] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory can be used to store the data used by the processor when performing operations.
[0109] Another embodiment of the present application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement a multi-source target detection method based on light perception and a mixture of experts model as described in the foregoing method embodiments.
[0110] That is, those skilled in the art can understand that all or part of the steps in the foregoing method embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium, including several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of a multi-source target detection method based on light perception and a mixture of experts model as described in the method embodiments of the present application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0111] Those of ordinary skill in the art can understand that the foregoing embodiments are specific embodiments for implementing the present application, and in actual applications, various changes can be made to them in form and details without departing from the spirit and scope of the present application.
Claims
1. A multi-source object detection method based on light perception and a mixture of experts model, characterized in that Comprising: The input paired RGB images and infrared images are downsampled and multi-scale feature extraction is performed using a two-stream backbone network to obtain RGB features at different scales and infrared features ; among them, , corresponding to downsampling resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 respectively; Input and into the first light-adaptive hybrid expert dynamic feature decomposition and combination module. Through light perception and multi-expert decision-making, the recombined features and are obtained. Input and into the second light-adaptive hybrid expert dynamic feature decomposition and combination module. Through light perception and multi-expert decision-making, the recombined features and are obtained. Concatenate and along the channels to obtain the concatenated feature Input into the spatial pyramid pooling module to obtain the fused feature ; Input , and into the RGB neck network for multi-path feature aggregation to generate RGB aggregation features , and . Input , and into the infrared neck network for multi-path feature aggregation to generate infrared aggregation features , and ; Input , and into the RGB detection network for detection. After performing non-maximum suppression processing, the RGB detection result is obtained. Input , and into the infrared detection network for detection. After performing non-maximum suppression processing, the infrared detection result is obtained; The network structures of the first light-adaptive hybrid expert dynamic feature decomposition and combination module and the second light-adaptive hybrid expert dynamic feature decomposition and combination module are the same, both consisting of a light perception network and three sparse gating hybrid expert models. The three sparse gating hybrid expert models are the daytime sparse gating hybrid expert model , the nighttime sparse gating hybrid expert model and the dark-night sparse gating hybrid expert model ; After the light perception network of the first light adaptive hybrid expert dynamic feature decomposition combination module receives , based on predicts the image light value , and at the same time obtains the maximum light value index , based on determines the currently belonging light scene, and starts the corresponding sparse gated mixture-of-experts model based on the currently belonging light scene; wherein, the currently belonging light scene is day, night or dark night; The sparse gated mixture-of-experts model launched by the first light-adaptive hybrid expert dynamic feature decomposition and combination module will and perform channel concatenation to obtain the concatenated feature , then use the noise Top-K sparse gating network to calculate the gating weights and activate the corresponding Top-K expert networks for comprehensive decision-making. Finally, the gating weights and the outputs of the expert networks participating in the calculation are weighted to obtain the recombined features and ; After the light perception network of the second light adaptive hybrid expert dynamic feature decomposition combination module receives , based on , it predicts the image light value , and at the same time obtains the maximum light value index . Based on , it determines the current light scene it belongs to, and starts the corresponding sparse gated mixture of experts model based on the current light scene it belongs to; among them, the current light scene it belongs to is daytime, night or dark night. The sparse gating mixture-of-experts model launched by the second light-adaptive hybrid expert dynamic feature decomposition and combination module will and perform channel concatenation to obtain the concatenated features . Subsequently, the noise Top-K sparse gating network is used to calculate the gating weights, and the corresponding Top-K expert networks are activated for comprehensive decision-making. Finally, the gating weights and the outputs of the expert networks participating in the calculation are weighted to obtain the reorganized features and .
2. The multi-source target detection method based on light perception and mixture of experts model according to claim 1, wherein The dual-stream backbone network consists of an RGB feature extraction branch and an infrared feature extraction branch; The RGB feature extraction branch consists of five sequentially connected visible light convolution modules. The input of the first visible light convolution module is the RGB image, and the input of the subsequent visible light convolution module is the output of the previous visible light convolution module. All five visible light convolution modules are used to downsample and extract features from their own inputs, successively obtaining RGB features with resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the RGB image. The five RGB features are sequentially denoted as , , , and ; The infrared feature extraction branch consists of five sequentially connected thermal infrared convolution modules. The input of the first thermal infrared convolution module is the infrared image, and the input of the next thermal infrared convolution module is the output of the previous thermal infrared convolution module. The five thermal infrared convolution modules are used to downsample and extract features of their own inputs, and obtain infrared features with resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the infrared image, respectively. The five infrared features are recorded in the order of the thermal infrared convolution modules as , , , and .
3. The multi-source target detection method based on light perception and mixture of experts model according to claim 1, characterized in that, The light perception network consists of a first convolutional module, a max pooling module, a second convolutional module, an adaptive average pooling module, a first linear module, a second linear module, a linear layer, and a ReLU activation function layer. The network structures of the first convolutional module and the second convolutional module are the same, and both consist of a 1×1 convolutional layer with a stride of 1, a batch normalization layer, and a ReLU activation function layer. The network structures of the first linear module and the second linear module are the same, and both consist of a linear layer, a ReLU activation function layer, and Droupout; Input of the light perception network It sequentially passes through the first convolutional module, the max pooling module, and the second convolutional module for channel dimensionality reduction and spatial downsampling. Subsequently, it undergoes adaptive average pooling by the adaptive average pooling module, and then light prediction is performed by the first linear module, the second linear module, and the linear layer. Finally, the predicted image light value is output by the ReLU activation function layer and the maximum light value index ; among which, .
4. A multi-source target detection method based on light perception and a mixture of experts model according to claim 3, characterized in that Each expert network in the sparse gating mixture-of-experts model adopts an independent feature decomposition and combination module. The feature decomposition and combination module consists of a decomposition unit, four feature extractors with the same network structure, and a multi-source feature recombination unit. The four feature extractors are a visible light basic feature extractor, a visible light detail feature extractor, a thermal infrared basic feature extractor, and a thermal infrared detail feature extractor; Input to the expert network First, it is decomposed by the decomposition unit into and , and the visible light basic feature extractor and the visible light detailed feature extractor respectively extract features from to obtain the RGB basic feature and the RGB detailed feature . The thermal infrared basic feature extractor and the thermal infrared detailed feature extractor respectively extract features from to obtain the infrared basic feature and the infrared detailed feature ; among which, ; The multi-source feature recombination unit performs multi-source feature recombination based on , and to obtain . At the same time, based on , and perform multi-source feature recombination to obtain .
5. The multi-source target detection method based on light perception and hybrid expert model according to claim 4, characterized in that The feature extractor consists of a first 1×1 convolutional unit, a channel splitting unit, a first 3×3 convolutional unit, a second 3×3 convolutional unit, a channel concatenation unit, and a second 1×1 convolutional unit. The network structures of the first 1×1 convolutional unit and the second 1×1 convolutional unit are the same, and both consist of a 1×1 convolutional layer with a stride of 1, a batch normalization layer, and a SiLU activation function layer. The network structures of the first 3×3 convolutional unit and the second 3×3 convolutional unit are the same, and both consist of a 3×3 convolutional layer with a stride of 1, a batch normalization layer, and a SiLU activation function layer; Let the input feature of the feature extractor be , After being processed by the first 1×1 convolutional unit, it is split into and along the channel dimension by the channel splitting unit. After being processed by the first 3×3 convolutional unit and the second 3×3 convolutional unit, it is concatenated with and along the channel dimension by the channel concatenation unit to obtain the concatenated feature . After being processed by the second 1×1 convolutional unit, dimensionality reduction is achieved to obtain the output feature .
6. The multi-source target detection method based on light perception and a mixture of experts model according to claim 5, characterized in that, The network structures of the RGB neck network and the infrared neck network are both top-down and bottom-up multi-path aggregation networks; In the RGB neck network, the top-down path is composed of a general PAN module, a first visible light PAN module, and a second visible light PAN module connected in sequence. The bottom-up path is composed of a third visible light PAN module and a fourth visible light PAN module connected in sequence. The output of the general PAN module is also used as the input of the fourth visible light PAN module, and the output of the first visible light PAN module is also used as the input of the third visible light PAN module. , and are respectively input into the general PAN module, the first visible light PAN module, and the second visible light PAN module. The second visible light PAN module, the third visible light PAN module, and the fourth visible light PAN module respectively output three different scales of RGB aggregation features 、 and ; In the infrared neck network, the top-down path is composed of sequentially connected ordinary PAN modules, the first thermal infrared PAN module, and the second thermal infrared PAN module. The bottom-up path is composed of sequentially connected third thermal infrared PAN module and fourth thermal infrared PAN module. The output of the ordinary PAN module also serves as the input of the fourth thermal infrared PAN module, and the output of the first thermal infrared PAN module also serves as the input of the third thermal infrared PAN module. , and are respectively input into the ordinary PAN module, the first thermal infrared PAN module, and the second thermal infrared PAN module. The second thermal infrared PAN module, the third thermal infrared PAN module, and the fourth thermal infrared PAN module respectively output three different scales of infrared aggregation features 、 and .
7. A multi-source target detection method based on light perception and a mixture of experts model according to claim 6, characterized in that, The dual-stream backbone network, the first light-adaptive mixture-of-experts dynamic feature decomposition and combination module, the second light-adaptive mixture-of-experts dynamic feature decomposition and combination module, the spatial pyramid pooling module, the RGB neck network, the infrared neck network, the RGB detection network, and the infrared detection network together constitute a multi-source object detection model. The overall loss function used when training the multi-source object detection model is expressed by the formula: ; Among them, represents the overall loss function used when training the multi-source object detection model, represents the classification loss, represents the th category, represents the regression loss, represents the distribution focusing loss, represents the multi-source feature decomposition loss, which is used to increase the correlation between basic features and reduce the correlation between detailed features, represents the illumination perception loss, which is calculated based on and obtained. represents the importance loss, which is used to measure the dispersion degree of the batch weights of each expert network to ensure that each expert network is equally important, represents the load balancing loss, which is used to avoid the imbalance of the number of samples assigned to each expert network, , , , are respectively the loss weights of the classification loss, the regression loss, the distribution focusing loss, and the multi-source feature decomposition loss, is the binary cross entropy loss, is the CIoU loss, is the cross entropy loss.
8. An electronic device, characterized in that, Comprising: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor. The instructions are executed by the at least one processor so that the at least one processor can execute a multi-source object detection method according to any one of claims 1 to 7.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it can implement a multi-source object detection method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Infrared visible light fusion recognition system and method based on modal difference feature guidance
CN114898189A
RGB-infrared multi-source image target detection method based on dynamic network feature fusion and YOLOv5
CN116311363A