Multi-source target detection method based on illumination perception and hybrid expert model
By introducing light perception and hybrid expert models into the multi-source object detection method, the problem of insufficient accuracy in the RGB infrared multi-source object detection of drone viewing angle is solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510542977.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The existing multi-source object detection methods have insufficient accuracy in complex application scenarios, especially in the RGB infrared multi-source object detection task of drone viewing angle. Rapid changes in lighting, morphology, shooting perspective and background lead to insufficient detection accuracy.
The multi-source object detection method based on light perception and hybrid expert model is adopted to determine the current lighting scene through light perception technology, and comprehensive decision-making is made using the corresponding hybrid expert model to improve detection accuracy and scene robustness.
It effectively improves the accuracy and robustness of RGB infrared multi-source object detection, and is especially suitable for drone perspective tasks with rapidly changing lighting, morphology, shooting perspective and background.
Smart Images

Figure CN120088465A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the technical fields of artificial intelligence and deep learning, and particularly to a multi-source object detection method based on light perception and a mixture of experts model. Background Art
[0002] Object detection is a core task in computer vision technology, the purpose of which is to identify and locate objects of interest in an image or video frame. Object detection methods based on deep learning have currently been widely applied in fields such as security monitoring, intelligent transportation, intelligent retail, and industrial manufacturing. RGB images (visible light images) can provide rich texture information and detail information under good lighting conditions, but their imaging effects are poor at night or in bad weather. In order to improve the accuracy and robustness of object detection, research teams at home and abroad have focused their research on RGB-infrared multi-source object detection technology. RGB-infrared multi-source object detection technology makes full use of the complementary characteristics of two different modalities of images and can adapt to different environmental conditions and object characteristics.
[0003] In order to achieve complementary fusion of multi-source features, a large number of attention mechanism modules have been designed in related research to improve object detection performance. Zhang L et al. proposed the CIAN network (Cross-modality interactive attention network), which fuses RGB features and infrared features by using a cross-modal interactive fusion module and further enhances the network's ability to extract and transmit context information by using a context enhancement module. Zhang H et al. designed an intra-modal attention module and an inter-modal attention module to enhance single-modal features and fuse complementary multi-modal features respectively. Yuan et al. proposed a network structure called Calibration and Complementary Transformer based on the powerful correlation modeling ability of Transformer, and calibrated modal features by calculating the cross-attention relationship between the RGB modality and the infrared modality.
[0004] However, most of the multi-source object detection methods proposed currently adopt a static network structure, which has the problem of insufficient accuracy in complex application scenarios. For example, in the RGB-infrared multi-source object detection task from the perspective of an unmanned aerial vehicle (UAV), the objects captured by the UAV change rapidly in terms of lighting, shape, shooting angle, and background. The feature performances of RGB samples and infrared samples are different at different times and in different scenarios. A single and fixed feature fusion strategy and a static network structure cannot adapt to the differences brought about by external environmental changes, which results in seriously insufficient detection accuracy for RGB-infrared multi-source object detection. Summary of the Invention
[0005] In view of this, an embodiment of the present application proposes a multi-source object detection method based on light perception and a mixture of experts model. By using light perception technology to determine the current light scene, and then using the corresponding mixture of experts model for comprehensive decision-making, the accuracy and scene robustness of RGB-infrared multi-source object detection are effectively improved.
[0006] To achieve the above object, an embodiment of the present application proposes a multi-source object detection method based on light perception and a mixture of experts model, and the method includes the following steps: Use a two-stream backbone network to perform downsampling and multi-scale feature extraction on the input paired RGB images and infrared images, so as to obtain RGB features of different scales and infrared features ; where , corresponding to downsampling resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 respectively; Input and into the first light-adaptive mixture of experts dynamic feature decomposition and combination module, and through light perception and multi-expert decision-making, obtain the recombined features and , input and into the second light-adaptive mixture of experts dynamic feature decomposition and combination module, and through light perception and multi-expert decision-making, obtain the recombined features and , perform channel splicing on and to obtain the spliced feature , input into the spatial pyramid pooling module to obtain the fused feature ; Input , and into the RGB neck network for multi-path feature aggregation to generate the RGB aggregated feature , and , input , and into the infrared neck network for multi-path feature aggregation to generate the infrared aggregated feature , and ; Input , and into the RGB detection network for detection, and after using non-maximum suppression processing, obtain the RGB detection result, and input , and are input into an infrared detection network for detection. After non-maximum suppression processing, infrared detection results are obtained.
[0007] To achieve the above object, an embodiment of the present application further provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, instructions executable by the at least one processor are stored in the memory, and the instructions are executed by the at least one processor so that the at least one processor can execute a multi-source target detection method based on light perception and a mixture of experts model as described above.
[0008] To achieve the above object, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which can implement a multi-source target detection method based on light perception and a mixture of experts model as described above when executed by a processor.
[0009] The multi-source target detection method based on light perception and a mixture of experts model proposed in the embodiment of the present application uses a two-stream backbone network, a first light-adaptive mixture of experts dynamic feature decomposition and combination module, a second light-adaptive mixture of experts dynamic feature decomposition and combination module, a spatial pyramid pooling module, an RGB neck network, an infrared neck network, an RGB detection network, and an infrared detection network to work together to achieve multi-source target detection. The light-adaptive mixture of experts dynamic feature decomposition and combination module integrates light perception technology and configures a mixture of experts model. Through the light perception technology, the currently belonging light scene can be accurately predicted, and then the mixture of experts model corresponding to the currently belonging light scene is used for comprehensive decision-making, effectively improving the accuracy and scene robustness of RGB infrared multi-source target detection. The method proposed in the present application is particularly applicable to the RGB infrared multi-source target detection task from the perspective of an unmanned aerial vehicle with rapid changes in aspects such as light, morphology, shooting angle, and background, and can well adapt to the differences brought by external environmental changes, thereby providing a reliable basic support for unmanned aerial vehicle tasks.
[0010] Optionally, the two-stream backbone network is composed of an RGB feature extraction branch and an infrared feature extraction branch; The RGB feature extraction branch consists of five sequentially connected visible light convolution modules. The input of the first visible light convolution module is the RGB image, and the input of the subsequent visible light convolution module is the output of the previous visible light convolution module. All five visible light convolution modules are used to downsample and extract features from their own inputs, successively obtaining RGB features with resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the RGB image. The five RGB features are sequentially denoted as , , , and ; The infrared feature extraction branch consists of five sequentially connected thermal infrared convolution modules. The input of the first thermal infrared convolution module is the infrared image, and the input of the subsequent thermal infrared convolution module is the output of the previous thermal infrared convolution module. All five thermal infrared convolution modules are used to downsample and extract features from their own inputs, successively obtaining infrared features with resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the infrared image. The five infrared features are sequentially denoted as , , , and .
[0011] Optionally, the first light-adaptive mixture-of-experts dynamic feature decomposition and combination module and the second light-adaptive mixture-of-experts dynamic feature decomposition and combination module have the same network structure, both consisting of a light perception network and three sparse gating mixture-of-experts models, namely the daytime sparse gating mixture-of-experts model , the nighttime sparse gating mixture-of-experts model and the dark-night sparse gating mixture-of-experts model ; After receiving , the light perception network of the first light-adaptive mixture-of-experts dynamic feature decomposition and combination module predicts the image light value based on and simultaneously obtains the maximum light value index . It determines the current belonging light scene based on and activates the corresponding sparse gating mixture-of-experts model based on the current belonging light scene; The activated sparse gating mixture-of-experts model of the first light-adaptive mixture-of-experts dynamic feature decomposition and combination module concatenates and along the channel dimension to obtain the concatenated feature, then use the noise Top-K sparse gating network to calculate the gating weights, activate the corresponding Top-K expert networks for comprehensive decision-making, and finally weight the gating weights and the outputs of the expert networks participating in the calculation to obtain the recombined features and ; After the light perception network of the second light-adaptive hybrid expert dynamic feature decomposition and combination module receives , based on predict the image light value , and at the same time obtain the maximum light value index , based on judge the current light scene, and start the corresponding sparse gating hybrid expert model based on the current light scene; among them, the current light scene is day, night or dark night; The started sparse gating hybrid expert model of the second light-adaptive hybrid expert dynamic feature decomposition and combination module will and perform channel splicing to obtain the spliced features , then use the noise Top-K sparse gating network to calculate the gating weights, activate the corresponding Top-K expert networks for comprehensive decision-making, and finally weight the gating weights and the outputs of the expert networks participating in the calculation to obtain the recombined features and .
[0012] Optionally, the light perception network consists of a first convolutional module, a max pooling module, a second convolutional module, an adaptive average pooling module, a first linear module, a second linear module, a linear layer, and a ReLU activation function layer. The network structures of the first convolutional module and the second convolutional module are the same, both consisting of a 1×1 convolutional layer with a stride of 1, a batch normalization layer, and a ReLU activation function layer. The network structures of the first linear module and the second linear module are the same, both consisting of a linear layer, a ReLU activation function layer, and Droupout; The input of the light perception network successively passes through the first convolutional module, the max pooling module, and the second convolutional module for channel dimensionality reduction and spatial downsampling, then is adaptively averaged by the adaptive average pooling module, and then the light is predicted by the first linear module, the second linear module, and the linear layer. Finally, the predicted image light value and the maximum light value index are output by the ReLU activation function layer; among them, .
[0013] Optionally, each expert network in the sparse gating mixture-of-experts model adopts an independent feature decomposition and combination module, which consists of a decomposition unit, four feature extractors with the same network structure, and a multi-source feature recombination unit. The four feature extractors are a visible light basic feature extractor, a visible light detail feature extractor, a thermal infrared basic feature extractor, and a thermal infrared detail feature extractor; Input to the expert network is first decomposed by the decomposition unit into and . The visible light basic feature extractor and the visible light detail feature extractor respectively perform feature extraction on to obtain the RGB basic feature and the RGB detail feature . The thermal infrared basic feature extractor and the thermal infrared detail feature extractor respectively perform feature extraction on to obtain the infrared basic feature and the infrared detail feature ; where ; The multi-source feature recombination unit is based on , and to perform multi-source feature recombination to obtain , and at the same time is based on , and to perform multi-source feature recombination to obtain .
[0014] Optionally, the feature extractor consists of a first 1×1 convolution unit, a channel splitting unit, a first 3×3 convolution unit, a second 3×3 convolution unit, a channel concatenation unit, and a second 1×1 convolution unit. The first 1×1 convolution unit and the second 1×1 convolution unit have the same network structure, both consisting of a 1×1 convolution layer with a stride of 1, a batch normalization layer, and a SiLU activation function layer. The first 3×3 convolution unit and the second 3×3 convolution unit have the same network structure, both consisting of a 3×3 convolution layer with a stride of 1, a batch normalization layer, and a SiLU activation function layer; Let the input feature of the feature extractor be , After being processed by the first 1×1 convolution unit, it is split by the channel splitting unit along the channel dimension into and , After being processed by the first 3×3 convolution unit and the second 3×3 convolution unit, it is concatenated by the channel concatenation unit with and along the channel dimension to obtain the concatenated feature , After being processed by the second 1×1 convolution unit, dimensionality reduction is achieved to obtain the output features .
[0015] Optionally, the network structures of the RGB neck network and the infrared neck network are both top-down and bottom-up multi-path aggregation networks; In the RGB neck network, the top-down path is composed of a common PAN module, a first visible light PAN module, and a second visible light PAN module connected in sequence. The bottom-up path is composed of a third visible light PAN module and a fourth visible light PAN module connected in sequence. The output of the common PAN module is also used as the input of the fourth visible light PAN module, and the output of the first visible light PAN module is also used as the input of the third visible light PAN module, , and are respectively input into the common PAN module, the first visible light PAN module, and the second visible light PAN module. The second visible light PAN module, the third visible light PAN module, and the fourth visible light PAN module respectively output three different scales of RGB aggregation features 、 and ; In the infrared neck network, the top-down path is composed of a common PAN module, a first thermal infrared PAN module, and a second thermal infrared PAN module connected in sequence. The bottom-up path is composed of a third thermal infrared PAN module and a fourth thermal infrared PAN module connected in sequence. The output of the common PAN module is also used as the input of the fourth thermal infrared PAN module, and the output of the first thermal infrared PAN module is also used as the input of the third thermal infrared PAN module, , and are respectively input into the common PAN module, the first thermal infrared PAN module, and the second thermal infrared PAN module. The second thermal infrared PAN module, the third thermal infrared PAN module, and the fourth thermal infrared PAN module respectively output three different scales of infrared aggregation features 、 and .
[0016] Optionally, the multi-source object detection model is composed of a two-stream backbone network, a first light-adaptive hybrid expert dynamic feature decomposition and combination module, a second light-adaptive hybrid expert dynamic feature decomposition and combination module, a spatial pyramid pooling module, an RGB neck network, an infrared neck network, an RGB detection network, and an infrared detection network. The overall loss function used when training the multi-source object detection model is expressed by the formula: ; Among them, represents the overall loss function used when training the multi-source object detection model, represents the classification loss, represents the th category, represents the regression loss, represents the distribution focus loss, represents the multi-source feature decomposition loss, which is used to increase the correlation between basic features and reduce the correlation between detailed features, represents the light perception loss, which is calculated based on and obtained. represents the importance loss, which is used to measure the dispersion degree of the batch weights of each expert network to ensure that each expert network is equally important, represents the load balancing loss, which is used to avoid the imbalance of the number of samples assigned to each expert network, , , , are respectively the loss weights of the classification loss, the regression loss, the distribution focus loss, and the multi-source feature decomposition loss, is the binary cross loss, is the CIoU loss, is the cross-entropy loss. Description of the Drawings
[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the related art, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or the related art. The following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. The drawings described here are only used to explain the present application and are not used to limit the present application.
[0018] Figure 1 is provided in an embodiment of the present application, which is a flowchart of a multi-source object detection method based on light perception and a hybrid expert model; Figure 2 is provided in an embodiment of the present application, which is a schematic structural diagram of a multi-source object detection model; Figure 3 is provided in an embodiment of the present application, which is a schematic structural diagram of a light-adaptive hybrid expert dynamic feature decomposition combination module; Figure 4 is provided in an embodiment of the present application, which is a schematic structural diagram of a light perception network; Figure 5It is a schematic structural diagram of a sparse gating mixture-of-experts model provided in an embodiment of the present application; Figure 6 It is a schematic structural diagram of a feature extractor provided in an embodiment of the present application; Figure 7 It is a schematic diagram of the experimental results of a comparative experiment conducted on the Drone Vehicle dataset provided in an embodiment of the present application; Figure 8 It is a schematic structural diagram of an electronic device provided in another embodiment of the present application. Detailed implementation manners
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below with reference to the accompanying drawings. Those of ordinary skill in the art can understand that in the embodiments of the present application, many technical details are proposed to help readers better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions required to be protected by the present application can still be implemented. The division of the following embodiments is only for the convenience of description and should not constitute any limitation to the specific implementation manners of the present application. The embodiments can be combined with each other and cross-referenced on the premise of not conflicting with each other.
[0020] To solve the problem that the existing multi-source object detection methods cannot adapt to the differences brought by changes in the external environment, an embodiment of the present application proposes a multi-source object detection method based on light perception and a mixture-of-experts model, which is applied to an electronic device. The electronic device can be a terminal or a server. In this embodiment and the following embodiments, the electronic device is described by taking the server as an example. The implementation details of a multi-source object detection method based on light perception and a mixture-of-experts model proposed in this embodiment are specifically described below. The following content is only implementation details provided for convenience of understanding and is not necessary for implementing this solution.
[0021] The specific process of a multi-source object detection method based on light perception and a mixture-of-experts model proposed in this embodiment can be as Figure 1 shown and includes: Step 101: Use a two-stream backbone network to perform downsampling and multi-scale feature extraction on the input paired RGB images and infrared images, so as to obtain RGB features and infrared features .
[0022] In a specific implementation, when multi-source object detection is required, pairs of RGB images and infrared images for multi-source object detection are obtained, and a two-stream backbone network is used to downsample and extract multi-scale features from the input pairs of RGB images and infrared images, thereby obtaining RGB features of different scales. and infrared features , , corresponding to downsampling resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 respectively.
[0023] In one example, the two-stream backbone network, the first light-adaptive mixture-of-experts dynamic feature decomposition combination module, the second light-adaptive mixture-of-experts dynamic feature decomposition combination module, the spatial pyramid pooling module, the RGB neck network, the infrared neck network, the RGB detection network, and the infrared detection network together constitute a multi-source object detection model. The multi-source object detection task is actually implemented by the multi-source object detection model, and the structure of the multi-source object detection model can be as Figure 2 shown.
[0024] In one example, as Figure 2 shown, the two-stream backbone network consists of an RGB feature extraction branch and an infrared feature extraction branch. The network structures of the RGB feature extraction branch and the infrared feature extraction branch are basically the same.
[0025] The RGB feature extraction branch consists of five sequentially connected visible-light convolution modules. The input of the first visible-light convolution module is the RGB image, and the input of the subsequent visible-light convolution module is the output of the previous visible-light convolution module. The five visible-light convolution modules are all used to downsample and extract features from their own inputs, successively obtaining RGB features with resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the RGB image. The five RGB features are sequentially denoted as , , , and .
[0026] The infrared feature extraction branch consists of five sequentially connected thermal-infrared convolution modules. The input of the first thermal-infrared convolution module is the infrared image, and the input of the subsequent thermal-infrared convolution module is the output of the previous thermal-infrared convolution module. The five thermal-infrared convolution modules are all used to downsample and extract features from their own inputs, successively obtaining infrared features with resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the infrared image. The five infrared features are sequentially denoted as , , , and 。
[0027] It should be noted that 、 、 、 will not be input into the subsequent network of the multi-source target detection model and will not participate in the calculation of the subsequent network.
[0028] In one example, the network structures of the visible light convolution module and the thermal infrared convolution module are the same, and both are composed of a basic convolution unit and a cross-stage local convolution unit connected in series. The basic convolution unit is composed of a 3×3 convolution, a batch normalization layer, and a SiLU activation function layer connected in series. The cross-stage local convolution unit first uses a basic convolution unit to perform dimensionality increase on its own input features, and then splits them into two feature maps according to the channel dimension and , for residual connection, and then feature extraction is performed through multiple basic convolution units in sequence. Finally, the two parts of the feature maps will be concatenated along the channel dimension, and a basic convolution unit will be used for dimensionality reduction output.
[0029] Step 102: Input and into the first light-adaptive hybrid expert dynamic feature decomposition and combination module. Through light perception and multi-expert decision-making, the recombined features and are obtained. Input and into the second light-adaptive hybrid expert dynamic feature decomposition and combination module. Through light perception and multi-expert decision-making, the recombined features and are obtained. Concatenate and along the channel dimension to obtain the concatenated feature . Input into the spatial pyramid pooling module to obtain the fused feature .
[0030] In specific implementation, 、 、 、 、 、 will participate in the calculation of the subsequent network. and are input into the first light-adaptive hybrid expert dynamic feature decomposition and combination module. Through light perception and multi-expert decision-making, the recombined features and are obtained. Similarly, and It is input into the second light-adaptive hybrid expert dynamic feature decomposition and combination module, and through light perception and multi-expert decision-making, the recombined features are obtained. and , and will be concatenated in channels to obtain the concatenated features . It is input into the spatial pyramid pooling module to obtain the fused features .
[0031] In one example, the network structures of the first light-adaptive hybrid expert dynamic feature decomposition and combination module and the second light-adaptive hybrid expert dynamic feature decomposition and combination module are the same, and both are composed of a light perception network and three sparse gated hybrid expert models. The three sparse gated hybrid expert models are the daytime sparse gated hybrid expert model , the night sparse gated hybrid expert model and the dark night sparse gated hybrid expert model . The specific structure of the light-adaptive hybrid expert dynamic feature decomposition and combination module can be as shown in Figure 3 as follows.
[0032] After the light perception network of the first light-adaptive hybrid expert dynamic feature decomposition and combination module receives , based on it predicts the image light value , and at the same time obtains the maximum light value index . Based on it determines the current light scene to which it belongs, and starts the corresponding sparse gated hybrid expert model based on the current light scene to which it belongs; among them, the current light scene to which it belongs is daytime, night or dark night.
[0033] The started sparse gated hybrid expert model of the first light-adaptive hybrid expert dynamic feature decomposition and combination module will and be concatenated in channels to obtain the concatenated features , and then use the noise Top-K sparse gating network to calculate the gating weights and activate the corresponding Top-K expert networks for comprehensive decision-making. Finally, the gating weights and the outputs of the expert networks participating in the calculation are weighted to obtain the recombined features and .
[0034] After the light perception network of the second light-adaptive hybrid expert dynamic feature decomposition and combination module receives , based on it predicts the image light value , and at the same time obtains the maximum light value index , based on judge the current illumination scene to which it belongs, and start the corresponding sparse gating mixture-of-experts model based on the current illumination scene to which it belongs; wherein, the current illumination scene to which it belongs is daytime, night or dark night.
[0035] The started sparse gating mixture-of-experts model of the second illumination adaptive mixture-of-experts dynamic feature decomposition and combination module will and perform channel splicing to obtain the spliced features , then use the noise Top-K sparse gating network to calculate the gating weights, activate the corresponding Top-K expert networks for comprehensive decision-making, and finally weight the gating weights and the outputs of the participating expert networks to obtain the recombined features and .
[0036] In one example, the structure of the illumination perception network is as Figure 4 shown. The illumination perception network consists of a first convolution module, a max pooling module, a second convolution module, an adaptive average pooling module, a first linear module, a second linear module, a linear layer, and a ReLU activation function layer. The network structures of the first convolution module and the second convolution module are the same, and both consist of a 1×1 convolution layer with a stride of 1, a batch normalization layer, and a ReLU activation function layer. The network structures of the first linear module and the second linear module are the same, and both consist of a linear layer, a ReLU activation function layer, and Droupout.
[0037] The input of the illumination perception network successively undergoes channel dimensionality reduction and spatial downsampling through the first convolution module, the max pooling module, and the second convolution module, and then undergoes adaptive average pooling by the adaptive average pooling module, and the size becomes , and then the first linear module, the second linear module, and the linear layer perform illumination prediction, and finally the ReLU activation function layer outputs the predicted image illumination value with a length of 3 and the maximum illumination value index . is used to calculate the illumination loss together with the illumination ground truth during the training phase , supervise the illumination perception network to perform effective illumination prediction on the input features, is used to judge the current illumination scene to which it belongs. Among them, , , , respectively represent height, width, and number of channels.
[0038] In one example, the working principle of the sparse gating mixture-of-experts model can be expressed by the formula: ; ; ; ; Among them, represents the channel splicing operation, represents the th expert network, represents the gating weight corresponding to the th expert network.
[0039] Taking the input sample as an example, the calculation process of the noise Top-K sparse gating network used in this embodiment is shown in the following formula: ; ; ; First, use the trainable weight to calculate an initial weight for the input sample . Then, use the trainable weight to introduce noise, and use the Softplus activation function and the StandardNormal standard normal distribution to generate tunable Gaussian noise, the purpose of which is to achieve load balancing. After adding, finally obtain the expert weight with noise. After that, use the method to retain the first values and set the remaining values to (the corresponding gating value is 0 after being processed by the Softmax function). Finally, use the Softmax function to generate the sparse gating weight.
[0040] In an example, the specific structure of the sparse gating mixture-of-experts model can be as Figure 5 shown. Each expert network in the sparse gating mixture-of-experts model adopts an independent feature decomposition and combination module, and the feature decomposition and combination module consists of a decomposition unit, four feature extractors with the same network structure, and a multi-source feature recombination unit. The four feature extractors with the same network structure are respectively a visible light basic feature extractor, a visible light detail feature extractor, a thermal infrared basic feature extractor, and a thermal infrared detail feature extractor.
[0041] The input to the expert network is first decomposed by the decomposition unit into and . The visible light basic feature extractor and the visible light detail feature extractor respectively perform feature extraction on to obtain the RGB basic feature and RGB detail features , the thermal infrared basic feature extractor and the thermal infrared detail feature extractor respectively extract features from to obtain the infrared basic features and infrared detail features . Among them, .
[0042] The multi-source feature recombination unit is based on , and to perform multi-source feature recombination to obtain . At the same time, based on , and to perform multi-source feature recombination to obtain .
[0043] and The calculation process is expressed by the formula as: ; .
[0044] In an example, the specific structure of the feature extractor is as shown in Figure 6 . The feature extractor is specifically composed of a first 1×1 convolutional unit, a channel splitting unit, a first 3×3 convolutional unit, a second 3×3 convolutional unit, a channel concatenation unit, and a second 1×1 convolutional unit. The network structures of the first 1×1 convolutional unit and the second 1×1 convolutional unit are the same, both consisting of a 1×1 convolutional layer with a stride of 1, a batch normalization layer, and a SiLU activation function layer. The network structures of the first 3×3 convolutional unit and the second 3×3 convolutional unit are the same, both consisting of a 3×3 convolutional layer with a stride of 1, a batch normalization layer, and a SiLU activation function layer.
[0045] Let the input feature of the feature extractor be , After being processed by the first 1×1 convolutional unit, it is split by the channel splitting unit along the channel dimension into and , After being processed by the first 3×3 convolutional unit and the second 3×3 convolutional unit, it is then concatenated with and along the channel dimension by the channel concatenation unit to obtain the concatenated feature . Finally After being processed by the second 1×1 convolutional unit, dimensionality reduction is achieved to obtain the output feature .
[0046] Step 103, take , and are input into the RGB neck network for multi-path feature aggregation to generate RGB aggregated features , and , and , and are input into the infrared neck network for multi-path feature aggregation to generate infrared aggregated features , and .
[0047] In a specific implementation, after the first light-adaptive mixture-of-experts dynamic feature decomposition and combination module, the second light-adaptive mixture-of-experts dynamic feature decomposition and combination module, and the spatial pyramid pooling module are completed, , and will be input into the RGB neck network for multi-path feature aggregation, and the RGB neck network generates RGB aggregated features , and , , and will be input into the infrared neck network for multi-path feature aggregation, and the infrared neck network generates infrared aggregated features , and .
[0048] In one example, the network structures of the RGB neck network and the infrared neck network are as Figure 2 shown, both are top-down and bottom-up multi-path aggregation networks (Path Aggregation Network, PAN).
[0049] In the RGB neck network, the top-down path is composed of a general PAN module, a first visible light PAN module, and a second visible light PAN module connected in sequence. The bottom-up path is composed of a third visible light PAN module and a fourth visible light PAN module connected in sequence. The output of the general PAN module also serves as the input of the fourth visible light PAN module, and the output of the first visible light PAN module also serves as the input of the third visible light PAN module. , and are respectively input into the general PAN module, the first visible light PAN module, and the second visible light PAN module. The second visible light PAN module, the third visible light PAN module, and the fourth visible light PAN module respectively output three different scales of RGB aggregated features 、 and 。
[0050] In the infrared neck network, the top-down path is composed of sequentially connected ordinary PAN modules, the first thermal infrared PAN module, and the second thermal infrared PAN module. The bottom-up path is composed of sequentially connected third thermal infrared PAN module and fourth thermal infrared PAN module. The output of the ordinary PAN module is also used as the input of the fourth thermal infrared PAN module, and the output of the first thermal infrared PAN module is also used as the input of the third thermal infrared PAN module. 、 and are respectively input into the ordinary PAN module, the first thermal infrared PAN module, and the second thermal infrared PAN module. The second thermal infrared PAN module, the third thermal infrared PAN module, and the fourth thermal infrared PAN module respectively output three different scales of infrared aggregation features 、 and 。
[0051] Step 104, input 、 and into the RGB detection network for detection. After using non-maximum suppression processing, the RGB detection result is obtained. Input 、 and into the infrared detection network for detection. After using non-maximum suppression processing, the infrared detection result is obtained.
[0052] In a specific implementation, after the RGB neck network and the infrared neck network are calculated, 、 and will be input into the RGB detection network for detection, and at the same time, non-maximum suppression processing is used to obtain the RGB detection result. 、 and will be input into the infrared detection network for detection, and after using non-maximum suppression processing, the infrared detection result can be obtained.
[0053] In an example, the overall loss function used when training the multi-source target detection model can be expressed by the formula: ; where, represents the overall loss function used when training the multi-source target detection model, represents the classification loss, represents the category represents the regression loss represents the distribution focusing loss represents the multi-source feature decomposition loss, which is used to increase the correlation between basic features and reduce the correlation between detailed features represents the illumination perception loss, based on calculated represents the importance loss, which is used to measure the dispersion of the batch weights of each expert network to ensure that each expert network is equally important represents the load balancing loss, which is used to avoid the imbalance of the number of samples assigned to each expert network and and and are respectively the loss weights of the classification loss, the regression loss, the distribution focusing loss, and the multi-source feature decomposition loss is the binary cross loss is the CIoU loss is the cross-entropy loss
[0054] In one example and and and are generally set to 0.5, 7.5, 1.5, and 20
[0055] In one example, the multi-source feature decomposition loss has the following calculation formula ; where is the Pearson correlation coefficient operator , which is used to ensure that the calculation result is positive
[0056] In one example, the purpose of setting the importance loss and the load balancing loss is to avoid the phenomenon of "self-reinforcement" in the training process of the gating network. "Self-reinforcement" means that larger weights are always assigned to a fixed few expert networks, while another part of the expert networks are always in an inactive state. This phenomenon will cause the gating network to lose its dynamics
[0057] The importance mentioned in this embodiment refers to the batch sum of the gating values of each expert network for the input sample , and the calculation formula is .
[0058] The importance loss Calculate the coefficient of variation (CV) of the importance set to measure the dispersion of the batch weights of each expert, aiming to ensure that all expert networks are equally important and avoid the problem of unbalanced weights of expert networks. The calculation formula is expressed as: .
[0059] To avoid the problem of unbalanced sample numbers assigned to each expert network, use the smoothing estimator to smoothly estimate the number of samples assigned to each expert network and calculate the load balancing loss based on this . The calculation formula is expressed as: .
[0060] A multi-source target detection method based on light perception and a mixture of experts model proposed in this embodiment uses a two-stream backbone network, a first light-adaptive mixture of experts dynamic feature decomposition and combination module, a second light-adaptive mixture of experts dynamic feature decomposition and combination module, a spatial pyramid pooling module, an RGB neck network, an infrared neck network, an RGB detection network, and an infrared detection network to work together to achieve multi-source target detection. The light-adaptive mixture of experts dynamic feature decomposition and combination module integrates light perception technology and configures a mixture of experts model. Through the light perception technology, the currently belonging light scene can be accurately predicted, and then the mixture of experts model corresponding to the currently belonging light scene is used for comprehensive decision-making, effectively improving the accuracy and scene robustness of RGB-infrared multi-source target detection. The method proposed in this embodiment is particularly suitable for RGB-infrared multi-source target detection tasks from the perspective of drones with rapid changes in aspects such as light, morphology, shooting angle, and background, and can well adapt to the differences brought by external environmental changes, thus providing a scientific and reliable basic support for drone tasks.
[0061] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step, or some steps can be split into multiple steps. As long as the same logical relationship is included, it is within the protection scope of this application. Minor modifications added to the algorithm or process or minor designs introduced, but without changing the core design of the algorithm and process, are within the protection scope of this application.
[0062] In one embodiment, the multi-source target detection model is obtained by modifying the lightweight single-source one-stage YOLOv8 network, which will be hereinafter referred to as IMoE-DMDet. To verify the effectiveness of the method proposed in this application, we iteratively trained and tested IMoE-DMDet using the publicly available RGB-infrared multi-source target detection dataset DroneVehicle from the perspective of drones, and compared it with other mainstream methods. Among them, the DroneVehicle training set contains 17,990 pairs of RGB-infrared images, and the test set contains 8,980 pairs of RGB-infrared images, including 5 types of aerial detection targets: cars, vans, buses, trucks, and box trucks.
[0063] The algorithm evaluation metrics involve multiple aspects such as accuracy, speed, and the number of parameters. We use the mean average precision (mAP@0.5 and mAP@0.5:0.95) to evaluate the algorithm accuracy, use frames per second (FPS) to evaluate the algorithm running time, and use Parms and GFLOPs to evaluate the number of model parameters and complexity. The relevant experimental results are as Figure 7 shown, where IMoE-DMDet uses the configuration E4K2, that is, there are 4 expert networks in the light-adaptive hybrid expert dynamic feature decomposition combination module, and the TOP-2 networks are selected to participate in the decision-making.
[0064] Compared with the mainstream method I 2 MDet, IMoE-DMDet has improved by 3.73% and 13.16% respectively in terms of the accuracy metrics of mAP@0.5 and mAP@0.5:0.95. The number of model parameters is less than I 2 MDet's 1 / 6. On a single NVIDIA GeForce RTX3090 graphics card, the detection speed reaches an amazing 25.63 frames per second, basically meeting the requirements of the RGB-infrared multi-source real-time target detection task from the perspective of drones.
[0065] Another embodiment of this application proposes an electronic device, the structure of which is as Figure 8 shown, including: at least one processor 201; and a memory 202 communicatively connected to the at least one processor 201; wherein, the memory 202 stores instructions executable by the at least one processor 201, and the instructions are executed by the at least one processor 201 to enable the at least one processor 201 to execute a multi-source target detection method based on light perception and a hybrid expert model as described in the above method embodiment.
[0066] Among them, the memory and the processor are connected in a bus manner. The bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and the memory together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be further described in this application. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single component or multiple components, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor is transmitted over the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor.
[0067] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory can be used to store the data used by the processor when executing operations.
[0068] Another embodiment of the present application proposes a computer-readable storage medium storing a computer program, which when executed by a processor, can implement a multi-source target detection method based on light perception and a mixture of experts model as described in the above method embodiment.
[0069] That is, those skilled in the art can understand that all or part of the steps in the above method embodiment can be completed by a program instructing relevant hardware. The program is stored in a storage medium, including several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of a multi-source target detection method based on light perception and a mixture of experts model as described in the method embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0070] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application, and in actual applications, various changes can be made in form and details without departing from the spirit and scope of the present application.
Claims
1. A multi-source target detection method based on illumination perception and hybrid expert model, characterized in that: include: A two-stream backbone network is used to downsample and extract multi-scale features from the input paired RGB images and infrared images to obtain RGB features of different scales. and infrared characteristics ;in, , corresponding to downsampling resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32, respectively; Will and Input to the first illumination adaptive hybrid expert dynamic feature decomposition combination module, and obtain the reorganized features through illumination perception and multi-expert decision making and ,Will and The input is sent to the second illumination adaptive hybrid expert dynamic feature decomposition combination module, and the reorganized features are obtained through illumination perception and multi-expert decision making. and ,right and Perform channel splicing to obtain the spliced features ,Will Input to the spatial pyramid pooling module to obtain fusion features ; Will , and Input to the RGB neck network for multi-path feature aggregation to generate RGB aggregate features , and ,Will , and Input to the infrared neck network for multi-path feature aggregation to generate infrared aggregate features , and ; Will , and Input to the RGB detection network for detection, and use non-maximum value suppression to obtain the RGB detection result. , and The input is sent to the infrared detection network for detection, and the infrared detection result is obtained after the non-maximum value is used for suppression processing.
2. According to claim 1, a multi-source target detection method based on illumination perception and hybrid expert model is characterized in that: The two-stream backbone network consists of an RGB feature extraction branch and an infrared feature extraction branch; The RGB feature extraction branch consists of five sequentially connected visible light convolution modules. The input of the first visible light convolution module is the RGB image, and the input of the next visible light convolution module is the output of the previous visible light convolution module. The five visible light convolution modules are used to downsample and extract features of their own inputs, and obtain RGB features with resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the RGB image, respectively. The five RGB features are recorded in the order of the visible light convolution modules as , , , and ; The infrared feature extraction branch consists of five sequentially connected thermal infrared convolution modules. The input of the first thermal infrared convolution module is the infrared image, and the input of the next thermal infrared convolution module is the output of the previous thermal infrared convolution module. The five thermal infrared convolution modules are used to downsample and extract features of their own inputs, and obtain infrared features with resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the infrared image, respectively. The five infrared features are recorded in the order of the thermal infrared convolution modules as , , , and .
3. The multi-source target detection method based on illumination perception and hybrid expert model according to claim 2 is characterized in that: The network structures of the first illumination adaptive hybrid expert dynamic feature decomposition combination module and the second illumination adaptive hybrid expert dynamic feature decomposition combination module are the same, both of which consist of an illumination perception network and three sparse gated hybrid expert models. The three sparse gated hybrid expert models are daytime sparse gated hybrid expert models, , Night Sparse Gated Hybrid Expert Model and Dark Night Sparse Gated Hybrid Expert Model ; The illumination perception network of the first illumination adaptive hybrid expert dynamic feature decomposition combination module receives Afterwards, based on Predict image illumination values , and get the maximum illumination value index ,based on Determine the current lighting scene, and start the corresponding sparse gated hybrid expert model based on the current lighting scene; wherein the current lighting scene is daytime, nighttime or dark night; The sparse gated hybrid expert model activated by the first illumination adaptive hybrid expert dynamic feature decomposition combination module will and Perform channel splicing to obtain the spliced features , then use the noise Top-K sparse gating network to calculate the gating weight, and activate the corresponding Top-K expert networks for comprehensive decision-making. Finally, the gating weight and the output of the expert network involved in the calculation are weighted to obtain the reorganized feature and ; The illumination perception network of the second illumination adaptive hybrid expert dynamic feature decomposition combination module receives Afterwards, based on Predict image illumination values , and get the maximum illumination value index ,based on Determine the current lighting scene, and start the corresponding sparse gated hybrid expert model based on the current lighting scene; wherein the current lighting scene is daytime, nighttime or dark night; The sparse gated hybrid expert model activated by the second illumination adaptive hybrid expert dynamic feature decomposition combination module will and Perform channel splicing to obtain the spliced features , then use the noise Top-K sparse gating network to calculate the gating weight, and activate the corresponding Top-K expert networks for comprehensive decision-making. Finally, the gating weight and the output of the expert network involved in the calculation are weighted to obtain the reorganized feature and .
4. The multi-source target detection method based on illumination perception and hybrid expert model according to claim 3 is characterized in that: The light perception network consists of a first convolution module, a maximum pooling module, a second convolution module, an adaptive average pooling module, a first linear module, a second linear module, a linear layer, and a ReLU activation function layer. The first convolution module and the second convolution module have the same network structure, both consisting of a 1×1 convolution layer with a step size of 1, a batch normalization layer, and a ReLU activation function layer. The first linear module and the second linear module have the same network structure, both consisting of a linear layer, a ReLU activation function layer, and Dropout. Input to the Light-Aware Network The first convolution module, the maximum pooling module, and the second convolution module are used to perform channel dimension reduction and spatial downsampling, followed by the adaptive average pooling module for adaptive average pooling, and then the first linear module, the second linear module, and the linear layer are used for illumination prediction. Finally, the ReLU activation function layer outputs the predicted image illumination value. and the maximum light value index ;in, .
5. The multi-source target detection method based on illumination perception and hybrid expert model according to claim 4, characterized in that: Each expert network in the sparse gated hybrid expert model adopts an independent feature decomposition and combination module, which consists of a decomposition unit, four feature extractors with the same network structure, and a multi-source feature recombination unit. The four feature extractors are respectively a visible light basic feature extractor, a visible light detail feature extractor, a thermal infrared basic feature extractor, and a thermal infrared detail feature extractor. Input to the expert network First, the decomposition unit is decomposed into and , the visible light basic feature extractor and the visible light detail feature extractor are respectively Perform feature extraction to obtain RGB basic features and RGB detail features , the thermal infrared basic feature extractor and the thermal infrared detail feature extractor are respectively Perform feature extraction to obtain infrared basic features and infrared detail features ;in, ; Multi-source feature recombination unit based on , and Perform multi-source feature recombination to obtain , and based on , and Perform multi-source feature recombination to obtain .
6. The multi-source target detection method based on illumination perception and hybrid expert model according to claim 5, characterized in that: The feature extractor consists of the first 1×1 convolution unit, the channel splitting unit, the first 3×3 convolution unit, the second 3×3 convolution unit, the channel splicing unit and the second 1×1 convolution unit. The network structure of the first 1×1 convolution unit and the second 1×1 convolution unit is the same, both consisting of a 1×1 convolution layer with a stride of 1, a batch normalization layer and a SiLU activation function layer. The network structure of the first 3×3 convolution unit and the second 3×3 convolution unit is the same, both consisting of a 3×3 convolution layer with a stride of 1, a batch normalization layer and a SiLU activation function layer. Assume that the input features of the feature extractor are , After being processed by the first 1×1 convolution unit, the channel splitting unit splits it into and , After being processed by the first 3×3 convolution unit and the second 3×3 convolution unit, the channel concatenation unit and and Splice along the channel dimension to obtain the spliced features , After the second 1×1 convolution unit is processed, the dimension is reduced and the output feature is obtained. .
7. The multi-source target detection method based on illumination perception and hybrid expert model according to claim 6, characterized in that: The network structures of RGB neck network and infrared neck network are both top-down and bottom-up multi-path aggregation networks; In the RGB neck network, the top-down path consists of a common PAN module, a first visible light PAN module, and a second visible light PAN module connected in sequence, and the bottom-up path consists of a third visible light PAN module and a fourth visible light PAN module connected in sequence. The output of the common PAN module also serves as the input of the fourth visible light PAN module, and the output of the first visible light PAN module also serves as the input of the third visible light PAN module. , and The RGB features are respectively input into the common PAN module, the first visible light PAN module and the second visible light PAN module. The second visible light PAN module, the third visible light PAN module and the fourth visible light PAN module output RGB aggregation features of three different scales. 、 and ; In the infrared neck network, the top-down path consists of a common PAN module, a first thermal infrared PAN module, and a second thermal infrared PAN module connected in sequence, and the bottom-up path consists of a third thermal infrared PAN module and a fourth thermal infrared PAN module connected in sequence. The output of the common PAN module also serves as the input of the fourth thermal infrared PAN module, and the output of the first thermal infrared PAN module also serves as the input of the third thermal infrared PAN module. , and The infrared data are respectively input into the common PAN module, the first thermal infrared PAN module and the second thermal infrared PAN module. The second thermal infrared PAN module, the third thermal infrared PAN module and the fourth thermal infrared PAN module respectively output infrared aggregation features of three different scales. 、 and .
8. The multi-source target detection method based on illumination perception and hybrid expert model according to claim 7, characterized in that: The dual-stream backbone network, the first illumination adaptive hybrid expert dynamic feature decomposition combination module, the second illumination adaptive hybrid expert dynamic feature decomposition combination module, the spatial pyramid pooling module, the RGB neck network, the infrared neck network, the RGB detection network and the infrared detection network together constitute the multi-source target detection model. The overall loss function used in training the multi-source target detection model is expressed by the formula: ; in, represents the overall loss function used when training the multi-source target detection model, represents the classification loss, Indicates categories, represents the regression loss, represents the distribution focusing loss, Represents the multi-source feature decomposition loss, which is used to increase the correlation between basic features and reduce the correlation between detail features. Represents the light perception loss, based on Calculated, Represents the importance loss, which is used to measure the discreteness of the batch weights of each expert network to ensure that each expert network is equally important. Represents the load balancing loss, which is used to avoid the imbalance in the number of samples allocated to each expert network. , , , They are the loss weight of classification loss, the loss weight of regression loss, the loss weight of distribution focus loss, and the loss weight of multi-source feature decomposition loss. is the binary cross loss, is the CIoU loss, is the cross entropy loss.
9. An electronic device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a multi-source target detection method based on illumination perception and a hybrid expert model as described in any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it is possible to implement a multi-source target detection method based on illumination perception and a hybrid expert model as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Infrared visible light fusion recognition system and method based on modal difference feature guidance
CN114898189A
RGB-infrared multi-source image target detection method based on dynamic network feature fusion and YOLOv5
CN116311363A
Salient target detection method based on multiband visual image perception and fusion
CN117132759A
Unmanned aerial vehicle multi-modal remote sensing image target detection method and device based on hybrid Mamb-CNN network
CN119540786A
Unmanned aerial vehicle online target counting method based on multi-modal dynamic neural network
CN120032280A