Lightweight target detection method and system based on multi-modal image

By adopting a multi-weight adjustment module in object detection, combined with weight-guided feature optimization, local lighting perception and parameterless channel attention mechanism, the problems of information loss in single-modal object detection under low-light conditions and insufficient complementarity utilization of early fusion methods and information redundancy are solved, and the accuracy and stability of object detection are significantly improved.

CN120107565AActive Publication Date: 2025-06-06BERTE DIGITAL INTELLIGENCE (HEBEI) TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510574739.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-06-06
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

The existing single-modal object detection is lost in low-light conditions, making it difficult to accurately identify the target, and early fusion methods cannot fully explore the complementarity between modes, resulting in key features loss and information redundancy.

Method used

A lightweight object detection method based on multimodal images is adopted, through weight-guided feature optimization, local light perception mechanism and parameterless channel attention mechanism, a multi-weight adjustment module is formed to effectively fuse the features of infrared and visible images, dynamically adjust the fusion weight, reduce redundant information, and enhance detection accuracy.

Benefits of technology

It significantly improves the accuracy and stability of the object detection task, overcomes the limitations of a single mode in complex environments, enhances the model's deep learning ability of modal relationships, and improves the performance of complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107565A_ABST
    Figure CN120107565A_ABST
Patent Text Reader

Abstract

The invention provides a lightweight target detection method and system based on a multi-modal image, and relates to the field of computer vision, and the method comprises the steps: respectively carrying out the convolution, batch normalization and ReLU activation operation of an infrared image and a visible light image to obtain a feature map, then generating a feature weight mask through the convolution operation and a Sigmoid function, and carrying out the recognition of the feature weight mask; performing pixel-by-pixel weighting on the feature map based on the feature weight mask to generate an optimized feature map; extracting illumination perception information through a local illumination perception module so as to perform weighted calculation on the optimized feature map to obtain a weighted feature map, and splicing the visible light weighted feature map and the infrared weighted feature map in the channel dimension to generate a spliced feature map; and performing subsequent feature extraction operation through the residual connection and the parameter-free channel attention module, and inputting the feature extraction operation to the target detection network for target detection so as to output a detection result. According to the method, target detection with higher precision can be realized on the premise of not introducing larger parameter quantity and calculation quantity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a lightweight target detection method and system based on multimodal images. Background Art

[0002] Object detection technology plays an important role in the field of computer vision and is widely used in many fields such as autonomous driving, video surveillance, and intelligent security. In recent years, with the development of deep learning technology, object detection algorithms based on convolutional neural networks (CNNs) have made significant progress. However, these methods still face challenges in some complex scenarios, especially in low-light environments, bad weather, and object detection under occlusion, where their performance usually drops significantly. Traditional single-modal object detection, such as detection based on visible light images, is heavily dependent on the illumination conditions of the image, so the detection accuracy is often not ideal at night or in low-light environments. In order to make up for the shortcomings of single-modal detection, multimodal fusion detection has gradually become an effective means. By combining data sources of different modalities (such as visible light and infrared images), the detectability of the target can be significantly improved, especially under complex lighting conditions. This multimodal information fusion method can rely on infrared thermal images to supplement information when visible light information is insufficient, thereby enhancing the robustness and accuracy of detection. At present, multimodal object detection can be divided into three categories according to the fusion strategy: early fusion, mid-term fusion, and late fusion. Among them, early fusion is the most intuitive fusion method. The infrared image and the visible light image are spliced ​​to generate a four-channel image, which is then input into the conventional target detection architecture. Compared with the other two methods, the early fusion model is simpler, easier to implement and integrate. In the process of realizing the present invention, the applicant found that the existing single-modal target detection usually faces the problem of information loss under low-light conditions, making it difficult to accurately identify the target, and thus significantly reducing the performance. In addition, although fusion target detection can make up for the shortcomings of single modality to a certain extent, the existing early fusion methods often face the problem of being unable to fully explore the complementarity between modalities, resulting in the loss of key features, limiting the model's deep learning of modal relationships, and thus affecting the performance in complex scenes. Secondly, the problem of information redundancy. After the repeated features of different modalities are fused, redundant information will be generated, which will increase the computational burden and interfere with the model's extraction of effective information. In addition, when there are differences in illumination in the image, the model cannot accurately capture the image detail information, thereby affecting the detection accuracy. Therefore, how to improve the detection accuracy of target detection tasks has become a technical problem that needs to be solved. Summary of the invention

[0003] The present invention aims to solve at least one of the technical problems existing in the prior art or related technology, and discloses a lightweight target detection method and system based on multimodal images, which can effectively integrate the features of infrared images and visible light images to achieve higher-precision target detection.

[0004] The first aspect of the present invention discloses a lightweight target detection method based on multimodal images, comprising: receiving an image to be detected: receiving a paired infrared image and a visible light image; weight-guided feature optimization: performing convolution, batch normalization and ReLU activation operations on the infrared image and the visible light image respectively to generate an enhanced infrared feature map and an enhanced visible light feature map; using a convolution operation to reduce the number of channels of the enhanced infrared feature map to 1, and limiting the output value range to between 0 and 1 through a Sigmoid function to generate an infrared feature weight mask; using a convolution operation to reduce the number of channels of the enhanced visible light feature map to 1, and limiting the output value range to between 0 and 1 through a Sigmoid function to generate a visible light feature weight mask; weighting the enhanced infrared feature map pixel by pixel based on the infrared feature weight mask to generate an optimized The infrared feature map after the enhancement is obtained; the enhanced visible light feature map is weighted pixel by pixel based on the visible light feature weight mask to generate an optimized visible light feature map; local illumination perception mechanism: the enhanced infrared feature map and the enhanced visible light feature map are input into the local illumination perception module to extract illumination perception information; the weights of the optimized infrared feature map and the optimized visible light feature map are adjusted according to the illumination perception information to generate a visible light weighted feature map and an infrared weighted feature map; the visible light weighted feature map and the infrared weighted feature map are spliced ​​in the channel dimension to generate a spliced ​​feature map; parameter-free channel attention mechanism: the spliced ​​feature map is subjected to subsequent feature extraction operations through residual connection and parameter-free channel attention module to obtain a multimodal fusion feature map; target detection: the multimodal fusion feature map is input into the target detection network for target detection to output the detection result.

[0005] According to the lightweight target detection method based on multimodal images disclosed in the present invention, preferably, the target detection network is a YOLO network.

[0006] According to the lightweight target detection method based on multimodal images disclosed in the present invention, preferably, the YOLO network specifically includes: a backbone network, a neck and a detection head; the multimodal fusion feature map is further subjected to feature extraction in the backbone network part, and features of multiple resolutions are aggregated in the neck of the target detection network to enhance detection accuracy, and finally the final target bounding box and category prediction are generated by the detection head.

[0007] According to the lightweight target detection method based on multimodal images disclosed in the present invention, preferably, the calculation process of the local illumination perception module specifically includes: the input features are subjected to convolution calculation, batch normalization and ReLU activation operations through a 3×3 convolution kernel, and a local illumination perception feature map is output; the number of channels of the local illumination perception feature map is compressed to 1 through convolution calculation, and the output value range is limited to between 0 and 1 through the Sigmoid function as illumination perception information.

[0008] According to the lightweight target detection method based on multimodal images disclosed in the present invention, preferably, the calculation process of the parameter-free channel attention module specifically includes: performing global average pooling operations and global average pooling operations on the spliced ​​feature map; adding the outputs of the global average pooling and the global maximum pooling to obtain a fused tensor; limiting the value range of the fused tensor to between 0 and 1 through the Sigmoid function to generate channel-level attention weights; expanding the attention weights to the same shape as the spliced ​​feature map, and multiplying them element by element to adjust the feature response of each channel of the spliced ​​feature map.

[0009] The second aspect of the present invention discloses a lightweight target detection system based on multimodal images, comprising: a memory for storing program instructions; a processor for calling the program instructions stored in the memory to implement a lightweight target detection method based on multimodal images as any of the above-mentioned technical solutions.

[0010] The beneficial effects of the present invention include at least: the present invention is a lightweight target detection algorithm based on the dynamic adjustment and fusion of multiple weights of multimodal images, which combines weight-guided feature optimization, local illumination perception mechanism and parameter-free channel attention mechanism to form a multiple weight adjustment module (MWAF), which complementarily fuses the features of infrared and visible light images at the input end, effectively overcomes the limitations of a single modality in complex environments without generating a large amount of calculation, thereby significantly improving the accuracy and stability of target detection tasks. Specifically: The feature optimization method based on weight guidance can generate a weight matrix through feature masks according to the importance of image features of different modalities, and perform weighted processing, which can make full use of the complementarity between the modalities and avoid the loss of key features caused by simple feature superposition. Secondly, selective weighted processing can alleviate the problem of information redundancy, highlight only the features that contribute to the detection task, thereby avoiding the interference of invalid or repeated information and improving the computational efficiency and detection accuracy of the model. The Local Illumination Perception Module (LIPM) captures local illumination information, perceives local illumination changes in the input image, and dynamically adjusts the fusion weights of multimodal features based on this information, effectively reducing feature distortion caused by large illumination differences in the image, ensuring that the location information and detail features of the target can still be accurately extracted under complex illumination conditions, thereby significantly improving the performance of the model in multimodal tasks. The Parameter-Free Channel Attention (PFCA) module uses two parameter-free operations, global average pooling (GAP) and global maximum pooling (GMP), to recalibrate the features of different channels, enhance the expressiveness of important features, and suppress the interference of irrelevant or redundant information, thereby improving the model's ability to extract deep semantic information of images and small target details. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 A schematic diagram of the network structure of a lightweight object detection method based on multimodal images according to an embodiment of the present invention is shown.

[0012] Figure 2 A schematic diagram of the network structure of a multiple weight adjustment module according to an embodiment of the present invention is shown.

[0013] Figure 3 A schematic diagram of the network structure of a local illumination perception module according to an embodiment of the present invention is shown.

[0014] Figure 4 A schematic diagram of the network structure of a parameter-free channel attention module according to an embodiment of the present invention is shown.

[0015] Figure 5 A schematic block diagram of a lightweight object detection system based on multimodal images according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0016] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention can also be implemented in other ways different from those described herein, and therefore, the present invention is not limited to the limitations of the specific embodiments disclosed below.

[0017] According to an embodiment of the present invention, in view of the fact that the performance of the current target detection scheme will be greatly reduced in a low-light environment, and the early fusion method will have the shortcomings of insufficient complementary utilization and information redundancy, a lightweight target detection method based on multimodal images is proposed, comprising the following steps:

[0018] Step 1: Build a lightweight object detection network based on dynamic adjustment and fusion of multiple weights of multimodal images;

[0019] Step 2: Input the paired infrared and visible light images into the network as training sets, perform network training, and adjust network parameters;

[0020] Step 3: Verify the network model obtained in step 2 and use the test set for final evaluation to ensure the robustness and generalization ability of the model in different scenarios;

[0021] Step 4: Input the infrared and visible light image pairs into the model verified in step 3 for real-time target detection.

[0022] like Figure 1 and Figure 2As shown, the lightweight target detection network based on dynamic adjustment and fusion of multiple weights of multimodal images proposed in this embodiment includes: a multiple weight adjustment module (MWAF), a backbone network (Backbone) of the target detection network, a neck (Neck) and a detection head (Head). The present invention can effectively fuse the features of infrared and visible light images by adding a multiple weight adjustment module before the backbone network of the detection network. The fused feature map obtained after processing by the multiple weight adjustment module is sent to the backbone network of the target detection network for further feature extraction. In the Neck part of the target detection network, features of multiple resolutions are aggregated to enhance the detection accuracy. Finally, the final target bounding box and category prediction are generated through the detection head part. To achieve higher-precision target detection.

[0023] like Figure 2 As shown in the figure, the multiple weight adjustment module (MWAF) specifically includes a weight-guided feature optimization method, a local illumination perception module, and a parameter-free attention module. In the multiple weight adjustment module, the input image is first processed by the weight-guided feature optimization method: first, the input 6-channel tensor is divided into two parts, the first 3 channels are visible light images (RGB images), and the last 3 channels are infrared images (IR images). Then, two convolution blocks are used to convolve, batch normalize, and ReLU activate the RGB and IR features respectively to generate enhanced features. This process can be expressed as:

[0024]

[0025]

[0026] Among them, f R_en is the enhanced RGB feature, f I_en is the enhanced IR feature, X R is the original RGB feature, X I is the original IR feature.

[0027] Then, a 1×1 convolution operation is used to reduce the number of channels of the feature map to 1, and the Sigmoid function is used to limit the value range of its output to between 0 and 1 to generate a weight mask. The process of generating a feature map mask is:

[0028]

[0029]

[0030] Among them, mask R , mask I are the weight masks of RGB image and IR image respectively, and σ represents the Sigmoid function.

[0031] Finally, the generated RGB mask and IR mask weight the original feature map pixel by pixel to generate the optimized RGB feature map and IR feature map. The process is as follows:

[0032]

[0033]

[0034] Among them, f R_m , f I_m They are the optimized RGB features and IR features respectively. It is an element-wise multiplication operation.

[0035] The above-mentioned weight-guided feature optimization method can weight the key information in the RGB feature map and the IR feature map, ensure that important areas are highlighted, and effectively filter out unimportant information.

[0036] like Figure 3 As shown in Figure 1, the calculation process of the local illumination perception module includes: the input features are first extracted through a 3×3 convolution kernel. The convolution kernel operation is expressed as:

[0037]

[0038]

[0039] Among them, Y R and Y I Represent the lighting information of RGB and IR images respectively.

[0040] Then, in order to further reduce the channel dimension and extract local illumination perception information, a 1×1 convolution operation is introduced to compress the number of channels to 1. The specific operation is as follows:

[0041]

[0042]

[0043] Among them, W R and W I Represents the illumination weights of RGB and IR images respectively. The local illumination value is reduced by 1 in order to concentrate the illumination information, so as to intuitively evaluate whether the illumination intensity of the current area is stronger or weaker than the baseline. The Sigmoid function compresses the result value to between 0 and 1 for subsequent adjustment of the weight factor.

[0044] After obtaining the local illumination characteristics (W R and W I), and then use this feature to adjust the image features of different modalities. In this implementation, the local illumination feature is used to adjust the weights between the IR image and the RGB image.

[0045] The weighted result of RGB features is:

[0046]

[0047] The weighted result of infrared features is:

[0048]

[0049] This weighting mechanism is based on local illumination features, which enhances the expression of infrared features in low-light areas, while suppressing the influence of RGB features in overly bright areas, ensuring the maintenance of illumination consistency during multimodal image fusion.

[0050] Subsequently, the fused feature map passes through residual connection and parameter-free channel attention module to further improve the representation ability of important features.

[0051] like Figure 4 As shown in the parameter-free channel attention module, the input feature map is first subjected to GAP (global average pooling) and GMP (global maximum pooling) operations: the average value X of all pixel values ​​on each channel is calculated avg and the maximum value X max ;

[0052] Add the outputs of global average pooling and global maximum pooling to get a tensor F that combines the two feature descriptions c :

[0053] F c =X avg +X max

[0054] The fused feature description Fc is used to generate the channel-level attention weight W through the Sigmoid activation function c :

[0055]

[0056] Among them, σ represents the Sigmoid function, which maps the input value to between (0,1).

[0057] Finally, the attention weight W c Expanded to the same shape as the input feature map, and multiplied element by element to adjust the feature response of each channel to obtain the output feature map F e ;

[0058] Through the above-mentioned parameter-free attention method, without introducing additional parameters, the feature response between channels can be adaptively adjusted, so that the model has stronger generalization ability in different scenarios, improving the overall performance while maintaining computational efficiency, and successfully enhancing the model's attention to key features.

[0059] like Figure 5 As shown, according to another embodiment of the present invention, a lightweight target detection system 500 based on multimodal images is also disclosed, including: a memory 501, used to store program instructions; a processor 502, used to call the program instructions stored in the memory to implement the lightweight target detection method based on multimodal images as in the above embodiment.

[0060] In summary, the present invention extracts features from the input infrared image and visible light image respectively, and ensures that effective fusion features can be performed under different lighting conditions through weight-guided feature optimization methods and local illumination perception modules. Subsequently, the fused feature map is further improved through residual connection and parameter-free channel attention module to further enhance the representation ability of important features. The fused features processed by the MWAF module are sent to the YOLO network (for example, YOLOv8) for target detection. The present invention successfully achieves stable and accurate target detection in complex environments (such as low light, occlusion, etc.) through an innovative multiple weight dynamic adjustment mechanism (weight guidance, illumination perception weight, attention weight). Compared with traditional single-modality and static weight fusion methods, the present invention significantly improves the robustness and reliability of detection. The present invention achieves excellent detection accuracy and subjective feeling with a simple and efficient network structure. In addition, the entire method is designed with special consideration of the lightweight characteristics of the network, so while ensuring performance improvement, it does not introduce a larger amount of parameters and calculations, and maintains a low resource occupancy. This enables the present invention to be widely applicable in embedded and real-time application scenarios.

[0061] All or part of the steps in the various methods of the above embodiments can be completed by controlling the relevant hardware through a program, and the program can be stored in a readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other readable medium that can be used to carry or store data.

[0062] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A lightweight target detection method based on multimodal images, characterized in that: include: Receiving images to be detected: receiving a pair of infrared images and visible light images; Feature optimization based on weight guidance: performing convolution, batch normalization and ReLU activation operations on the infrared image and the visible light image respectively to generate an enhanced infrared feature map and an enhanced visible light feature map; The number of channels of the enhanced infrared feature map is reduced to 1 by using a convolution operation, and the output value range is limited to between 0 and 1 by using a Sigmoid function to generate an infrared feature weight mask; The number of channels of the enhanced visible light feature map is reduced to 1 by using a convolution operation, and the output value range is limited to between 0 and 1 by using a Sigmoid function to generate a visible light feature weight mask; Based on the infrared feature weight mask, the enhanced infrared feature map is weighted pixel by pixel to generate an optimized infrared feature map; based on the visible light feature weight mask, the enhanced visible light feature map is weighted pixel by pixel to generate an optimized visible light feature map; Local illumination perception mechanism: inputting the enhanced infrared feature map and the enhanced visible light feature map into a local illumination perception module to extract illumination perception information; Adjusting the weights of the optimized infrared feature map and the optimized visible light feature map according to the light perception information to generate a visible light weighted feature map and an infrared weighted feature map; splicing the visible light weighted feature map and the infrared weighted feature map in the channel dimension to generate a spliced ​​feature map; Parameter-free channel attention mechanism: The concatenated feature map is subjected to subsequent feature extraction operations through residual connection and parameter-free channel attention module to obtain a multimodal fusion feature map; Target detection: The multimodal fusion feature map is input into the target detection network for target detection to output the detection result.

2. The lightweight target detection method based on multimodal images according to claim 1, characterized in that: The target detection network is a YOLO network.

3. The lightweight target detection method based on multimodal images according to claim 2, characterized in that: The YOLO network specifically includes: a backbone network, a neck and a detection head; the multimodal fusion feature map is further subjected to feature extraction in the backbone network part, and features of multiple resolutions are aggregated in the neck of the target detection network to enhance detection accuracy, and finally the final target bounding box and category prediction are generated by the detection head.

4. The lightweight target detection method based on multimodal images according to claim 1, characterized in that: The calculation process of the local illumination perception module specifically includes: The input features are convolved through a 3×3 convolution kernel, batch normalized, and ReLU activated to output a local illumination perception feature map. The number of channels of the local illumination perception feature map is compressed to 1 through convolution calculation, and the output value range is limited to between 0 and 1 through the Sigmoid function as illumination perception information.

5. The lightweight target detection method based on multimodal images according to claim 1, characterized in that: The calculation process of the parameter-free channel attention module specifically includes: Performing a global average pooling operation and a global average pooling operation on the spliced ​​feature map; Add the outputs of global average pooling and global maximum pooling to get the fused tensor; The value range of the fused tensor is limited to between 0 and 1 through the Sigmoid function to generate channel-level attention weights; The attention weights are expanded to the same shape as the concatenated feature map and multiplied element-wise to adjust the feature response of each channel of the concatenated feature map.

6. A lightweight target detection system based on multimodal images, characterized in that: include: A memory for storing program instructions; A processor, configured to call the program instructions stored in the memory to implement the lightweight target detection method based on multimodal images as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Target detection method and system using illumination guidance and attention mechanism

    CN115131640A

  • RGB-T target tracking method based on multi-modal hierarchical relation modeling

    CN116580275A

  • Characteristic decomposition-based infrared and visible light image fusion method for power grid transmission line

    CN118587545A

  • Infrared and visible light image fusion method based on illumination adaptation and attention guidance

    CN118918019A

  • Denoising method and system for inspection image of power transmission line

    CN119540566A

Cited By

  • Circuit board tiny short circuit detection method and system based on machine vision

    CN121708386A

  • Method and system for detecting fine short circuit of circuit board based on machine vision

    CN121708386B