Target detection method and system based on visible light and infrared light feature fusion

By extracting and fusing the features of visible and infrared light images captured by the drone in a low-light environment, and using the U-shaped network to segment the image features, the problem of low target detection accuracy and stability in the prior art is solved, and a more accurate and reliable target detection effect is achieved.

CN120219408AInactive Publication Date: 2025-06-27CHINA TELECOM UNMANNED TECHNOLOGY (JIANGSU) CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510695738.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, multi-sensor fusion method has low target detection accuracy and low detection results stability in low light environments of drones, and cannot provide accurate and reliable target detection capabilities.

Method used

By extracting the initial features of the visible light image and the infrared light image, differential features are generated, and the differential features are added and fused with the initial features to generate fusion features. Then, the dual-light fusion image features are segmented using a U-shaped network to generate the outline of the target.

Benefits of technology

It improves the accuracy and stability of target detection, enhances the target detection capabilities of the drone in low-light environments, and provides more accurate and reliable detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219408A_ABST
    Figure CN120219408A_ABST
Patent Text Reader

Abstract

The invention relates to a target detection method and system based on visible light and infrared light feature fusion, and relates to the technical field of computer vision, and the method comprises the steps: determining the difference feature of the initial features of a visible light image and an infrared light image; adding and fusing the difference features with the initial features of the visible light image and the infrared light image to generate fusion features of the visible light image and fusion features of the infrared light image; respectively extracting target features of fusion features of the visible light image and the infrared light image; fusing the target features of the fusion features of the visible light image and the infrared light image to generate dual-light fusion image features; and segmenting the features of the dual-light fusion image to generate the contour of the target. The feature extraction module is utilized to deeply mine deep association among different modal features, and the dual-light fusion image features are optimized on the basis of improving the description capability of the features to the target, so that the dual-light fusion image features are more accurately used for target detection, and the accuracy and reliability of unmanned aerial vehicle target detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and particularly to an object detection method and system based on the feature fusion of visible light and infrared light. Background Art

[0002] In the field of drone aerial photography, when facing weak light conditions such as fog, cloudy days, and nights, the images captured by the visible light cameras carried by drones have poor imaging quality, with problems such as low brightness and poor contrast. In this case, conventional object detection methods are difficult to effectively detect objects, greatly limiting the application of drones in these complex environments.

[0003] Currently, drones on the market often carry visible light cameras and infrared cameras. Visible light cameras work by capturing natural light, and the information they obtain is the same as what is seen by the human eye, capable of providing rich texture and color information. However, in low light environments, their performance is severely restricted. In contrast, infrared cameras belong to active imaging. Although their resolution is relatively low and they lack texture and color details, they have unique advantages. That is, in low light environments such as at night, they can still obtain the thermal radiation information of the target, successfully breaking through the limitation of lighting conditions. Thus, it can be seen that visible light cameras and infrared cameras have a certain degree of complementarity in capturing imaging characteristics.

[0004] In the prior art, for the problem of difficult object detection by drones under low light conditions, there are mainly the following two technical solutions. One is from the perspective of image enhancement. Its principle is to use deep learning algorithms to learn a large number of images captured under different lighting conditions, thereby training a model that can adaptively adjust for low light images. The other is from the perspective of multi-sensor fusion. Specifically, the images of the visible light camera and the infrared camera are subjected to feature extraction, and then the image features of the two cameras are superimposed through methods such as superposition and merging. Finally, the feature vector is input into a neural network model, and the powerful classification and recognition capabilities of the neural network are used to achieve object detection.

[0005] The method based on image enhancement mainly relies on constructing a neural network to enhance the brightness and contrast of images. However, in actual applications, there are many interference factors in complex backgrounds, and the reflected light of different objects is intertwined, making the image features extremely complex. In a low signal-to-noise ratio environment, the signal is weak and vulnerable to noise interference, resulting in blurred image details. This method based on image enhancement only operates on visible light images, has poor robustness in such complex scenarios, is difficult to accurately and stably extract target features, has a poor effect on improving the accuracy of object detection, and cannot reliably cope with complex and changeable low light environments.

[0006] Based on the multi-sensor fusion method, the features of visible light and infrared camera images are superimposed and a neural network is used to detect targets. However, the existing fusion methods only simply superimpose, without deeply exploring the complementarity of the two imaging methods in characteristics. For example, during the fusion process, it is difficult to effectively match the low-resolution features of infrared images with the high-resolution texture features of visible light images, resulting in the loss of some key information. This directly affects the accuracy of the final target detection, and the stability of the detection results is also greatly reduced, making it impossible to provide accurate and reliable target detection capabilities for drones under low-light conditions. Summary of the Invention

[0007] Based on this, aiming at the problems of low accuracy and low stability and reliability of the final target detection in the multi-sensor fusion method, a target detection method and system based on the fusion of visible light and infrared light features are provided, which can deeply explore the complementarity of the features of visible light images and infrared camera images, so that the fused features contain more comprehensive and accurate target information, and help to improve the accuracy and stability of target detection.

[0008] This application provides a target detection method based on the fusion of visible light and infrared light features, including: Extract the initial features of the visible light image and the initial features of the infrared light image; Generate the difference features of the visible light image and the infrared light image according to the initial features of the visible light image and the initial features of the infrared light image; Add and fuse the difference features with the initial features of the visible light image and the initial features of the infrared light image respectively to generate the fusion features of the visible light image and the fusion features of the infrared light image; Extract the target features of the fusion features of the visible light image and the target features of the fusion features of the infrared light image; Fuse the target features of the fusion features of the visible light image and the target features of the fusion features of the infrared light image to generate dual-light fusion image features; Based on the U-shaped network, segment the dual-light fusion image features to generate the contour of the target.

[0009] In one embodiment, the process of extracting the features of an image or the features of the features includes: extracting the local features of the input image or the input features through convolution operations, normalizing the local features, and introducing non-linear factors to the normalized local features using the rectified linear unit activation function.

[0010] In one embodiment, generating the difference features of the visible light image and the infrared light image according to the initial features of the visible light image and the initial features of the infrared light image includes: Perform a subtraction operation on the initial features of the visible light image and the initial features of the infrared light image to generate a feature difference value; Based on the preset convolutional block attention module, global average pooling and global maximum pooling operations are performed on the feature difference values in the channel dimension and the spatial dimension respectively to generate difference features.

[0011] In one embodiment, adding and fusing the difference features with the initial features of the visible light image and the initial features of the infrared light image respectively includes: Using a convolutional layer with a convolution kernel of 1×1 to adjust the dimension of the difference features; Adding and fusing the difference features with adjusted dimensions to the initial features of the visible light image and the initial features of the infrared light image respectively.

[0012] In one embodiment, after fusing the target features of the fused features of the visible light image and the target features of the fused features of the infrared light image, it further includes: Extracting and fusing the features of the dual - light fused image features to generate target dual - light fused image features.

[0013] In one embodiment, the U - shaped network includes an improved U - shaped network with an SE attention module embedded at the output end of the down - sampling process of the U - shaped network.

[0014] In a second aspect, the present application further provides an object detection system based on visible - light and infrared - light feature fusion, including: An image feature extraction module, configured to extract the initial features of the visible light image and the initial features of the infrared light image; A difference feature generation module, configured to generate difference features between the visible light image and the infrared light image according to the initial features of the visible light image and the initial features of the infrared light image; A first feature fusion module, configured to add and fuse the difference features with the initial features of the visible light image and the initial features of the infrared light image respectively to generate the fused features of the visible light image and the fused features of the infrared light image; A fused - feature extraction module, configured to extract the target features of the fused features of the visible light image and the target features of the fused features of the infrared light image; A second feature fusion module, configured to fuse the target features of the fused features of the visible light image and the target features of the fused features of the infrared light image to generate dual - light fused image features; A segmentation module, configured to segment the dual - light fused image features based on the U - shaped network to generate the contour of the object.

[0015] In a third aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented: Extracting the initial features of the visible light image and the initial features of the infrared light image; Generate the difference features of the visible light image and the infrared light image based on the initial features of the visible light image and the initial features of the infrared light image; Add and fuse the difference features with the initial features of the visible light image and the initial features of the infrared light image respectively to generate the fusion features of the visible light image and the fusion features of the infrared light image; Extract the target features of the fusion features of the visible light image and the target features of the fusion features of the infrared light image; Fuse the target features of the fusion features of the visible light image and the target features of the fusion features of the infrared light image to generate the dual-light fusion image features; Segment the dual-light fusion image features based on the U-shaped network to generate the contour of the target.

[0016] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. The computer program, when executed by a processor, implements the following steps: Extract the initial features of the visible light image and the initial features of the infrared light image; Generate the difference features of the visible light image and the infrared light image based on the initial features of the visible light image and the initial features of the infrared light image; Add and fuse the difference features with the initial features of the visible light image and the initial features of the infrared light image respectively to generate the fusion features of the visible light image and the fusion features of the infrared light image; Extract the target features of the fusion features of the visible light image and the target features of the fusion features of the infrared light image; Fuse the target features of the fusion features of the visible light image and the target features of the fusion features of the infrared light image to generate the dual-light fusion image features; Segment the dual-light fusion image features based on the U-shaped network to generate the contour of the target.

[0017] In a fifth aspect, the present application provides a computer program product, including a computer program. The computer program, when executed by a processor, implements the following steps: Extract the initial features of the visible light image and the initial features of the infrared light image; Generate the difference features of the visible light image and the infrared light image based on the initial features of the visible light image and the initial features of the infrared light image; Add and fuse the difference features with the initial features of the visible light image and the initial features of the infrared light image respectively to generate the fusion features of the visible light image and the fusion features of the infrared light image; Extract the target features of the fusion features of the visible light image and the target features of the fusion features of the infrared light image; Fuse the target features of the fusion features of the visible light image and the target features of the fusion features of the infrared light image to generate the dual-light fusion image features; Based on the U-shaped network to segment the dual-light fusion image features and generate the contour of the target.

[0018] A target detection method and system based on the fusion of visible light and infrared light features provided by this application have the following specific technical effects: 1. By obtaining the differential features of the visible light image and the infrared light image and fusing them with the first initial feature and the second initial feature, then using the feature extraction module to extract the target features of the fusion features of the visible light image and the target features of the fusion features of the infrared light image and fuse them into dual-light fusion image features, and finally using the feature extraction module to optimize the dual-light fusion image features, so as to optimize the dual-light fusion image features on the basis of deeply mining the deep-level association between different modal features by using the feature extraction module and improving the description ability of the features for the target, and further making the dual-light fusion image features more accurate for subsequent target detection, improving the accuracy and reliability of UAV target detection.

[0019] 2. In this application, when performing feature extraction, it is extracted by using a preset feature extraction module. The feature extraction module is an improved neural network, including CNN, BN, and ReLU, which can improve the semantic segmentation accuracy, enhance the response sensitivity to small targets, reduce missed detections and misjudgments, and when facing image samples with multiple interference noises, it can adaptively adjust the feature channel weights, suppress interference, and ensure the stable and accurate detection of small target contours.

[0020] 3. The improved U-shaped network in this application has a stronger ability to capture the details of small targets from the UAV perspective, further improving the reliability of overall target detection. Description of the Drawings

[0021] Figure 1 It is a framework diagram of a dual-light fusion neural network based on visible light and infrared light image features in an embodiment; Figure 2 It is a schematic diagram of a target detection method based on the fusion of visible light and infrared light features in an embodiment; Figure 3 It is a flowchart of a target detection method based on the fusion of visible light and infrared light features in an embodiment; Figure 4 It is a framework diagram of a convolutional block attention module in an embodiment; Figure 5 It is a framework diagram of an improved U-shaped segmentation network in an embodiment; Figure 6 It is a schematic diagram of an SE module in an embodiment. Detailed Embodiment

[0022] In order to make the objectives, technical solutions and advantages of this application clearer, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely used to explain this application and are not used to limit this application.

[0023] Referring to Figure 1 , this patent constructs a dual - light fusion neural network framework based on visible - light and infrared - light image features. This framework scientifically divides the process of drone target detection under low - light conditions into four parts: image input, feature fusion, semantic segmentation, and target contour output. In the image input stage, the framework receives the images captured by the visible - light camera and the infrared camera carried by the drone, making full use of the information obtained by the two cameras under different imaging characteristics. In the feature fusion link, through the constructed feature extraction module, different feature information of the visible - light image and the infrared - light image is extracted and organically fused together, deeply exploring their complementarity and avoiding information conflicts or losses caused by simple superposition. The semantic segmentation process identifies the targets to be detected on the fused feature map and carefully divides the target contours. Finally, in the target contour output stage, the framework outputs the pixel coordinates corresponding to the target contours, providing accurate data support for subsequent drone decision - making and operations.

[0024] The feature extraction module provided by this application includes: a convolutional neural network layer, a batch normalization layer, and a non - linear activation layer. Among them, the convolutional neural network layer is used to capture local image features through convolutional operations; the batch normalization layer is used to normalize local features; the non - linear activation layer is used to introduce non - linear factors into the normalized local features using the rectified linear unit activation function.

[0025] The process of extracting features of an image or features of features using the feature extraction module includes: extracting local features of the input image or input features through convolutional operations, normalizing the local features, and introducing non - linear factors into the normalized local features using the rectified linear unit activation function.

[0026] The convolutional neural network layer (CNN) captures local image features, such as edges, textures, etc., through convolutional operations, providing basic feature information for subsequent feature analysis and target recognition. The batch normalization layer (BN) normalizes data, accelerates network convergence, makes network training more stable, and reduces training difficulties caused by changes in data distribution. The non - linear activation layer introduces non - linear factors using the rectified linear unit activation function (ReLU), enhancing the network's ability to express complex patterns and enabling the network to learn more complex image feature relationships.

[0027] The feature extraction module provided in this application is an improved neural network, which helps to improve the accuracy of semantic segmentation, and can more accurately segment small targets at the pixel level and accurately outline their contours. Furthermore, the feature extraction module improves the response sensitivity to small targets, effectively responds to small targets with rich details and irregular shapes, and reduces missed judgments and misjudgments. In addition, the feature extraction module enhances the generalization ability of the model. In complex environments such as different lighting and weather, it can adaptively adjust the feature channel weights, suppress interference, and ensure stable and accurate detection of small target contours in the face of image samples containing multiple interference noises.

[0028] Reference Figure 2 and Figure 3 Based on the above-mentioned dual-light fusion neural network framework based on visible light and infrared light image features, the present application provides a target detection method based on visible light and infrared light feature fusion, including: S100, extracting initial features of the visible light image and initial features of the infrared light image.

[0029] The visible light images and infrared light images taken by the drone are used as the input of the feature extraction module. After being processed by the feature extraction module, the initial features of the visible light images and the initial features of the infrared images are generated, laying the foundation for subsequent feature fusion and processing.

[0030] This application uses visible light images and infrared images taken by drones as input, with the aim of obtaining information about the target scene under different spectra. Visible light images present visual features such as texture and color of the target, and infrared images can capture the thermal radiation information of the target. The combination of the two provides a multi-source feature basis for subsequent fusion processing, thereby enhancing the comprehensive understanding and analysis capabilities of the target.

[0031] S200 , generating difference features between the visible light image and the infrared light image according to initial features of the visible light image and initial features of the infrared light image.

[0032] S210 , performing a subtraction operation on the initial features of the visible light image and the initial features of the infrared light image to generate a feature difference value.

[0033] The initial features of the generated visible light image and the initial features of the infrared image are subtracted to obtain the difference information between the visible light image and the infrared image at the feature level. These differences reflect the differences in target characteristics under different spectra and help to mine complementary features.

[0034] S220, based on the preset convolutional block attention module, perform global average pooling and global maximum pooling operations on the feature difference values ​​in the channel dimension and the spatial dimension respectively to generate difference features.

[0035] Reference Figure 4, the feature difference value is input into the Convolutional Block Attention Module (CBAM) to generate difference features. In the channel dimension, spatial information is aggregated through global average pooling and global max pooling, and the channel attention weights are learned by a multi-layer perceptron to highlight key channel features and suppress redundant channels; in the spatial dimension, spatial attention weights are obtained through channel-based max pooling and average pooling to focus on important spatial regions. The CBAM module enables the network to pay more attention to key features, suppress unimportant features, improve the feature utilization efficiency, and enhance the network's ability to extract effective information.

[0036] S300, add and fuse the difference features with the initial features of the visible light image and the initial features of the infrared light image respectively to generate the fused features of the visible light image and the fused features of the infrared light image.

[0037] S310, use a convolutional layer with a convolution kernel of 1×1 to adjust the dimension of the difference features.

[0038] The difference features processed by the CBAM module are input into a convolutional layer with a convolution kernel of 1×1. Without changing the spatial dimension, the number of channels is adjusted to achieve feature dimensionality reduction or increase, further fuse the features, make the feature representation more compact and efficient, prepare for subsequent feature fusion, and optimize the feature structure to better adapt to the subsequent processing of the network.

[0039] S320, add and fuse the difference features with adjusted dimensions with the initial features of the visible light image and the initial features of the infrared light image respectively to generate the first fused features of the visible light image and the second fused features of the infrared light image.

[0040] The features processed by the convolutional layer are added to the initial features of the visible light image and the initial features of the infrared light image to generate the fused features of the visible light image and the fused features of the infrared light image. This skip connection method fuses the feature information of different stages and paths, retains the useful information of the original features, prevents gradient disappearance, enhances the network's comprehensive expression of multi-modal information, makes the fused features contain more comprehensive image information, and provides a rich feature basis for subsequent in-depth analysis.

[0041] S400, extract the target features of the fused features of the visible light image and the target features of the fused features of the infrared light image.

[0042] The fused features of the visible light image and the fused features of the infrared light image after preliminary fusion are input into the feature extraction module again. Through further convolutional operations, normalization processing, and non-linear activation, more advanced and abstract target features are extracted, the deep-level associations between different modal features are mined, and the ability of the features to describe the target is improved.

[0043] S500 fuses the target features of the fused features of the visible light image and the target features of the fused features of the infrared light image to generate dual-light fused image features.

[0044] Additively fuse the target features of the fused features of the visible light image and the target features of the fused features of the infrared light image extracted by the feature extraction module to generate the final feature representation, i.e., dual-light fused image features.

[0045] S510 extracts and fuses the features of the dual-light fused image features to generate target dual-light fused image features.

[0046] By re-inputting the dual-light fused image features into the feature extraction module for convolution operation, normalization processing, and non-linear activation operation, the features of the dual-light fused image features are optimized to generate target dual-light fused image features. By outputting the result of the target dual-light fused image features that have been fully fused and optimized, it can be accurately used for target detection in the subsequent segmentation network, improving the accuracy and reliability of target detection during task execution.

[0047] S600 segments the dual-light fused image features based on a U-shaped network to generate the contour of the target.

[0048] Refer to Figure 5 and Figure 6 , in one embodiment, the U-shaped network includes an improved U-shaped network with an SE attention module embedded at the output end of the downsampling process of the U-shaped network.

[0049] For the output target dual-light fused image features, the SE module first performs a compression operation through global average pooling to integrate the feature maps of each channel into a single value to obtain channel global information. Subsequently, an excitation operation is implemented through two fully connected layers, enabling the improved U-shaped network to automatically learn the importance of each feature channel in the small target contour detection task. During this process, non-linear features are incorporated through operations such as activation functions to adapt to the diversity and complexity of small target contours.

[0050] Setting the SE module in the U-shaped network, on the one hand, effectively eliminates irrelevant information such as background interference and noise caused by the shooting perspective in the UAV image. Using the channel weights learned by the SE module, it strengthens the features related to the small target contours under the UAV perspective, suppresses irrelevant features, and avoids noise mixing during the skip connection process, enabling the network to focus on small target contour feature extraction. On the other hand, it reduces the model calculation overhead, reasonably allocates computing resources according to the channel importance screening results, improves the operation efficiency, and quickly processes the image data collected by the UAV. At the same time, it enhances the model's learning ability for foreground small targets, highlights the feature channels related to small targets, and accurately captures small target contour information.

[0051] The improved U-shaped network provided by this application has stronger ability to capture details of small targets from the perspective of drones compared with the traditional U-shaped network, further improving the reliability of overall target detection.

[0052] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps in other steps.

[0053] Based on the same inventive concept, the embodiments of this application also provide a target detection system based on the fusion of visible light and infrared light features. The implementation solutions provided by this system to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the target detection system based on the fusion of visible light and infrared light features provided below can refer to the limitations on the target detection method based on the fusion of visible light and infrared light features in the above text, and will not be repeated here.

[0054] In one embodiment, a target detection system based on the fusion of visible light and infrared light features is provided, including: an image feature extraction module, a difference feature generation module, a first feature fusion module, a fused feature extraction module, a second feature fusion module, and a segmentation module, where: The image feature extraction module is used to extract the initial features of the visible light image and the initial features of the infrared light image.

[0055] The difference feature generation module is used to generate the difference features between the visible light image and the infrared light image according to the initial features of the visible light image and the initial features of the infrared light image.

[0056] The first feature fusion module is used to add and fuse the difference features with the initial features of the visible light image and the initial features of the infrared light image respectively to generate the fused features of the visible light image and the fused features of the infrared light image.

[0057] The fused feature extraction module is used to extract the target features of the fused features of the visible light image and the target features of the fused features of the infrared light image.

[0058] The second feature fusion module is used to fuse the target features of the fusion features of the visible light image and the target features of the fusion features of the infrared light image to generate dual-light fusion image features.

[0059] The segmentation module is used to segment the dual-light fusion image features based on the U-shaped network to generate the contour of the target.

[0060] In one embodiment, the difference feature generation module is further used to perform a subtraction operation on the initial features of the visible light image and the initial features of the infrared light image to generate a feature difference value; based on the preset convolutional block attention module, global average pooling and global maximum pooling operations are respectively performed on the feature difference value in the channel dimension and the spatial dimension to generate a difference feature.

[0061] In one embodiment, the first feature fusion module is further used to adjust the dimension of the difference feature by using a convolutional layer with a convolution kernel of 1×1; the difference feature after dimension adjustment is respectively added and fused with the initial features of the visible light image and the initial features of the infrared light image.

[0062] Each module in the above target detection system based on visible light and infrared light feature fusion can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0063] In one embodiment, a computer device is provided. The computer device can be a server. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a target detection method based on visible light and infrared light feature fusion.

[0064] In one embodiment, a computer device is provided, and the computer device may be a terminal. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a target detection method based on the fusion of visible light and infrared light features.

[0065] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0066] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0067] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0068] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0069] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0070] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A target detection method based on the fusion of visible light and infrared light features, characterized in that, Including: Extracting the initial features of the visible light image and the initial features of the infrared light image; Generating the difference features between the visible light image and the infrared light image according to the initial features of the visible light image and the initial features of the infrared light image; Adding and fusing the difference features with the initial features of the visible light image and the initial features of the infrared light image respectively to generate the fused features of the visible light image and the fused features of the infrared light image; Extracting the target features of the fused features of the visible light image and the target features of the fused features of the infrared light image; Fusing the target features of the fused features of the visible light image and the target features of the fused features of the infrared light image to generate the dual-light fused image features; Segmenting the dual-light fused image features based on the U-shaped network to generate the contour of the target.

2. The method according to claim 1, wherein The process of extracting the features of the image or the features of the features includes: extracting the local features of the input image or the input features through convolutional operations, normalizing the local features, and introducing non-linear factors to the normalized local features by using the rectified linear unit activation function.

3. The method according to claim 1 or 2, characterized in that, Generating the difference features between the visible light image and the infrared light image according to the initial features of the visible light image and the initial features of the infrared light image includes: Performing a subtraction operation on the initial features of the visible light image and the initial features of the infrared light image to generate a feature difference value; Based on the preset convolutional block attention module, performing global average pooling and global maximum pooling operations on the feature difference value in the channel dimension and the spatial dimension respectively to generate the difference features.

4. The method according to claim 1 or 2, characterized in that, Adding and fusing the difference features with the initial features of the visible light image and the initial features of the infrared light image respectively includes: Using a convolutional layer with a convolution kernel of 1×1 to adjust the dimension of the difference features; Adding and fusing the difference features with adjusted dimensions to the initial features of the visible light image and the initial features of the infrared light image respectively.

5. The method according to claim 1 or 2, characterized in that, After fusing the target features of the fused features of the visible light image and the target features of the fused features of the infrared light image, it further includes: Extracting and fusing the features of the dual-light fused image features to generate the target dual-light fused image features.

6. The method according to claim 1 or 2, characterized in that, The U-shaped network includes an improved U-shaped network with an SE attention module embedded at the output end of the downsampling process of the U-shaped network.

7. An object detection system based on the fusion of visible light and infrared light features, characterized in that, Including: An image feature extraction module for extracting the initial features of the visible light image and the initial features of the infrared light image; A difference feature generation module for generating the difference features between the visible light image and the infrared light image according to the initial features of the visible light image and the initial features of the infrared light image; A first feature fusion module for adding and fusing the difference features with the initial features of the visible light image and the initial features of the infrared light image respectively to generate the fused features of the visible light image and the fused features of the infrared light image; A fused feature extraction module for extracting the target features of the fused features of the visible light image and the target features of the fused features of the infrared light image; A second feature fusion module for fusing the target features of the fused features of the visible light image and the target features of the fused features of the infrared light image to generate the dual-light fused image features; A segmentation module for segmenting the dual-light fused image features based on the U-shaped network to generate the contour of the target.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Infrared and visible light image fused multispectral target detection method and system

    CN113688806A

  • Multispectral target detection model training method, target detection method and system

    CN117911710A

  • Unmanned aerial vehicle night vision target recognition system and method based on multi-modal image fusion

    CN117975308A

  • Camouflage target detection method and device, storage medium and computer equipment

    CN118196400A

  • Unmanned aerial vehicle multi-modal remote sensing image target detection method and device based on hybrid Mamb-CNN network

    CN119540786A