Visible light-infrared image target detection method and system based on frequency domain guidance

CN122416381BActive Publication Date: 2026-09-11JIANGXI LIANCHUANG SPECIAL MICROELECTRONICS
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610883254.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-11
Estimated Expiration
2046-06-18

AI Technical Summary

Technical Problem

[0007]本发明旨在解决可见光-红外双模态小目标检测中,频域信息与跨模态空间交互彼此割裂、模态特异性频域特性利用不充分的问题,提供一种能够在频域引导下实现跨模态空间特征协同增强的可见光-红外图像目标检测方法及系统,使之在维持线性计算开销的同时,显著提升复杂环境中远距离小目标的检测鲁棒性与精度

Benefits of technology

[0012] This application presents a frequency-domain guided visible-infrared image target detection method and system. By constructing a collaborative fusion paradigm of "frequency domain guidance and spatial execution," it organically combines a learnable spatial-frequency feature extraction module with a frequency-domain guided bidirectional cross-modal state-space collaborative fusion module. This significantly improves the detection accuracy and robustness of small targets at long distances while maintaining linear computational complexity. Specifically, the learnable spatial-frequency feature extraction module utilizes learnable wavelet bases and learnable frequency weights to adaptively enhance the inherent frequency domain differences between visible light and infrared modes. This effectively highlights the high-frequency texture details of visible light images and the low-frequency thermal contours of infrared images, overcoming the shortcomings of traditional fixed-base frequency domain transformations in adapting to modal differences. The frequency-domain guided bidirectional cross-modal Mamba interaction module drives the gated modulation and state-space modeling of spatial features with frequency-domain modulation signals as conditions. This achieves efficient modeling of cross-modal long-range dependencies, deeply integrating frequency domain priors into the spatial domain feature extraction process. This significantly improves the feature discrimination capability of small targets in complex scenarios such as low contrast, target camouflage, and severe weather.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122416381B_ABST
    Figure CN122416381B_ABST
Patent Text Reader

Abstract

The application discloses a visible light-infrared image target detection method and system based on frequency domain guidance, which comprises the following steps: inputting a visible light and infrared image pair into a double-flow backbone network, performing modal adaptive frequency domain enhancement in a bottleneck structure through a learnable space-frequency feature extraction module, and forming an enhanced feature pyramid; performing fusion by using a bidirectional cross-modal state space collaborative fusion module guided by the frequency domain: extracting a double-modal frequency domain modulation signal, performing gate modulation on spatial features as conditions for each other, and then inputting the spatial features into a state space model for bidirectional interaction, and then performing significant perception gate adaptive fusion to obtain fusion features; and finally, inputting a detection head to perform target classification and positioning. The frequency domain information is used as a guidance condition for cross-modal spatial interaction, a collaborative fusion paradigm of "frequency domain guidance and spatial execution" is constructed, the detection precision and robustness of a small target at a long distance in a complex scene are significantly improved while the linear computational complexity is maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image processing and computer vision technology, and particularly relates to a method and system for target detection in visible light-infrared images based on frequency domain guidance. Background Technology

[0002] Visible light images possess rich texture details, but their image quality deteriorates significantly at night or in low light conditions. Infrared images can operate in all weather conditions and are sensitive to thermal radiation, but they suffer from low spatial resolution and insufficient detail representation. Therefore, visible light-infrared dual-modal fusion target detection has become an important direction for improving all-weather surveillance capabilities. Feature-level fusion, balancing efficiency and characterization ability, is currently the mainstream technical approach.

[0003] The utilization of frequency domain information has been explored in existing technologies. For example, Chinese patent application CN120635411A discloses a multi-scale frequency-space fusion camouflage target detection network, which uses the wavelet transform result of the image and the original image as inputs to two branches respectively. Frequency domain features and spatial domain features are extracted by a two-stream network before fusion. This method has reference value in introducing frequency domain processing, but it treats the frequency domain branch and the spatial domain branch as relatively independent paths processed in parallel, with information interaction only occurring during the fusion stage. This belongs to the "independent extraction first, then fusion" paradigm. Under this paradigm, frequency domain information is only used as one of the supplementary features and fails to substantially play a role in the modeling process of spatial domain feature extraction. The guiding role of frequency domain priors for cross-modal complementarity is not fully utilized.

[0004] On the other hand, state-space sequence models, represented by Mamba, have begun to be introduced into the field of multimodal fusion due to their linear computational complexity and ability to model long sequences. Their core technology involves dynamically adjusting information transmission paths through selective state mechanisms, taking into account both local details and global perception. However, current methods for using Mamba in cross-modal fusion mostly focus on hidden state interactions in the spatial domain. An organic connection has not yet been established between frequency domain analysis and Mamba state-space modeling, resulting in insufficient mining of complementary information between modes when facing challenges such as low contrast, severe weather, and target camouflage.

[0005] Furthermore, most existing frequency domain transforms employ fixed basis functions (such as Haar wavelets) or standard Fourier transforms, which cannot adaptively adjust to the distinct frequency domain characteristics of the two modes: visible light (rich in texture information and prominent high-frequency components) and infrared (smooth thermal profiles and prominent low-frequency components), thus weakening the expression of mode-specific advantages.

[0006] In summary, how to integrate frequency domain priors into the cross-modal spatial interaction process while maintaining linear computational complexity, and form an efficient synergy between frequency domain information and spatial domain modeling, is a key issue that urgently needs to be addressed in the current field of bimodal small target detection. Summary of the Invention

[0007] This invention aims to address the problems in visible-infrared dual-modal small target detection, such as the fragmentation of frequency domain information and cross-modal spatial interaction, and the insufficient utilization of modal-specific frequency domain characteristics. It provides a visible-infrared image target detection method and system that can achieve synergistic enhancement of cross-modal spatial features under frequency domain guidance, thereby significantly improving the robustness and accuracy of detecting long-distance small targets in complex environments while maintaining linear computational overhead.

[0008] In a first aspect, the present invention provides a visible-infrared image target detection method based on frequency domain guidance, comprising: Acquire a registered pair of visible light and infrared images; A visible light image and a corresponding infrared image are fed into a dual-stream backbone network with symmetrical structure and independent parameters. In at least one bottleneck structure of the dual-stream backbone network, a learnable space-frequency feature extraction module is used to perform modally adaptive frequency domain enhancement on the feature map of the input at least one bottleneck structure, and multi-scale features are extracted layer by layer to form the enhanced visible light feature pyramid and infrared feature pyramid. At at least one level of the same scale in the visible light feature pyramid and the infrared feature pyramid, a frequency-domain guided bidirectional cross-modal state-space collaborative fusion module is used to fuse the visible light feature map and the infrared feature map at the at least one level of the same scale to obtain fused features at each scale level. The fusion process includes: Extracting visible light frequency domain modulation signals from visible light feature maps at a certain scale level, and extracting infrared frequency domain modulation signals from infrared feature maps at a certain scale level; Using the infrared frequency domain modulation signal as a condition, the spatial features of the visible light feature map at a certain scale level are gated and modulated, and the first modulation result is sent into the first state space model to obtain the enhanced target visible light features. Using the visible light frequency domain modulation signal as a condition, the spatial features of the infrared feature map at a certain scale level are gated and modulated, and the second modulation result is sent into the second state space model to obtain the enhanced target infrared features. The visible light features and infrared features of the target are adaptively fused to obtain fused features at a certain scale level; The fused features at each scale level are input into the detection head network of the dual-stream backbone network for target classification and localization, and the detection result corresponding to a certain visible light image and infrared image pair is output.

[0009] Secondly, the present invention provides a visible-infrared image target detection system based on frequency domain guidance, comprising: The acquisition module is configured to acquire a registered pair of visible light and infrared images; The extraction module is configured to send a visible light image and a corresponding infrared image into a dual-stream backbone network with symmetrical structure and independent parameters. In at least one bottleneck structure of the dual-stream backbone network, the feature map of the input at least one bottleneck structure is enhanced in the frequency domain by a learnable space-frequency feature extraction module, and multi-scale features are extracted layer by layer to form an enhanced visible light feature pyramid and an infrared feature pyramid. A fusion module is configured to fuse the visible light feature map and the infrared feature map at at least one same-scale level of the visible light feature pyramid and the infrared feature pyramid using a frequency-domain-guided bidirectional cross-modal state-space collaborative fusion module, to obtain fused features at each scale level. The fusion process includes: Extracting visible light frequency domain modulation signals from visible light feature maps at a certain scale level, and extracting infrared frequency domain modulation signals from infrared feature maps at a certain scale level; Using the infrared frequency domain modulation signal as a condition, the spatial features of the visible light feature map at a certain scale level are gated and modulated, and the first modulation result is sent into the first state space model to obtain the enhanced target visible light features. Using the visible light frequency domain modulation signal as a condition, the spatial features of the infrared feature map at a certain scale level are gated and modulated, and the second modulation result is sent into the second state space model to obtain the enhanced target infrared features. The visible light features and infrared features of the target are adaptively fused to obtain fused features at a certain scale level; The output module is configured to input the fused features of each scale level into the detection head network of the dual-stream backbone network, perform target classification and localization, and output the detection result corresponding to a certain visible light image and infrared image pair.

[0010] Thirdly, an electronic device is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of the frequency-domain guided visible-infrared image target detection method according to any embodiment of the present invention.

[0011] Fourthly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein when the program instructions are executed by a processor, the processor performs the steps of the visible-infrared image target detection method based on frequency domain guidance according to any embodiment of the present invention.

[0012] This application presents a frequency-domain guided visible-infrared image target detection method and system. By constructing a collaborative fusion paradigm of "frequency domain guidance and spatial execution," it organically combines a learnable spatial-frequency feature extraction module with a frequency-domain guided bidirectional cross-modal state-space collaborative fusion module. This significantly improves the detection accuracy and robustness of small targets at long distances while maintaining linear computational complexity. Specifically, the learnable spatial-frequency feature extraction module utilizes learnable wavelet bases and learnable frequency weights to adaptively enhance the inherent frequency domain differences between visible light and infrared modes. This effectively highlights the high-frequency texture details of visible light images and the low-frequency thermal contours of infrared images, overcoming the shortcomings of traditional fixed-base frequency domain transformations in adapting to modal differences. The frequency-domain guided bidirectional cross-modal Mamba interaction module drives the gated modulation and state-space modeling of spatial features with frequency-domain modulation signals as conditions. This achieves efficient modeling of cross-modal long-range dependencies, deeply integrating frequency domain priors into the spatial domain feature extraction process. This significantly improves the feature discrimination capability of small targets in complex scenarios such as low contrast, target camouflage, and severe weather. Attached Figure Description

[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 A flowchart illustrating a visible-infrared image target detection method based on frequency domain guidance, as provided in an embodiment of the present invention; Figure 2 This is a structural block diagram of a visible light-infrared image target detection system based on frequency domain guidance, provided in an embodiment of the present invention. Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0015] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0016] Please see Figure 1The diagram shows a flowchart of a visible-infrared image target detection method based on frequency domain guidance according to this application.

[0017] like Figure 1 As shown, the frequency-domain guided visible-infrared image target detection method specifically includes the following steps: Step S101: Obtain a registered pair of visible light and infrared images.

[0018] In this step, images of the same scene are simultaneously acquired using a dual-light camera (visible light camera + infrared thermal imager), and the two images are registered based on camera calibration parameters to ensure a one-to-one correspondence between pixels in the visible light image and the infrared image. Registration can be performed using methods such as affine transformation or perspective transformation.

[0019] Step S102: A visible light image and a corresponding infrared image are respectively fed into a dual-stream backbone network with symmetrical structure and independent parameters. In at least one bottleneck structure of the dual-stream backbone network, the feature map of the input at least one bottleneck structure is enhanced in the frequency domain by a learnable space-frequency feature extraction module, and multi-scale features are extracted layer by layer to form the enhanced visible light feature pyramid and infrared feature pyramid.

[0020] In this step, at a certain scale level of the dual-stream backbone network, the following operations are performed on the feature map of the bottleneck structure input at that scale level: Perform 1×1 convolution on the feature map of a bottleneck structure at a certain scale to compress the channels and obtain the base features; The basic features are subjected to a two-dimensional discrete wavelet transform using a learnable decomposition filter to obtain multiple sub-band features. Each sub-band feature is multiplied by a learnable non-negative scaling factor to obtain weighted sub-band features. The weighted sub-band features are then subjected to an inverse wavelet transform using a learnable reconstruction filter to obtain wavelet-enhanced features. The Fourier transform of the base features is performed to obtain the spectrum. A weight map controlled by learnable parameters is generated based on the frequency coordinates. The weight map is multiplied element by element with the spectrum and then an inverse Fourier transform is performed to obtain the Fourier enhanced features. The wavelet enhancement features and the Fourier enhancement features are respectively weighted by channel attention and then fused through learnable fusion weights to obtain fused frequency domain features. These features are then fused with the base features through residual connections to output space-frequency enhancement features, which serve as the output feature map at a certain scale level. The output feature maps at different scales of each scale level are combined according to their resolution from high to low to form the enhanced visible light feature pyramid and infrared feature pyramid.

[0021] In one specific embodiment, the dual-stream backbone network employs two branches with symmetrical structure but independent parameters and no shared weights. Each branch is based on the YOLOv12 framework, replacing the C3k2 bottleneck module with the learnable space-frequency feature extraction module (LWF-FFT module) designed in this invention. The working process of this module is described in detail below: For a feature map of a certain bottleneck structure as input , For the set of real numbers, The height of the feature map, The width of the feature map. The number of channels in the feature map is first compressed using a 1×1 convolution. , The number of intermediate channels after 1×1 convolution compression is used to obtain the basic features. .

[0022] Then the two enhancement branches are executed in parallel: Learnable wavelet enhancement branch: A two-dimensional discrete wavelet transform is performed on the basic features using a learnable decomposition filter to obtain four sub-band features: low-frequency sub-band. Horizontal high-frequency subband Vertical high-frequency subband Diagonal high-frequency subband For low-frequency subband Horizontal high-frequency subband Vertical high-frequency subband Diagonal high-frequency subband Multiply by the learnable nonnegative scaling factor of the low-frequency subband respectively Horizontal high-frequency subband learnable non-negative scaling factor Learnable nonnegative scaling factor for vertical high-frequency subbands Learnable non-negative scaling factor for diagonal high-frequency subbands The weighted subbands are obtained. Then, an inverse wavelet transform is performed using a learnable reconstruction filter (forming a biorthogonal pair with the decomposition filter) to recover the spatial domain and obtain the wavelet enhancement features. .

[0023] Fourier frequency domain enhancement branch: Perform a two-dimensional real Fourier transform on the basic features to obtain the spectrum. According to frequency coordinates Calculate frequency distance Generate a weighted graph, where, , For the frequency coordinates in the weighted graph The weight value at that location, , All parameters are learnable. After multiplying the weight map and the spectrum element-wise, an inverse Fourier transform is performed to obtain the Fourier enhanced features. .

[0024] The outputs of the two branches are weighted by efficient channel attention (ECA) and then processed by learnable scalar weights. , Summing yields the fused frequency domain features. Finally, the fused frequency domain features are added to the base features via residual connections, and then residually connected to the feature map of a certain bottleneck structure input to output the space-frequency enhanced features.

[0025] For the visible light mode, during module initialization, the initial value of the scaling factor of the high-frequency subband of the wavelet is made greater than that of the low-frequency subband, and the initial value of the high-frequency gain parameter in the Fourier branch is greater than that of the low-frequency gain parameter, in order to enhance the high-frequency texture details; for the infrared mode, the opposite is true, focusing on highlighting the low-frequency thermal profile.

[0026] The above enhancement operations are repeatedly performed at multiple scale levels of the backbone network (such as levels with sampling multiples of 8, 16, and 32), and the feature maps output from each scale level constitute the enhanced visible light feature pyramid and infrared feature pyramid.

[0027] Step S103: At at least one level of the same scale in the visible light feature pyramid and the infrared feature pyramid, the visible light feature map and the infrared feature map at the at least one level of the same scale are fused using a frequency domain-guided bidirectional cross-modal state space collaborative fusion module to obtain the fused features at each scale level.

[0028] In this step, the fusion process includes: extracting the visible light frequency domain modulation signal from the visible light feature map at a certain scale level, and extracting the infrared frequency domain modulation signal from the infrared feature map at a certain scale level.

[0029] Specifically, frequency domain modulation modules with non-shared parameters are configured for the visible light mode and the infrared mode, respectively. The frequency domain modulation module is composed of a learnable wavelet basis and a cascaded 1×1 convolutional layer. The visible light feature map at a certain scale level is input into the frequency domain modulation module of the visible light mode. The number of channels is first expanded by wavelet decomposition and then compressed back to the original number of channels by the convolutional layer to generate the visible light frequency domain modulation signal. The infrared feature map at a certain scale level is input into the frequency domain modulation module of the infrared mode. The number of channels is first expanded by wavelet decomposition and then compressed back to the original number of channels by the convolutional layer to generate the infrared frequency domain modulation signal.

[0030] Using the infrared frequency domain modulation signal as a condition, the spatial features of the visible light feature map at a certain scale level are gated and modulated, and the first modulation result is sent into the first state space model to obtain the enhanced target visible light features.

[0031] Specifically, the visible light feature map at a certain scale level is flattened into a sequence along the spatial dimension and then layer-normalized to obtain a normalized visible light spatial sequence; the infrared frequency domain modulation signal is flattened into a sequence along the spatial dimension and then layer-normalized to obtain a normalized infrared frequency domain sequence; the normalized infrared frequency domain sequence is linearly projected and activated with a sigmoid function, and then multiplied element-wise with the normalized visible light spatial sequence to obtain a first gated modulation result; the first gated modulation result is fed into a first state space model to obtain a state enhancement sequence, which is then layer-normalized and reshaped into a spatial feature of the same size as the visible light feature map at a certain scale level, and residually connected with the visible light feature map at a certain scale level to output the enhanced target visible light feature.

[0032] Using the visible light frequency domain modulation signal as a condition, the spatial features of the infrared feature map at a certain scale level are gated and modulated, and the second modulation result is sent into the second state space model to obtain the enhanced target infrared features.

[0033] Specifically, the infrared feature map at a certain scale level is flattened into a sequence along the spatial dimension and then layer-normalized to obtain a normalized infrared spatial sequence; the visible light frequency domain modulation signal is flattened into a sequence along the spatial dimension and then layer-normalized to obtain a normalized visible light frequency domain sequence; the normalized visible light frequency domain sequence is linearly projected and activated with a sigmoid function, and then multiplied element-wise with the normalized infrared spatial sequence to obtain a second gated modulation result; the second gated modulation result is fed into a second state space model to obtain a state enhancement sequence, which is then layer-normalized and reshaped into a spatial feature of the same size as the infrared feature map at a certain scale level, and residually connected with the infrared feature map at a certain scale level to output the enhanced target infrared feature.

[0034] The visible light features and infrared features of the target are adaptively fused to obtain fused features at a certain scale level.

[0035] Specifically, the visible light features and infrared features of the target are concatenated along the channel dimension to obtain concatenated features; a 1×1 convolution and sigmoid activation are performed on the concatenated features to generate a spatial saliency map, wherein the value of each spatial position in the spatial saliency map is between 0 and 1; the visible light features and infrared features of the target are weighted and summed using the spatial saliency map to obtain a fused feature at a certain scale level.

[0036] Step S104: Input the fusion features of each scale level into the detection head network of the dual-stream backbone network to perform target classification and localization, and output the detection result corresponding to a certain visible light image and infrared image pair.

[0037] In this step, the fusion features of each scale level are used to construct a fusion feature pyramid; The fused feature pyramid is input into the detection head network, which includes a classification branch and a regression branch. The classification branch outputs the class probability of each candidate target, and the regression branch outputs the bounding box coordinates of each candidate target. Based on the category probability and bounding box coordinates, a final detection result is generated, which includes the target category label and location information.

[0038] In summary, the method of this application acquires registered visible light and infrared image pairs, feeds them into a dual-stream backbone network, and performs modally adaptive frequency domain enhancement through a learnable space-frequency feature extraction module in the bottleneck structure to form an enhanced feature pyramid. At the same scale level of the pyramid, a frequency-domain guided bidirectional cross-modal state space collaborative fusion module is used for fusion: extracting dual-modal frequency domain modulation signals, gating and modulating spatial features with each other as conditions, and then feeding them into the state space model for bidirectional interaction, followed by saliency-aware adaptive fusion to obtain fused features; finally, the fused features are input into the detection head for target classification and localization. By using frequency domain information as a guiding condition for cross-modal spatial interaction, a collaborative fusion paradigm of "frequency domain guidance and spatial execution" is constructed, which significantly improves the detection accuracy and robustness of small targets at long distances in complex scenes while maintaining linear computational complexity.

[0039] Please see Figure 2 The diagram shows a structural block diagram of a visible-infrared image target detection system based on frequency domain guidance according to this application.

[0040] like Figure 2 As shown, the visible light-infrared image target detection system 200 includes an acquisition module 210, an extraction module 220, a fusion module 230, and an output module 240.

[0041] The acquisition module 210 is configured to acquire a registered visible light image and an infrared image pair; the extraction module 220 is configured to send a visible light image and a corresponding infrared image into a structurally symmetrical and parameter-independent dual-stream backbone network, and in at least one bottleneck structure of the dual-stream backbone network, perform modally adaptive frequency domain enhancement on the feature map of the input at least one bottleneck structure through a learnable space-frequency feature extraction module, and extract multi-scale features layer by layer to form enhanced visible light feature pyramids and infrared feature pyramids; the fusion module 230 is configured to fuse the visible light feature map and infrared feature map of the at least one same scale level of the visible light feature pyramid and the infrared feature pyramid using a frequency domain-guided bidirectional cross-modal state space collaborative fusion module to obtain fused features at each scale level, wherein the fusion process includes: from a certain Visible light frequency domain modulation signals are extracted from visible light feature maps at a certain scale level, and infrared frequency domain modulation signals are extracted from infrared feature maps at a certain scale level. Using the infrared frequency domain modulation signals as conditions, the spatial features of the visible light feature maps at a certain scale level are gated and modulated, and the first modulation result is fed into a first state space model to obtain enhanced target visible light features. Using the visible light frequency domain modulation signals as conditions, the spatial features of the infrared feature maps at a certain scale level are gated and modulated, and the second modulation result is fed into a second state space model to obtain enhanced target infrared features. The target visible light features and the target infrared features are adaptively fused to obtain fused features at a certain scale level. The output module 240 is configured to input the fused features at each scale level into the detection head network of the dual-stream backbone network for target classification and localization, and output detection results corresponding to a certain visible light image and infrared image pair.

[0042] It should be understood that Figure 2 The modules and references described in the document Figure 1 The steps described in the text correspond to those in the method described above. Therefore, the operations, features, and corresponding technical effects described above also apply to the method described in the text. Figure 2 The various modules in the document will not be described in detail here.

[0043] In other embodiments, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein when the program instructions are executed by a processor, the processor performs the frequency-domain guided visible-infrared image target detection method in any of the above method embodiments. In one embodiment, the computer-readable storage medium of the present invention stores computer-executable instructions, which are configured as follows: Acquire a registered pair of visible light and infrared images; A visible light image and a corresponding infrared image are fed into a dual-stream backbone network with symmetrical structure and independent parameters. In at least one bottleneck structure of the dual-stream backbone network, a learnable space-frequency feature extraction module is used to perform modally adaptive frequency domain enhancement on the feature map of the input at least one bottleneck structure, and multi-scale features are extracted layer by layer to form the enhanced visible light feature pyramid and infrared feature pyramid. At at least one level of the same scale in the visible light feature pyramid and the infrared feature pyramid, a frequency-domain guided bidirectional cross-modal state-space collaborative fusion module is used to fuse the visible light feature map and the infrared feature map at the at least one level of the same scale to obtain fused features at each scale level. The fusion process includes: Extracting visible light frequency domain modulation signals from visible light feature maps at a certain scale level, and extracting infrared frequency domain modulation signals from infrared feature maps at a certain scale level; Using the infrared frequency domain modulation signal as a condition, the spatial features of the visible light feature map at a certain scale level are gated and modulated, and the first modulation result is sent into the first state space model to obtain the enhanced target visible light features. Using the visible light frequency domain modulation signal as a condition, the spatial features of the infrared feature map at a certain scale level are gated and modulated, and the second modulation result is sent into the second state space model to obtain the enhanced target infrared features. The visible light features and infrared features of the target are adaptively fused to obtain fused features at a certain scale level; The fused features at each scale level are input into the detection head network of the dual-stream backbone network for target classification and localization, and the detection result corresponding to a certain visible light image and infrared image pair is output.

[0044] Computer-readable storage media may include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application program required for at least one function; the data storage area may store data created based on the use of the frequency-domain guided visible-infrared image target detection system, etc. Furthermore, the computer-readable storage medium may include high-speed random access memory, and may also include memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the computer-readable storage medium may optionally include memory remotely configured relative to a processor, which can be connected to the frequency-domain guided visible-infrared image target detection system via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0045] Figure 3This is a schematic diagram of the structure of the electronic device provided in the embodiment of the present invention, such as... Figure 3 As shown, the device includes a processor 310 and a memory 320. The electronic device may also include an input device 330 and an output device 340. The processor 310, memory 320, input device 330, and output device 340 can be connected via a bus or other means. Figure 3 Taking a bus connection as an example, the memory 320 is the computer-readable storage medium described above. The processor 310 executes various server functions and data processing by running non-volatile software programs, instructions, and modules stored in the memory 320, thereby implementing the frequency-domain guided visible-infrared image target detection method described in the above embodiment. The input device 330 can receive input digital or character information and generate key signal inputs related to user settings and function control of the frequency-domain guided visible-infrared image target detection system. The output device 340 may include a display screen or other display device.

[0046] The aforementioned electronic device can execute the method provided in the embodiments of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in this embodiment can be found in the method provided in the embodiments of the present invention.

[0047] In one implementation, the above-described electronic device is applied to a frequency-domain guided visible-infrared image target detection system as a client, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: Acquire a registered pair of visible light and infrared images; A visible light image and a corresponding infrared image are fed into a dual-stream backbone network with symmetrical structure and independent parameters. In at least one bottleneck structure of the dual-stream backbone network, a learnable space-frequency feature extraction module is used to perform modally adaptive frequency domain enhancement on the feature map of the input at least one bottleneck structure, and multi-scale features are extracted layer by layer to form the enhanced visible light feature pyramid and infrared feature pyramid. At at least one level of the same scale in the visible light feature pyramid and the infrared feature pyramid, a frequency-domain guided bidirectional cross-modal state-space collaborative fusion module is used to fuse the visible light feature map and the infrared feature map at the at least one level of the same scale to obtain fused features at each scale level. The fusion process includes: Extracting visible light frequency domain modulation signals from visible light feature maps at a certain scale level, and extracting infrared frequency domain modulation signals from infrared feature maps at a certain scale level; Using the infrared frequency domain modulation signal as a condition, the spatial features of the visible light feature map at a certain scale level are gated and modulated, and the first modulation result is sent into the first state space model to obtain the enhanced target visible light features. Using the visible light frequency domain modulation signal as a condition, the spatial features of the infrared feature map at a certain scale level are gated and modulated, and the second modulation result is sent into the second state space model to obtain the enhanced target infrared features. The visible light features and infrared features of the target are adaptively fused to obtain fused features at a certain scale level; The fused features at each scale level are input into the detection head network of the dual-stream backbone network for target classification and localization, and the detection result corresponding to a certain visible light image and infrared image pair is output.

[0048] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.

[0049] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A visible-infrared image target detection method based on frequency domain guidance, characterized in that, include: Acquire a registered pair of visible light and infrared images; A visible light image and its corresponding infrared image are fed into a symmetrical and parameter-independent dual-stream backbone network. In at least one bottleneck structure of the dual-stream backbone network, a learnable space-frequency feature extraction module performs modally adaptive frequency domain enhancement on the feature map of the input bottleneck structure, and extracts multi-scale features layer by layer to construct enhanced visible light and infrared feature pyramids. Specifically, this includes: At a certain scale level of the dual-stream backbone network, perform the following operations on the feature map of the bottleneck structure input at that scale level: Perform 1×1 convolution on the feature map of a bottleneck structure at a certain scale to compress the channels and obtain the base features; The basic features are subjected to a two-dimensional discrete wavelet transform using a learnable decomposition filter to obtain multiple sub-band features. Each sub-band feature is multiplied by a learnable non-negative scaling factor to obtain weighted sub-band features. The weighted sub-band features are then subjected to an inverse wavelet transform using a learnable reconstruction filter to obtain wavelet-enhanced features. The Fourier transform of the base features is performed to obtain the spectrum. A weight map controlled by learnable parameters is generated based on the frequency coordinates. The weight map is multiplied element by element with the spectrum and then an inverse Fourier transform is performed to obtain the Fourier enhanced features. The wavelet enhancement features and the Fourier enhancement features are respectively weighted by channel attention and then fused through learnable fusion weights to obtain fused frequency domain features. These features are then fused with the base features through residual connections to output space-frequency enhancement features, which serve as the output feature map at a certain scale level. The output feature maps at different scales of each scale level are combined according to the resolution from high to low to form the enhanced visible light feature pyramid and infrared feature pyramid. At at least one level of the same scale in the visible light feature pyramid and the infrared feature pyramid, a frequency-domain guided bidirectional cross-modal state-space collaborative fusion module is used to fuse the visible light feature map and the infrared feature map at the at least one level of the same scale to obtain fused features at each scale level. The fusion process includes: Extracting visible light frequency domain modulation signals from visible light feature maps at a certain scale level, and extracting infrared frequency domain modulation signals from infrared feature maps at a certain scale level; Using the infrared frequency domain modulation signal as a condition, the spatial features of the visible light feature map at a certain scale level are gated and modulated, and the first modulation result is sent into the first state space model to obtain the enhanced target visible light features. Using the visible light frequency domain modulation signal as a condition, the spatial features of the infrared feature map at a certain scale level are gated and modulated, and the second modulation result is sent into the second state space model to obtain the enhanced target infrared features. The visible light features and infrared features of the target are adaptively fused to obtain fused features at a certain scale level; The fused features at each scale level are input into the detection head network of the dual-stream backbone network for target classification and localization, and the detection result corresponding to a certain visible light image and infrared image pair is output.

2. The visible-infrared image target detection method based on frequency domain guidance according to claim 1, characterized in that, The extraction of visible light frequency domain modulation signals from a visible light feature map at a certain scale level, and the extraction of infrared frequency domain modulation signals from an infrared feature map at a certain scale level, include: For the visible light mode and the infrared mode, frequency domain modulation modules with non-shared parameters are configured respectively. The frequency domain modulation modules are composed of learnable wavelet bases and cascaded 1×1 convolutional layers. The visible light feature map at a certain scale level is input into the frequency domain modulation module of the visible light mode. The number of channels is first expanded by wavelet decomposition, and then compressed back to the original number of channels by the convolutional layer to generate the visible light frequency domain modulation signal. An infrared feature map at a certain scale level is input into the frequency domain modulation module of the infrared mode. The number of channels is first expanded by wavelet decomposition, and then compressed back to the original number of channels by the convolutional layer to generate an infrared frequency domain modulation signal.

3. The visible-infrared image target detection method based on frequency domain guidance according to claim 1, characterized in that, The step of using the infrared frequency domain modulation signal as a condition to perform gated modulation on the spatial features of a visible light feature map at a certain scale level, and feeding the first modulation result into the first state space model to obtain enhanced target visible light features includes: The visible light feature map at a certain scale level is flattened into a sequence along the spatial dimension and then normalized to obtain a normalized visible light spatial sequence. The infrared frequency domain modulation signal is flattened into a sequence along the spatial dimension and then normalized to obtain a normalized infrared frequency domain sequence. The normalized infrared frequency domain sequence is linearly projected and activated by Sigmoid, and then multiplied element-wise with the normalized visible light spatial sequence to obtain the first gated modulation result. The first gated modulation result is fed into the first state space model to obtain the state enhancement sequence. After layer normalization, it is reshaped into a spatial feature with the same size as the visible light feature map at a certain scale level. The enhanced target visible light feature is then output by residual connection with the visible light feature map at a certain scale level.

4. The visible-infrared image target detection method based on frequency domain guidance according to claim 1, characterized in that, The step of using the visible light frequency domain modulation signal as a condition to perform gated modulation on the spatial features of an infrared feature map at a certain scale level, and then feeding the second modulation result into the second state space model to obtain enhanced target infrared features includes: The infrared feature map at a certain scale level is flattened into a sequence along the spatial dimension and then normalized to obtain a normalized infrared spatial sequence. The visible light frequency domain modulation signal is flattened into a sequence along the spatial dimension and then normalized to obtain a normalized visible light frequency domain sequence. The normalized visible light frequency domain sequence is linearly projected and activated by Sigmoid, and then multiplied element-wise with the normalized infrared spatial sequence to obtain the second gated modulation result. The second gated modulation result is fed into the second state space model to obtain the state enhancement sequence. After layer normalization, it is reshaped into a spatial feature with the same size as the infrared feature map of a certain scale level. The enhanced target infrared feature is then output by residual connection with the infrared feature map of a certain scale level.

5. The visible-infrared image target detection method based on frequency domain guidance according to claim 1, characterized in that, The adaptive fusion of the target's visible light features and the target's infrared features to obtain fused features at a certain scale level includes: The visible light feature and the infrared feature of the target are stitched together along the channel dimension to obtain the stitched feature; Perform a 1×1 convolution and Sigmoid activation on the spliced ​​features to generate a spatial saliency map, wherein the value of each spatial location in the spatial saliency map is between 0 and 1; The visible light features and infrared features of the target are weighted and summed using the spatial saliency map to obtain a fused feature at a certain scale level.

6. The visible-infrared image target detection method based on frequency domain guidance according to claim 1, characterized in that... The process of inputting the fused features at each scale level into the detection head of the dual-stream backbone network for target classification and localization, and outputting the detection result corresponding to a certain visible light image and infrared image pair, includes: The fusion features at each scale level are used to construct a fusion feature pyramid; The fused feature pyramid is input into the detection head network, which includes a classification branch and a regression branch. The classification branch outputs the class probability of each candidate target, and the regression branch outputs the bounding box coordinates of each candidate target. Based on the category probability and bounding box coordinates, a final detection result is generated, which includes the target category label and location information.

7. A visible-infrared image target detection system based on frequency domain guidance, characterized in that, include: The acquisition module is configured to acquire a registered pair of visible light and infrared images; The extraction module is configured to input a visible light image and a corresponding infrared image into a structurally symmetrical and parameter-independent dual-stream backbone network. In at least one bottleneck structure of the dual-stream backbone network, a learnable space-frequency feature extraction module performs modally adaptive frequency domain enhancement on the feature map of the input bottleneck structure, and extracts multi-scale features layer by layer to construct enhanced visible light feature pyramids and infrared feature pyramids. Specifically, this includes: At a certain scale level of the dual-stream backbone network, perform the following operations on the feature map of the bottleneck structure input at that scale level: Perform 1×1 convolution on the feature map of a bottleneck structure at a certain scale to compress the channels and obtain the base features; The basic features are subjected to a two-dimensional discrete wavelet transform using a learnable decomposition filter to obtain multiple sub-band features. Each sub-band feature is multiplied by a learnable non-negative scaling factor to obtain weighted sub-band features. The weighted sub-band features are then subjected to an inverse wavelet transform using a learnable reconstruction filter to obtain wavelet-enhanced features. The Fourier transform of the base features is performed to obtain the spectrum. A weight map controlled by learnable parameters is generated based on the frequency coordinates. The weight map is multiplied element by element with the spectrum and then an inverse Fourier transform is performed to obtain the Fourier enhanced features. The wavelet enhancement features and the Fourier enhancement features are respectively weighted by channel attention and then fused through learnable fusion weights to obtain fused frequency domain features. These features are then fused with the base features through residual connections to output space-frequency enhancement features, which serve as the output feature map at a certain scale level. The output feature maps at different scales of each scale level are combined according to the resolution from high to low to form the enhanced visible light feature pyramid and infrared feature pyramid. A fusion module is configured to fuse the visible light feature map and the infrared feature map at at least one same-scale level of the visible light feature pyramid and the infrared feature pyramid using a frequency-domain-guided bidirectional cross-modal state-space collaborative fusion module, to obtain fused features at each scale level. The fusion process includes: Extracting visible light frequency domain modulation signals from visible light feature maps at a certain scale level, and extracting infrared frequency domain modulation signals from infrared feature maps at a certain scale level; Using the infrared frequency domain modulation signal as a condition, the spatial features of the visible light feature map at a certain scale level are gated and modulated, and the first modulation result is sent into the first state space model to obtain the enhanced target visible light features. Using the visible light frequency domain modulation signal as a condition, the spatial features of the infrared feature map at a certain scale level are gated and modulated, and the second modulation result is sent into the second state space model to obtain the enhanced target infrared features. The visible light features and infrared features of the target are adaptively fused to obtain fused features at a certain scale level; The output module is configured to input the fused features of each scale level into the detection head network of the dual-stream backbone network, perform target classification and localization, and output the detection result corresponding to a certain visible light image and infrared image pair.

8. An electronic device, characterized in that, include: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the method described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-scale frequency-space fusion camouflage target detection method and system

    CN120635411A

  • Visible light-infrared target detection method based on adaptive frequency domain feature fusion

    CN121904528A

  • Vehicle target detection system and method based on Mama and double-domain interaction

    CN122135319A