Multi-modal fusion target detection method and system suitable for dark and weak environment

Through the multimodal fusion object detection method, combined with thermal imaging, ultraviolet light, low-light channel imaging and image fusion models, the imaging problem of night vision instruments in dark and weak environments is solved, real color display and object detection are realized, and the imaging quality and environmental adaptability of night vision instruments are improved.

CN120495824APending Publication Date: 2025-08-15ARMY ENG UNIV OF PLA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510631068.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing night vision instruments have poor imaging quality in dark environments and cannot effectively transmit target color information. They have problems with parallax and insufficient algorithm timeliness, and cannot effectively image in light-free environments.

Method used

The multimodal fusion object detection method is adopted to obtain the image to be detected for state judgment, and thermal image, ultraviolet light, and low light channel imaging is selectively performed, and image fusion is used to fusion, combining polarized light sources and near-infrared light sources for adjustments to realize true color display and object detection.

Benefits of technology

It realizes clear imaging under different lighting and environmental conditions, improves the versatility and practicality of night vision devices, enhances target recognition capabilities and real-time performance, and improves imaging quality and environmental adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495824A_ABST
    Figure CN120495824A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal fusion target detection method and system suitable for a dark environment, and belongs to the technical field of image processing. The method comprises the following steps: acquiring a to-be-detected image; performing state judgment on the to-be-detected image to obtain a state judgment result; respectively performing thermal image channel imaging, ultraviolet light channel imaging and low-light channel imaging on the to-be-detected image to obtain a thermal image channel image, an ultraviolet light channel image and a low-light channel image; selecting one or more images from the thermal image channel image, the ultraviolet light channel image and the low-light channel image according to the state judgment result and a preset channel priority; performing channel imaging screening on the selected image to obtain a screened image; inputting the screened image into a pre-constructed image fusion model for image fusion to obtain a fused image; and performing target detection on the fused image to obtain a target detection result. According to the invention, multiple spectrums can be selectively fused, true color display can be realized, and target color information can be effectively transmitted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a multimodal fusion target detection method and system suitable for dim environments. Background Art

[0002] Infrared thermal imaging technology receives infrared radiation emitted by objects and converts it into visible thermal images, but the images are rough and blurry, with unclear details, and can be easily deceived by infrared camouflage nets. The image clarity is reduced when encountering a medium that can reflect heat.

[0003] Low-light-level night vision technology uses a low-light image intensifier to amplify weak light sources, forming images observable to the human eye. However, low-light-level night vision technology is easily affected by the surrounding environment, resulting in low image contrast and unclear layers. Furthermore, it relies on weak nighttime light sources and cannot form images in the absence of light.

[0004] As for visible light and infrared fusion night vision devices, most currently available are single-band visible light or infrared devices, as well as fusion devices that combine visible light and infrared. Fusion night vision devices use dual lenses to capture separate images and use algorithms to fuse the images. However, these devices suffer from issues such as grayscale or false color display, parallax between dual-lens images, artifacts in the fused image, and algorithmic inefficiencies that affect real-time display.

[0005] Fusion night vision systems combine low-light and infrared imaging technologies, achieving image fusion through optical components or digital image processing. Besides being bulky, heavy, and expensive, infrared cannot form temperature-differential images or penetrate transparent objects in dark, shadowy environments with small temperature differences. Dark environments also lack sufficient fill light for the low-light imaging tubes, rendering the fusion night vision system ineffective. Using active infrared for fill light would be detected by enemy infrared detection systems, exposing the user and potentially even causing the operation to fail. Furthermore, the image fusion process suffers from parallax errors and algorithmic timeliness issues.

[0006] The individual soldier's full-color night vision system uses an ultra-low-light digital image sensor to capture red, green, and blue light sources separately, and then synthesizes a full-color image through software. However, it has problems such as delays caused by software calculations and image processing, image freezes and delays caused by low frame rates, and inability to operate in dark environments.

[0007] External low-light-level thermal fusion night vision devices attach a thermal imager as a module to the device, enabling image overlay. However, the thermal image is too bright, impacting the device's performance. Furthermore, the fixed-focus system is unsuitable for close-focus observation, the image is monochrome, and it cannot display color text or images.

[0008] Currently, most fusion night vision devices use grayscale or pseudo-color displays, which do not conform to the habits of the human eye and cannot effectively transmit the target color information. There is parallax in the dual-lens imaging fusion, and the algorithm is not timely enough, and true color display cannot be achieved, affecting the realism and observation effect of the image. Low-light-level night vision devices cannot form images in lightless environments, and infrared thermal imagers have poor imaging effects in specific environments. They are not universal and practical, and the imaging quality, environmental adaptability, real-timeness and maintainability cannot be guaranteed. Summary of the Invention

[0009] The purpose of the present invention is to overcome the shortcomings of the existing technology and provide a multimodal fusion target detection method and system suitable for dark environments. The method and system can selectively fuse multiple spectra, realize true color display, effectively transmit target color information, improve the realism and observation effect of the image, have versatility and practicality, and can significantly improve imaging quality, environmental adaptability, real-time performance and maintainability.

[0010] To achieve the above object, the present invention is implemented by adopting the following technical solutions:

[0011] On the one hand, the present invention provides a multimodal fusion target detection method suitable for dim environments, comprising:

[0012] Obtain the image to be detected;

[0013] Performing a status judgment on the image to be detected to obtain a status judgment result;

[0014] Performing thermal imaging channel imaging, ultraviolet light channel imaging, and low-light channel imaging on the image to be detected, respectively, to obtain a thermal imaging channel image, an ultraviolet light channel image, and a low-light channel image;

[0015] Selecting one or more images from the thermal imaging channel image, the ultraviolet light channel image, and the low-light channel image according to the state judgment result and the preset channel priority;

[0016] Perform channel imaging screening on the selected image to obtain a screening image;

[0017] Inputting the screened image into a pre-built image fusion model for image fusion to obtain a fused image; the image fusion model includes an encoder, an attention module, and a decoder connected in sequence;

[0018] Performing target detection on the fused image to obtain a target detection result.

[0019] Optionally, the status judgment result is expressed as:

[0020] ;

[0021] in, Indicates the status judgment result; 、 、 、 、 They represent the dynamic weights of the light intensity parameter, polarization degree parameter, thermal radiation difference parameter, ultraviolet characteristic intensity parameter, and low-light signal-to-noise ratio parameter of the image to be detected respectively; 、 、 、 、 They represent the light intensity parameter, polarization degree parameter, thermal radiation difference parameter, ultraviolet characteristic intensity parameter, and low-light signal-to-noise ratio parameter of the image to be detected respectively; represents the feature extraction function; Represents the argmax activation function.

[0022] Optionally, the state judgment result includes a strong light environment, a dim light environment, a dark light environment, a no light environment, an environment blocked by a transparent object, a dark shadow environment, and an environment with similar temperature.

[0023] Optionally, also include:

[0024] When there is interference in the image to be detected, performing polarized light source imaging on the image to be detected to obtain a polarized light image;

[0025] One or more images are selected from the polarized light image, the thermal imaging channel image, the ultraviolet light channel image, and the low-light channel image.

[0026] Optionally, also include:

[0027] Use the acquired polarized light source to detect suspicious targets and determine whether the suspicious targets are enemy infrared detectors or heat sources;

[0028] If the suspicious target is not an enemy infrared detector or heat source, a feedback signal is sent to the near-infrared light source, and the visible light module image in the thermal imaging channel image is adjusted using the near-infrared light source until the intensity of the near-infrared light source is greater than the brightness of the visible light module image, thereby obtaining an adjusted visible light module image; and a thermal imaging channel adjusted image is obtained based on the adjusted visible light module image and the acquired far-infrared light module image;

[0029] If the suspected target is an enemy infrared detector or heat source, no feedback signal is sent to the near-infrared light source.

[0030] Optionally, the processing steps of the image fusion model include:

[0031] In the encoder, low-frequency feature extraction and high-frequency feature extraction are sequentially performed on the screened image to obtain mixed features;

[0032] In the attention module, importance weights are assigned to the mixed features according to the channel importance weights to obtain assigned features;

[0033] In the decoder, the assigned features are decoded and reconstructed to obtain a fused image.

[0034] Optionally, the channel importance weight is expressed as:

[0035] ;

[0036] in, represents the importance weight of the i-th channel; represents the signal-to-noise ratio of the i-th channel; Indicates the matching degree of the i-th channel environment; 、 Represents the adjustment coefficient.

[0037] Optionally, before performing target detection on the fused image, the method further includes:

[0038] The fused image is subjected to brightness adjustment, contrast adjustment, noise filtering, and edge detection to obtain a processed image.

[0039] In a second aspect, the present invention provides a multimodal fusion target detection system suitable for dim environments, comprising:

[0040] The image acquisition module is used to: acquire the image to be detected;

[0041] A state judgment module is used to: perform state judgment on the image to be detected and obtain a state judgment result;

[0042] The channel imaging module is used to perform thermal imaging channel imaging, ultraviolet light channel imaging, and low-light channel imaging on the image to be detected, thereby obtaining a thermal imaging channel image, an ultraviolet light channel image, and a low-light channel image;

[0043] An image selection module is used to select one or more images from the thermal imaging channel image, the ultraviolet light channel image, and the low-light channel image according to the state judgment result and the preset channel priority;

[0044] An image screening module is used to: perform channel imaging screening on the selected image to obtain a screening image;

[0045] An image fusion module is configured to input the screened image into a pre-built image fusion model for image fusion to obtain a fused image; the image fusion model includes an encoder, an attention module, and a decoder connected in sequence;

[0046] The target detection module is used to perform target detection on the fused image to obtain a target detection result.

[0047] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the multimodal fusion target detection method for dim environments as described in the first aspect.

[0048] Compared with the prior art, the present invention has the following beneficial effects:

[0049] 1. The present invention selectively fuses multiple spectra based on the state judgment results of the image to be detected, which can achieve true color display, effectively transmit target color information, improve the realism and observation effect of the image, and combine the imaging characteristics of multiple spectra to achieve clear imaging under different lighting and environmental conditions, thereby improving the versatility and practicality of night vision devices, and significantly improving the imaging quality, environmental adaptability, real-time performance and maintainability of night vision devices.

[0050] 2. The present invention utilizes special spectra such as ultraviolet light and polarized light to enhance the ability to identify specific targets and improve the detection and recognition performance of night vision devices. It introduces innovative technologies such as state judgment algorithm, channel priority, multi-spectral fusion network and channel importance weight to enhance the core performance of the system from the algorithm and mechanism level. Compared with traditional methods, it has significantly improved the target detection accuracy, real-time performance and anti-interference ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 FIG2 is a flow chart of a multimodal fusion target detection method applicable to dim environments in an embodiment of the present invention;

[0052] Figure 2 FIG2 is a schematic structural diagram of a multimodal fusion target detection system suitable for dim environments according to an embodiment of the present invention;

[0053] Figure 3 The figure shows a comparison diagram of brightness and polarization degree of an image to be detected in one embodiment of the present invention within a natural day;

[0054] Figure 4 FIG. 4 is a schematic structural diagram of an image fusion model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0055] The technical solution of the present invention is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.

[0056] The term "and / or" simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. Additionally, the character " / " generally indicates an "or" relationship between the related objects.

[0057] Example 1

[0058] like Figure 1 As shown, this embodiment introduces a multimodal fusion target detection method suitable for dim environments, which specifically includes the following steps:

[0059] Step 1: Use the incoming light detector to perform perception analysis and status judgment, and use polarized light to illuminate suspicious targets to find enemy infrared detectors. Specifically:

[0060] The light input detector analyzes and integrates light in various spectral ranges to achieve accurate perception of the environment.

[0061] The light input judger is used to judge the state of the acquired image to be detected and obtain the judgment result, wherein the judgment result includes strong light environment, low light environment, dark light environment, no light environment, environment blocked by transparent objects, dark shadow environment, and environment with similar temperature.

[0062] This embodiment introduces a deep learning classification model based on multi-parameter fusion, namely a lightweight convolutional neural network (CNN) model. The loss function of this CNN model is a cross-entropy loss combined with prior knowledge of the environment (for example, the low-light SNR has a higher weight in a dark environment).

[0063] The input of the CNN model is the characteristic parameters of multispectral data: light intensity (L), polarization (P), thermal radiation difference ( ), ultraviolet characteristic intensity (UV), low-light signal-to-noise ratio (SNR), the output is the probability distribution of the environment type, and the state judgment result is expressed as:

[0064] ;

[0065] in, Indicates the status judgment result; 、 、 、 、 They represent the dynamic weights of the light intensity parameter, polarization degree parameter, thermal radiation difference parameter, ultraviolet characteristic intensity parameter, and low-light signal-to-noise ratio parameter of the image to be detected respectively; 、 、 、 、 They represent the light intensity parameter, polarization degree parameter, thermal radiation difference parameter, ultraviolet characteristic intensity parameter, and low-light signal-to-noise ratio parameter of the image to be detected respectively; represents the feature extraction function; Represents the argmax activation function.

[0066] Perform thermal imaging channel imaging, ultraviolet light channel imaging, and low-light channel imaging on the image to be detected, respectively, to obtain a thermal imaging channel image, an ultraviolet light channel image, and a low-light channel image;

[0067] At the same time, the light input detector can use information from polarized light or infrared spectrum segments to locate possible enemy infrared detectors or heat sources. By analyzing thermal imaging data, the system can distinguish between natural heat sources and artificial heat sources, such as enemy equipment or personnel.

[0068] Use the acquired polarized light source or infrared spectrum information to detect suspicious targets and determine whether the suspicious targets are enemy infrared detectors or heat sources;

[0069] If the suspicious target is not an enemy infrared detector or heat source, a feedback signal is sent to the near-infrared light source, and the thermal imaging channel image is adjusted using the near-infrared light source to obtain a thermal imaging channel adjusted image;

[0070] If the suspected target is an enemy infrared detector or heat source, no feedback signal is sent to the near-infrared light source;

[0071] The thermal imaging channel image includes a visible light module image and a far-infrared light module image. In this embodiment, a near-infrared light source is used to adjust the visible light module image in the thermal imaging channel image until the intensity of the near-infrared light source is greater than the brightness of the visible light module image, thereby obtaining a visible light module adjusted image.

[0072] That is, since thermal imagers cannot observe targets blocked by transparent objects or targets whose temperature is similar to the surrounding dark environment (determined by the difference between visible light VIS and near-infrared light NIR in the thermal imager, that is, the brightness / intensity of visible light VIS is greater than that of near-infrared light NIR), it is necessary to use near-infrared light for fill light, so that the near-infrared light NIR is greater than the visible light VIS. At night or in low light conditions, the performance of the visible light module is enhanced, and near-infrared light is used to illuminate the target for fill light, thereby improving image quality and target recognition rate.

[0073] Step 2: Based on the judgment result and the light, temperature difference and other conditions, select one or more of the thermal imaging channel, ultraviolet light channel, and low light channel. Specifically:

[0074] The light entry discriminator can intelligently fuse multiple spectral information, including thermal imaging, ultraviolet light, and low-light, based on the battlefield situation. This fusion aims to compensate for the shortcomings of a single sensor, such as the limitations of thermal imaging in identifying details and the sensitivity of low-light sensors in bright light environments. Through algorithm optimization, the light entry discriminator can create a more comprehensive and clearer scene image. The thermal imaging channel is used to detect temperature differences to identify concealed enemy personnel and equipment; the low-light channel captures terrain details in weak light; and the ultraviolet channel is used to identify possible chemical markers or special substances.

[0075] Specifically, based on the judgment result and preset channel priorities, one or more images are selected from the thermal imaging channel, the ultraviolet channel, and the low-light channel. Channel priorities are defined based on the environment type. For example, in a low-light environment, the low-light channel image has a higher priority than the thermal imaging channel image.

[0076] In low-light environments, the light detector can detect the light level entering the objective lens and determine whether to activate additional light sources, such as polarized light sources. Polarized light can penetrate certain types of smoke and dust, improving the ability to identify distant objects, especially in environments with interference. Through initial detection, the system can determine whether it is necessary to activate a polarized light source to enhance target detection and recognition capabilities.

[0077] That is, when there is interference in the perception environment of the target to be detected, the image to be detected is imaged using a polarized light source to obtain a polarized light image; and one or more images are selected from the polarized light image and the thermal imaging channel image, the ultraviolet light channel image, and the low-light channel image.

[0078] Step 3: The image of the selected channel and the result of the polarization filter are displayed on the split screen. The light output decision device makes a decision and filters the channel with poor imaging effect. Specifically:

[0079] The selected image is subjected to channel imaging screening to obtain a screened image. The light output judge receives original signals from the thermal imaging split screen, ultraviolet light split screen, low light split screen and polarized light split screen. These signals carry environmental information in different spectral ranges. The judge's task is to analyze these diversified signals, evaluate their relative importance and applicability, and then decide how to optimally fuse these signals to generate the clearest and most complete target image. This fusion capability allows the system to provide the best visual effects in complex and changing environments.

[0080] Step 4: Based on the polarized light detection results and the thermal imager's imaging, the outgoing light determinator and incoming light determinator feedback signals. If the enemy does not have an infrared detector, the active near-infrared light source is activated. Specifically:

[0081] Referring to step one, if the suspicious target is not an enemy infrared detector or heat source, the outgoing light detector sends a feedback signal to the near-infrared light source, and the incoming light detector uses the near-infrared light source to adjust the visible light module image in the thermal imaging channel image until the intensity of the near-infrared light source is greater than the brightness of the visible light module image, thereby obtaining a visible light module adjusted image.

[0082] Step 5: Input the filtered image into the image fusion model for image fusion to obtain a fused image. Specifically:

[0083] This example proposes a channel confidence assessment method based on fuzzy logic to optimize the selection and fusion ratio of multispectral images. It also uses a Multi-Scale Fusion Network (MSF-Net) that combines wavelet transform with an attention mechanism to preserve details and suppress noise.

[0084] The image fusion model includes an encoder, an attention module, and a decoder connected in sequence; the loss function of the image fusion model is for:

[0085] ;

[0086] in, To optimize the structural similarity loss function, Preserve loss function for edge enhancement; 、 is a hyperparameter.

[0087] like Figure 4 As shown, in the encoder, the low-frequency (contour) features and high-frequency (detail) features of each channel of the filtered image are extracted in turn.

[0088] In the attention module, the importance weights of the mixed features are distributed according to the channel importance weights to obtain the distributed features. For example, the weight of the chemical labeling area of the ultraviolet channel is increased; the weight of the low-light channel in the low-light environment is 0.7, and the weight of the thermal imaging channel is 0.3; the channel importance weights are expressed as:

[0089] ;

[0090] in, represents the importance weight of the i-th channel; represents the signal-to-noise ratio of the i-th channel; Indicates the matching degree of the i-th channel environment; 、 represents the adjustment coefficient;

[0091] In the decoder, the assigned features are decoded and reconstructed to obtain a fused image that retains the multispectral complementary information.

[0092] Step 6: The light output decision device inputs the fused image into the image processing circuit to obtain a processed image, presents the processed image on the fused display screen, and uses the detection algorithm to perform target detection on the processed image to obtain the target detection result, which can be observed through the eyepiece group.

[0093] Example 2

[0094] Based on Example 1, this example introduces an experimental example of a multimodal fusion target detection method suitable for dim environments:

[0095] Detection method flow for strong light environment, low light environment, dark light environment, no light environment, environment with transparent objects blocking, dark shadow environment, and environment with similar target and ambient temperature:

[0096] In strong light conditions:

[0097] Step 1: The light detector behind the objective lens assembly detects high ambient light intensity and automatically turns off or reduces the power of the polarized light source to avoid overexposure or reflection interference;

[0098] Step 2: Due to sufficient light, only the thermal imaging channel and the ultraviolet light channel may be enabled to confirm potential heat sources or non-visible light activities;

[0099] Step 3: Split the screen to display the thermal image and UV channel images. The light input judge checks the image quality to confirm that no other channels are needed to supplement.

[0100] Step 4: Since enemy infrared detectors are difficult to detect under strong light conditions, the active near-infrared light source remains in standby mode;

[0101] Step 5: The clear image of the low-light channel is directly sent to the image processing circuit and presented to the observer through the fusion display screen.

[0102] In low-light and dim environments:

[0103] Step 1: The light input detector detects that the ambient light is weak, turns on the polarized light source, and uses the thermal imaging channel to detect the temperature difference;

[0104] Step 2: The thermal imaging channel and the low-light channel are integrated to compensate for insufficient light, and the ultraviolet channel is used as an auxiliary to detect specific substances;

[0105] Step 3: The light output decision device evaluates the imaging quality of each channel and may select the thermal imaging channel as the primary channel and the low-light channel as the secondary channel for fusion display;

[0106] Step 4: If no enemy infrared detectors are detected, activate the near-infrared light source to enhance illumination, but use caution to avoid exposing your position.

[0107] Step 5: The image processing circuit combines and optimizes the images from the thermal and low-light channels and displays the final result on a fusion display.

[0108] In dark and shadowy environments:

[0109] Step 1: In complete darkness or shadow, the light detector determines that natural light is unavailable, and the polarized light source is turned off. Since low-light conditions are unavailable, the low-light channel is turned off, and the thermal imaging channel and active near-infrared light source are used instead.

[0110] Step 2: Rely primarily on the thermal imaging channel to obtain the target outline, while activating the near-infrared light source to illuminate the environment. The low-light channel may not provide effective information.

[0111] Step 3: The images generated by the thermal imaging channel and the near-infrared light source are filtered by the light output decision device to ensure there is no overexposure or interference;

[0112] Step 4: The near-infrared light source continues to work to provide necessary illumination, while the thermal imaging channel continues to monitor the temperature difference;

[0113] Step 5: The image processing circuit combines the thermal image and the image under near-infrared illumination, and the fusion display shows the final high-contrast image.

[0114] When the target and ambient temperatures are almost similar:

[0115] Step 1: If the light input detector detects that the temperature difference is not obvious, it may increase the use of polarized light sources and use the ultraviolet light channel to detect special features;

[0116] Step 2: Selectively fuse all available channels. The thermal imaging channel alone may not provide sufficient discrimination.

[0117] Step 3: The light output detector may require more complex algorithms to filter and fuse information from multiple channels to enhance target recognition;

[0118] Step 4: If the temperature difference is not sufficient to identify the target, near-infrared light sources or other auxiliary means may be used to enhance the contrast;

[0119] Step 5: The image processing circuit deeply integrates all channel data, enhances the distinction between the target and the background through algorithms, and finally presents it on the fusion display.

[0120] In an environment with transparent objects blocking the view:

[0121] Step 1: The light detector detects a transparent obstruction. Polarized light sources may help penetrate certain types of transparent materials, such as water or mist.

[0122] Step 2: Use polarized light technology to try to penetrate the obstruction, while thermal imaging and ultraviolet light channels help identify activities behind the obstruction;

[0123] Step 3: The light output decision device selects the clearest image based on the effect of polarized light, which may require combining data from thermal imaging and ultraviolet light channels;

[0124] Step 4: Depending on the nature of the obstruction, you may need to adjust the wavelength or angle of the polarized light to achieve optimal penetration.

[0125] Step 5: The final image is processed by the image processing circuit and the target image after penetrating the obstruction is presented on the fusion display screen.

[0126] In a specific embodiment, during a night raid mission, a special operations team was dispatched to an enemy-controlled mountainous area with the goal of detecting and destroying the enemy's military base and its forward defenses deep in the valley. The terrain in the area was complex, including dense forests, narrow canyons, and cliffs partially made of rock. The light was extremely weak at night, and visibility was extremely low due to the light fog. The enemy also deployed advanced counter-reconnaissance measures, including camouflage, smoke screens, and infrared detectors.

[0127] Step 1: The special forces team's multimodal fusion target detection system begins operating. The light detector behind the objective lens analyzes the environment and detects extremely weak light and light fog. Simultaneously, the system emits polarized light, attempting to penetrate the fog and locate the presence of enemy infrared detectors.

[0128] Step 2: Based on the judgment result, the system automatically integrates the thermal imaging channel, the low-light channel, and the ultraviolet light channel; the thermal imaging channel is used to detect temperature differences to identify concealed enemy personnel and equipment; the low-light channel captures terrain details in weak light; and the ultraviolet light channel is used to identify possible chemical markers or special substances.

[0129] Step 3: The light output decision device analyzes the images from the three channels, removes the low-light channel image that is severely affected by fog, and fuses the images from the thermal and ultraviolet channels to form a clear battlefield situation map.

[0130] Step 4: The system detects that the enemy has deployed infrared detectors and decides not to activate the active near-infrared light source to avoid exposing its position. However, for specific areas, such as suspected enemy outposts or key facilities, the system will appropriately control the polarized light source to minimize exposure risk while providing necessary illumination.

[0131] Step 5: The image processing circuit integrates the filtered images, performs target detection and enhances the target outline through algorithms, reduces background noise, and ultimately presents a clear enemy base layout and defensive fortification locations on the fused display. Special forces members observe in real time through the eyepiece set and quickly plan their routes and attack points.

[0132] Example 3

[0133] Based on Example 1, this example introduces a multimodal fusion target detection system suitable for dim environments, including:

[0134] The image acquisition module is used to: acquire the image to be detected;

[0135] A state judgment module is used to: perform state judgment on the image to be detected and obtain a state judgment result;

[0136] The channel imaging module is used to perform thermal imaging channel imaging, ultraviolet light channel imaging, and low-light channel imaging on the image to be detected, thereby obtaining a thermal imaging channel image, an ultraviolet light channel image, and a low-light channel image;

[0137] An image selection module is used to select one or more images from the thermal imaging channel image, the ultraviolet light channel image, and the low-light channel image according to the state judgment result and the preset channel priority;

[0138] An image screening module is used to: perform channel imaging screening on the selected image to obtain a screening image;

[0139] An image fusion module is configured to input the screened image into a pre-built image fusion model for image fusion to obtain a fused image; the image fusion model includes an encoder, an attention module, and a decoder connected in sequence;

[0140] The target detection module is used to perform target detection on the fused image to obtain a target detection result.

[0141] Among them, such as Figure 2 As shown, each module realizes target detection through an objective lens group, an incoming light discriminator, a near-infrared light source, an enhanced thermal imager, an ultraviolet imager, a low-light-level night vision device, a polarized light source, an outgoing light discriminator, an image processing circuit, a fusion display screen, and an eyepiece group.

[0142] The image to be detected is acquired through the objective lens group.

[0143] Near-infrared light sources are often used for supplemental lighting, especially in low-light conditions, providing additional illumination without interfering with human visual perception. Near-infrared light (NIR) has a wavelength range of roughly 700nm to 1400nm. This light source is very effective for enhancing the contrast of certain materials or penetrating minor obstacles such as mist or smoke. In enhanced thermal imagers, near-infrared light sources can improve image quality, especially at night or in low visibility conditions.

[0144] By integrating with the visible light module in an enhanced thermal imager, it can provide richer and more detailed information. For example, in a battlefield environment, the visible light module can help warfighters see the entire battlefield, the far-infrared module can reveal hidden heat sources (such as the human body), and the near-infrared light source can provide additional contrast in low-light environments, helping to distinguish different objects or features.

[0145] The enhanced thermal imager includes an enhanced thermal imaging unit, a visible light module, a far-infrared light module, and a thermal imaging split screen. By integrating the visible light module and the far-infrared light module, it improves the detection and imaging capabilities. Specifically:

[0146] Visible light imaging can reveal basic sensory information such as the outline and appearance of a target object, especially in good lighting conditions, where it can accurately delineate the target. However, under extreme weather conditions, such as those caused by fog or smoke, the amount of image information available is extremely limited.

[0147] In contrast, although infrared imaging can overcome the serious obstacles caused by extremely harsh weather conditions and more accurately capture the thermal radiation of target objects, showing excellent performance in countering interference and dark night environments, infrared imaging has low resolution and blurred target texture details, resulting in high missed and false alarm rates in target detection. Therefore, target detection based on combining the advantages of both can not only extract and fuse complementary feature information in the original image to generate a fused image with richer single information and better scene expression capabilities, but also further improve the accuracy of high-level target detection and is widely used in complex scenarios such as combat.

[0148] The visible light module is a traditional optical component in a thermal imager. It uses principles similar to those of ordinary cameras to capture visual images of a scene, including lenses, sensors (such as CCD or CMOS), and other electronic components. The role of the visible light module is to provide clear visual images so that users can identify specific objects and details in the scene.

[0149] The far-infrared light module is the core of the thermal imager, which can detect and measure the infrared radiation (heat energy) emitted by objects. The far-infrared light module can generate a temperature distribution map, that is, a thermal image, which is used to show the temperature differences between different objects or different areas of the same object.

[0150] An ultraviolet imager is a device used to capture images in the ultraviolet (UV) band. Its main components include an ultraviolet imaging unit, a UV photoelectric sensor, and an ultraviolet split screen.

[0151] Ultraviolet light typically has wavelengths between 100nm and 400nm, making it invisible to the human eye. Ultraviolet imagers use specialized sensors to receive UV radiation and convert it into visible images. Sunlight ultraviolet light with wavelengths between 190nm and 285nm is almost completely absorbed by Earth's ozone layer as it passes through the atmosphere. Even below the ozone layer, scattering from other atmospheric components and surface ozone also absorb it, resulting in near-zero background radiation at Earth's surface. This creates a natural "dark chamber" near Earth's surface—a place where naturally occurring ultraviolet light in this wavelength band is virtually undetectable. Solar-blind ultraviolet light is typically detected in three key scenarios: natural meteorological anomalies, such as strong lightning; unnatural danger signals, such as gunfire, gunpowder explosions, fires, and corona generated by leakage from high-voltage power lines; or artificially generated solar-blind ultraviolet light sources. This means that detecting a solar-blind ultraviolet signal in this dark chamber indicates the occurrence of a specific event.

[0152] In the solar-blind ultraviolet spectrum region, the background noise is extremely low. Compared with infrared detection technology and visible light detection technology, solar-blind ultraviolet imaging technology has the following advantages:

[0153] 1) All-day, all-weather: Daytime detection is completely unaffected by strong sunlight, environmental changes (fog, haze, day and night changes), and other high-temperature interference sources;

[0154] 2) Intuitive, no image processing required: Sun-blind UV images have no background interference, and the imaging is intuitive, without the need for complex image processing. For example, when a missile strikes, no complex image processing is required at all, and "discovery is targeting";

[0155] 3) High precision: Strong anti-interference ability, suitable for various complex environments (nighttime, bad weather, electromagnetic interference, etc.), and can achieve centimeter-level high-precision positioning;

[0156] 4) High sensitivity: The response time including imaging reaches 10ms;

[0157] 5) Wide monitoring range: Real-time monitoring can be carried out from several kilometers away, and a large area can be covered by means of PTZ or helicopter observation;

[0158] 6) Simple structure: No need for cryogenic cooling or scanning, the detection imaging system is small in size, light in weight and has low power consumption.

[0159] In the civilian sector, it is being widely used in areas such as aircraft blind landing through fog, ship pilotage through fog, forest fire warnings, power grid security monitoring, maritime search and rescue, satellite navigation, and criminal investigation. Power systems and high-speed rail systems use solar-blind ultraviolet imaging technology to achieve highly sensitive corona and arc detection, identifying early local defects and issuing early warnings. It can also monitor equipment operation to prevent faulty operation and major accidents. Leveraging the principle that forest fires emit solar-blind ultraviolet signals, it enables remote, real-time monitoring of forest fires while shielding against interference from other high-temperature objects, addressing the high false alarm rate of existing infrared and visible fire alarms. In addition to long-range imaging, ultraviolet imaging can also be used for close-range criminal investigation photography, primarily to search for latent fingerprints, footprints, hidden bloodstains, and other physical evidence. UV imaging can capture latent fingerprints, weak contrast, or hidden fingerprints that are obscured by background interference. It is suitable for photographing fingerprints left on surfaces such as glass, plastic, enamel, paint, photos, and adhesive tape. Glass appears black due to its strong absorption of 254nm short-wave ultraviolet rays, while sweat fingerprints appear white due to their 50% reflectivity. The contrast between black and white is very large, so ultraviolet imaging can show subtle features very clearly.

[0160] In the military, ultraviolet imagers can be used for missile warning, space-based early warning, and the detection of gunfire and gunpowder explosions. The ultraviolet spectrum lies just outside the visible spectrum and carries a unique and critical piece of information: signs of corona discharge. Corona discharge is essentially a subtle charge release process triggered by a localized voltage exceeding the breakdown strength threshold of the medium (usually air). This phenomenon is particularly pronounced in the military, where static electricity accumulated by various factors (such as friction) from military infrastructure to large-scale combat equipment, and even soldiers during operations, can trigger corona discharge, flashover, or arcing.

[0161] During a corona discharge, electrons in air molecules undergo a rapid cycle of energy absorption and release. Each time these electrons release their accumulated energy, they instantly stimulate the emission of ultraviolet light. This characteristic is invaluable in military surveillance and security. UV imaging technology can keenly detect these subtle UV signals, even the most subtle corona discharge phenomena. This not only provides a new dimension of battlefield situational awareness but also provides early warning of potential electrical failures, preventing unplanned downtime of military equipment, ensuring operational effectiveness while protecting personnel safety.

[0162] The low-light-level night vision device includes a low-light-level night vision unit, an image intensifier tube (IIT) and a low-light-level split screen.

[0163] Low-light-level night vision utilizes the low light levels in the environment to produce clear images even under starlight. It provides visible images by enhancing existing light, including moonlight, starlight, or city background light.

[0164] In conditions where there is almost no visible light, low-light night vision devices can provide sufficient brightness and contrast to enable users to see objects in the dark.

[0165] Polarized light source is a device that produces polarized light. Polarized light refers to light whose light wave vibration direction is confined to a specific plane. It includes polarizing filters and polarized light splitters, the latter of which can convert natural light into polarized light.

[0166] Polarimetric imaging technology holds unique potential in battlefield environments, demonstrating its potential for countering concealment, enhancing anti-interference capabilities, extending target detection range, and accurately diagnosing target properties. Since 1985, Timothy J. Rogne and his colleagues have been exploring how to use polarimetric detection to optimize infrared target recognition in low signal-to-noise ratio environments and complex backgrounds. Their research has demonstrated its effectiveness through a series of exhaustive experiments. Their research targets include man-made structures such as vehicles, aircraft, and surfaces of varying materials, including asphalt, metal, and concrete. Furthermore, natural environments such as trees, grass, water, and clouds provide a diverse range of background environments.

[0167] When the target and the surrounding environment have similar radiation intensities or are obscured, traditional imaging methods often fall short and struggle to effectively distinguish the target from the background. In contrast, polarization imaging technology can accurately highlight the target from a complex background. Furthermore, polarization imaging technology can overcome a thorny problem in infrared detection, known as "thermal contrast decay." This phenomenon manifests as a sudden drop in the temperature contrast between the target and the background, almost to the point of being indistinguishable, at a specific moment. In the brightness image, there are two time periods where the target contrast is nearly zero. However, in the polarization image, the target contrast is maintained at a significantly higher level, demonstrating the excellent performance of polarization imaging technology in all-weather target recognition.

[0168] like Figure 3 As shown in the figure, this graph records 24 hours of observation data for the same scene and plots a trend line for the contrast of a selected target over time. In this data set, Brightness represents the curve of the target contrast over time based on the brightness image, while Polarization reflects the changing trend of the same indicator in the linear polarization image.

[0169] The light detector is the core component of the system, responsible for processing and integrating information from different sensors to optimize vision at night or in low-light conditions. The light detector works by analyzing and integrating light from various spectral ranges to achieve accurate perception of the environment. The following is a detailed description of the light detector's functions:

[0170] ①Activate the polarized light source for initial detection

[0171] In low-light environments, a light detector detects the light level entering the objective lens and determines whether to activate additional light sources, such as polarized light. Polarized light can penetrate certain types of smoke and dust, improving the ability to identify distant objects, especially in environments with interference. Through initial detection, the system can determine whether to activate a polarized light source to enhance target detection and recognition capabilities.

[0172] ②Find enemy infrared detectors

[0173] The light detector uses information from polarized light or infrared spectrum segments to locate possible enemy infrared detectors or heat sources. By analyzing thermal imaging data, the system can distinguish between natural and artificial heat sources, such as enemy equipment or personnel.

[0174] ③Integration of multiple spectral information

[0175] The light input discriminator intelligently integrates multiple spectral information, including thermal imaging, ultraviolet light, and low-light, based on the battlefield situation. This fusion aims to address the shortcomings of a single sensor, such as thermal imaging's limitations in identifying detail and the sensitivity of low-light sensors in bright sunlight. Through algorithmic optimization, the discriminator creates a more comprehensive and clearer image of the scene.

[0176] ④Dispatch near-infrared light supplement

[0177] Because thermal imagers cannot observe targets obscured by transparent objects or targets with temperature differences similar to those in the surrounding night environment (this is determined by the difference between visible light (VIS) and near-infrared light (NIR) in a thermal imager, i.e., the brightness / intensity of visible light (VIS) is greater than that of near-infrared light (NIR),), near-infrared light is required for fill illumination, making the near-infrared light greater than the visible light (VIS). At night or in low light conditions, the performance of the visible light module is enhanced, and near-infrared light is used to illuminate the target for fill illumination, thereby improving image quality and target recognition rate.

[0178] Ultimately, the goal of the light detector is to achieve perspective detection of targets by analyzing and fusing information from various light sources. This means that even in complex battlefield environments, the system can penetrate smoke, dust, and other obstacles to accurately locate and identify targets, providing the operator with clear visual information.

[0179] The optical decision maker is also a core component of the system, and its functions include:

[0180] ①Multimodal signal fusion

[0181] The optical discriminator receives raw signals from the thermal imaging screen, the ultraviolet light screen, the low-light light screen, and the polarized light screen. These signals carry environmental information from different spectral ranges. The discriminator's task is to analyze these diverse signals, assess their relative importance and applicability, and then determine how to optimally fuse them to produce the clearest and most complete target image. This fusion capability allows the system to provide optimal visual effects in complex and changing environments.

[0182] ②Control near-infrared light source

[0183] The light output discriminator is also responsible for sending feedback signals to the near-infrared light source. This function enables the system to dynamically adjust the light source's output to adapt to changing environmental conditions. For example, in low-light environments, the discriminator may instruct the near-infrared light source to increase its brightness to supplement the visible light module's insufficient brightness; while in high-light conditions, it may reduce the light source's intensity to avoid overexposure or interference with the normal operation of other sensors, ensuring system stability and efficiency under various lighting conditions.

[0184] ③Interface with image processing circuit

[0185] The output of the discriminator is directly connected to the image processing circuit, which not only transmits the fused signal, but may also contain additional instructions or parameters to guide the subsequent image processing process, including processing instructions such as brightness adjustment, contrast enhancement, noise filtering, edge detection, etc., aiming to optimize the quality of the final image and make it more suitable for human vision and further computer vision analysis.

[0186] Features of the Fusion Display include:

[0187] ①Multi-source data fusion display

[0188] The fused display screen displays the processed image signals received from the image processing circuitry. These signals originate from one or more of thermal imaging, ultraviolet light, low-light, polarized light, and possibly with the assistance of near-infrared light sources. The display screen's task is to integrate data from these various sources into a unified visual image, allowing the operator to gain a comprehensive understanding of the environment on a single screen. The fused display screen can detect objects in the visual image using technologies such as YOLOv5, YOLOv8, and YOLOv11.

[0189] ② Augmented reality function

[0190] The fused display screen has augmented reality (AR) features that not only show the actual captured image, but can also overlay additional graphics, text, or other information, such as target identification, distance measurements, temperature readings, etc., which help improve the operator's understanding and reaction speed.

[0191] ③User interface and control

[0192] Display screens are typically equipped with a user interface that allows the operator to adjust settings, such as changing the image contrast and brightness, selecting a different display mode, or switching to a specific sensor data source. This allows the operator to tailor the display content to specific mission requirements.

[0193] ④ Improve situational awareness

[0194] By integrating data from multiple sensors into a single view, the fused display helps operators quickly understand complex environments and enhance situational awareness. This is particularly important in scenarios requiring rapid decision-making, such as military reconnaissance and firefighting and rescue operations.

[0195] ⑤ Reduce information overload

[0196] The fusion display screen filters and prioritizes data to display only the most valuable information for the current task, filters out imaging sources with poor imaging effects, and reduces image noise, thereby reducing information redundancy and overload for the operator and improving efficiency and safety.

[0197] Observe the target detection results through the eyepiece set.

[0198] Example 4

[0199] This embodiment introduces a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of the multimodal fusion target detection method suitable for dim environments as described in Example 1 or 2 are implemented.

[0200] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the purpose of the present invention and the claims, which are all protected by the present invention.

Claims

1. A multimodal fusion target detection method suitable for dim environments, characterized in that: include: Obtain the image to be detected; Performing a status judgment on the image to be detected to obtain a status judgment result; Performing thermal imaging channel imaging, ultraviolet light channel imaging, and low-light channel imaging on the image to be detected, respectively, to obtain a thermal imaging channel image, an ultraviolet light channel image, and a low-light channel image; Selecting one or more images from the thermal imaging channel image, the ultraviolet light channel image, and the low-light channel image according to the state judgment result and the preset channel priority; Perform channel imaging screening on the selected image to obtain a screening image; Inputting the screened image into a pre-built image fusion model for image fusion to obtain a fused image; the image fusion model includes an encoder, an attention module, and a decoder connected in sequence; Performing target detection on the fused image to obtain a target detection result.

2. The multimodal fusion target detection method suitable for dim environments according to claim 1, characterized in that: The state judgment result is expressed as: ; in, Indicates the status judgment result; 、 、 、 、 They represent the dynamic weights of the light intensity parameter, polarization degree parameter, thermal radiation difference parameter, ultraviolet characteristic intensity parameter, and low-light signal-to-noise ratio parameter of the image to be detected respectively; 、 、 、 、 They represent the light intensity parameter, polarization degree parameter, thermal radiation difference parameter, ultraviolet characteristic intensity parameter, and low-light signal-to-noise ratio parameter of the image to be detected respectively; represents the feature extraction function; Represents the argmax activation function.

3. The multimodal fusion target detection method suitable for dim environments according to claim 1 or 2, characterized in that: The state judgment results include strong light environment, dim light environment, dark light environment, no light environment, environment blocked by transparent objects, dark shadow environment, and environment with similar temperature.

4. The multimodal fusion target detection method suitable for dim environments according to claim 1, characterized in that: Also includes: When there is interference in the image to be detected, performing polarized light source imaging on the image to be detected to obtain a polarized light image; One or more images are selected from the polarized light image, the thermal imaging channel image, the ultraviolet light channel image, and the low-light channel image.

5. The multimodal fusion target detection method suitable for dim environments according to claim 1, characterized in that: Also includes: Use the acquired polarized light source to detect suspicious targets and determine whether the suspicious targets are enemy infrared detectors or heat sources; If the suspicious target is not an enemy infrared detector or heat source, a feedback signal is sent to the near-infrared light source, and the visible light module image in the thermal imaging channel image is adjusted using the near-infrared light source until the intensity of the near-infrared light source is greater than the brightness of the visible light module image, thereby obtaining an adjusted visible light module image; and a thermal imaging channel adjusted image is obtained based on the adjusted visible light module image and the acquired far-infrared light module image; If the suspected target is an enemy infrared detector or heat source, no feedback signal is sent to the near-infrared light source.

6. The multimodal fusion target detection method suitable for dim environments according to claim 1, characterized in that: The processing steps of the image fusion model include: In the encoder, low-frequency feature extraction and high-frequency feature extraction are sequentially performed on the screened image to obtain mixed features; In the attention module, importance weights are assigned to the mixed features according to the channel importance weights to obtain assigned features; In the decoder, the assigned features are decoded and reconstructed to obtain a fused image.

7. The multimodal fusion target detection method suitable for dim environments according to claim 6, characterized in that: The channel importance weight is expressed as: ; in, represents the importance weight of the i-th channel; represents the signal-to-noise ratio of the i-th channel; Indicates the matching degree of the i-th channel environment; 、 Represents the adjustment coefficient.

8. The multimodal fusion target detection method suitable for dim environments according to claim 1, characterized in that: Before performing target detection on the fused image, the method further includes: The fused image is subjected to brightness adjustment, contrast adjustment, noise filtering, and edge detection to obtain a processed image.

9. A multimodal fusion target detection system suitable for dim environments, characterized by: include: The image acquisition module is used to: acquire the image to be detected; A state judgment module is used to: perform state judgment on the image to be detected and obtain a state judgment result; The channel imaging module is used to perform thermal imaging channel imaging, ultraviolet light channel imaging, and low-light channel imaging on the image to be detected, thereby obtaining a thermal imaging channel image, an ultraviolet light channel image, and a low-light channel image; An image selection module is used to select one or more images from the thermal imaging channel image, the ultraviolet light channel image, and the low-light channel image according to the state judgment result and the preset channel priority; An image screening module is used to: perform channel imaging screening on the selected image to obtain a screening image; An image fusion module is configured to input the screened image into a pre-built image fusion model for image fusion to obtain a fused image; the image fusion model includes an encoder, an attention module, and a decoder connected in sequence; The target detection module is used to perform target detection on the fused image to obtain a target detection result.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the multimodal fusion target detection method applicable to dim environments as described in any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Auxiliary transportation management system based on video sensing technology

    CN120952651A