An adaptive scene-switching target recognition method and system integrating visible light and infrared imaging
By dynamically adjusting the weights of visible light and infrared imaging, and combining a coupling model of light intensity and temperature difference, the problem of insufficient target recognition rate and positioning accuracy caused by changes in light and thermal radiation characteristics in existing technologies is solved, and efficient target recognition in complex scenarios is achieved.
Patent Information
- Application Number
- CN202510731445.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-06-03
AI Technical Summary
Existing fusion methods cannot adapt to dynamic changes in light intensity and thermal radiation characteristics, resulting in insufficient target recognition rate and positioning accuracy in complex scenes. In particular, in environments such as dawn-dusk transition, fog, haze, rain, and snow, existing technologies lack adaptive processing capabilities and cannot effectively fuse information from visible light and infrared imaging.
By acquiring the illumination intensity of visible light images and the temperature difference of infrared images, the weights are dynamically adjusted based on a coupled model. Combined with gradient guidance and local interference suppression strategies, cross-modal feature competition and adaptive scene switching are achieved. The thermal radiation characteristics of infrared images are used to enhance target features and eliminate noise superposition. Database joint analysis and confidence-driven decision-making mechanisms are adopted to improve recognition accuracy.
It significantly improves the signal-to-noise ratio and target contour clarity of fused images in complex and variable scenarios, enhances target recognition accuracy and environmental adaptability, and solves the problem of target recognition under severe light fluctuations and weather interference.
Smart Images

Figure CN120635387B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and specifically to an adaptive scene-switching target recognition method and system that integrates visible light and infrared imaging. Background Technology
[0002] In complex and ever-changing real-world scenarios, target recognition technology faces multiple challenges, including drastic fluctuations in lighting, weather interference, and dynamic changes in thermal characteristics. Existing fusion methods suffer from the following drawbacks: First, fusion strategies based on fixed weights or single thresholds cannot adapt to dynamic changes in light intensity and thermal radiation characteristics. Especially under critical environmental conditions such as dawn / dusk, fog, haze, rain, and snow, the effective information from visible light and infrared imaging exhibits a nonlinear complementary relationship. Existing methods lack a fine perception and dynamic coupling mechanism for environmental parameters, leading to inaccurate fusion weight allocation. Furthermore, existing technologies lack adaptive processing capabilities for local pixel interference (such as headlight glare and heat source noise). In strong backlight or complex thermal background scenes, the noise components of visible light and infrared images are linearly superimposed, causing a sharp drop in the signal-to-noise ratio of the fused target feature image. Most importantly, existing methods have not established a correlation mechanism between modal weights and feature enhancement, resulting in the weakening or masking of key target features during the fusion process. This essentially stems from a lack of pixel-level dynamic perception and nonlinear coupling capability for light intensity and temperature difference, failing to map changes to the feature enhancement mechanism. Ultimately, this results in target recognition rate, positioning accuracy, and environmental adaptability in complex scenes failing to meet practical application requirements. Summary of the Invention
[0003] To address the shortcomings of existing technologies, this invention provides an adaptive scene-switching target recognition method and system that integrates visible light and infrared imaging.
[0004] An adaptive scene-switching target recognition method integrating visible light and infrared imaging includes: acquiring visible light and infrared images of the target area, acquiring the light intensity corresponding to the visible light image, and acquiring the temperature difference corresponding to the infrared image; acquiring a coupling coefficient based on a coupling model, light intensity, and temperature difference, and acquiring the illumination weight of the visible light image and the infrared weight of the infrared image based on the coupling coefficient; acquiring a target feature image based on the illumination weight, infrared weight, visible light image, and infrared image; and acquiring a target recognition result based on the target feature image.
[0005] Optionally, obtaining the illumination intensity corresponding to the visible light image includes: collecting illumination information of the target area based on the photosensitive sensor array, and outputting an illumination intensity distribution map corresponding to the spatial position of the visible light image pixels based on the illumination information; and obtaining the illumination intensity corresponding to the visible light image based on the illumination intensity distribution map.
[0006] Optionally, obtaining the temperature difference corresponding to the infrared image includes: collecting radiation temperature information of the target area based on the infrared sensor array, and outputting a temperature distribution map corresponding to the spatial position of the infrared image pixels based on the radiation temperature information; setting a base background temperature, and obtaining the temperature difference corresponding to the infrared image based on the base background temperature and the temperature distribution map.
[0007] Optionally, obtaining the illumination weight of the visible light image and the infrared weight of the infrared image based on the coupling coefficient includes: directly mapping the coupling coefficient to the illumination weight of the visible light image; and generating the infrared weight of the infrared image based on the illumination weight and a preset weight constraint relationship.
[0008] Optionally, obtaining target recognition results based on target feature images includes: matching and comparing the target feature images with a pre-stored target feature database, determining multiple candidate target categories based on matching similarity, selecting target categories that meet preset confidence conditions from the multiple candidate target categories as target recognition results, and outputting target location and category information.
[0009] Optionally, the coupling model for obtaining the coupling coefficient based on the coupling model, light intensity, and temperature difference is expressed as follows: , ;in, For pixels The corresponding coupling coefficient, The minimum coupling coefficient, pixels in a visible light image Corresponding light intensity This is the critical threshold for light intensity. pixels in an infrared image The corresponding temperature difference This is the adjustment coefficient.
[0010] Optionally, the target feature image obtained based on illumination weights, infrared weights, visible light images, and infrared images can be represented as follows: ;in, For the pixel points in the target feature image The corresponding eigenvalues, For illumination weight, pixels in a visible light image The corresponding gradient magnitude, For infrared weights, pixels in an infrared image The corresponding gradient magnitude.
[0011] An adaptive scene-switching target recognition system integrating visible light and infrared imaging is also provided. The system includes: a data acquisition module for acquiring visible light and infrared images of the target area, acquiring the light intensity corresponding to the visible light image, and acquiring the temperature difference corresponding to the infrared image; a data adaptive coupling module for acquiring coupling coefficients based on a coupling model, light intensity, and temperature difference, and acquiring the illumination weight of the visible light image and the infrared weight of the infrared image based on the coupling coefficients; a data feature processing module for acquiring target feature images based on the illumination weights, infrared weights, visible light images, and infrared images; and a comparison and recognition module for acquiring target recognition results based on the target feature images.
[0012] Optionally, the data acquisition module is also used to: collect illumination information of the target area based on the photosensitive sensor array, and output an illumination intensity distribution map corresponding to the spatial position of the visible light image pixels based on the illumination information; and obtain the illumination intensity corresponding to the visible light image based on the illumination intensity distribution map.
[0013] Optionally, the data acquisition module is also used to: collect radiation temperature information of the target area based on the infrared sensor array, and output a temperature distribution map corresponding to the spatial position of the infrared image pixels based on the radiation temperature information; preset a base background temperature, and obtain the temperature difference corresponding to the infrared image based on the base background temperature and the temperature distribution map.
[0014] The beneficial effects of this invention are reflected in:
[0015] In the adaptive scene-switching target recognition method that integrates visible light and infrared imaging, firstly, based on a nonlinear coupling model of pixel-level illumination intensity and temperature difference, adaptive allocation of cross-modal weights is achieved. In areas with drastic illumination fluctuations (such as strong backlight or dawn-dusk transitions), the visible light weights are dynamically suppressed by the temperature difference to preserve texture details in highlight areas. In low-illuminance or weather-disturbed scenes (such as fog, haze, rain, and snow), the infrared weights are enhanced based on the thermal radiation penetration characteristics. Simultaneously, residual illumination information is used to correct the edge blurring of the thermal image, significantly improving the signal-to-noise ratio and target contour clarity of the fused image. Furthermore, the combination of a gradient-guided feature competition mechanism and a local interference suppression strategy solves the noise superposition problem—for visible light interference such as headlight glare, the dominant mode is dynamically switched through thermal radiation spectrum stability analysis to eliminate false edges. In complex thermal backgrounds, the spatial gradient constraint of the temperature difference is used to isolate thermal noise points and enhance the thermal conduction characteristics of the real target. Finally, the joint database analysis and confidence-driven decision-making mechanism further improve the recognition accuracy. Attached Figure Description
[0016] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.
[0017] Figure 1 This is a schematic diagram of the steps in one embodiment of the adaptive scene switching target recognition method that integrates visible light and infrared imaging of the present invention.
[0018] Figure 2 This is a schematic diagram of a portion of step S1 in the adaptive scene switching target recognition method that integrates visible light and infrared imaging of the present invention.
[0019] Figure 3 This is a schematic diagram of another part of the steps in S1 of the adaptive scene switching target recognition method that integrates visible light and infrared imaging of the present invention.
[0020] Figure 4 This is a schematic diagram of a portion of step S2 in the adaptive scene switching target recognition method that integrates visible light and infrared imaging of the present invention.
[0021] Figure 5 This is a schematic diagram of part of step S4 in the adaptive scene switching target recognition method that integrates visible light and infrared imaging of the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0023] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0024] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, the terms "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0025] like Figure 1 As shown, an adaptive scene-switching target recognition method integrating visible light and infrared imaging is provided, including:
[0026] S1. Obtain the visible light image and infrared image of the target area, and obtain the light intensity corresponding to the visible light image and the temperature difference corresponding to the infrared image;
[0027] S2. Based on the coupling model, light intensity and temperature difference, obtain the coupling coefficient, and obtain the illumination weight of the visible light image and the infrared weight of the infrared image according to the coupling coefficient;
[0028] S3. Obtain the target feature image based on the illumination weight, infrared weight, visible light image, and infrared image;
[0029] S4. Obtain the target recognition result based on the target feature image.
[0030] In this embodiment, it should be noted that in S1, environmental information is collected collaboratively by multimodal sensors to achieve refined perception of light intensity and thermal radiation characteristics. For acquiring light intensity in visible light images, a photosensitive sensor array is directly used to acquire regional block sampling, which is then combined with adaptive interpolation to calculate the global light distribution. For example, in scenes with drastic light changes, the extreme brightness points in the high dynamic range region of the image (such as the center of a strong light source or the edge of a shadow) are extracted first. A light intensity mapping model of sparse sampling points is established by combining local contrast analysis, and then gradient-guided interpolation is used to generate the overall light intensity distribution. This method can capture the large-scale light intensity gradient characteristics during dawn and dusk while avoiding the computational redundancy of independent pixel-by-pixel acquisition, significantly reducing the computational load while maintaining accuracy.
[0031] Furthermore, for infrared images, a temperature field mapping model is constructed based on the physical characteristics of thermal radiation. The target area is scanned using an infrared sensor array, and historical temperature data is statistically analyzed using a sliding window based on the temporal characteristics of infrared images to generate a background temperature baseline in real time (e.g., the persistently low temperature background of snow accumulation in a snow scene). Local temperature anomaly regions of stable heat source targets (e.g., human bodies, vehicle engines) are extracted through inter-frame difference analysis. For example, under headlight glare interference, the transient thermal noise (e.g., light reflection) is distinguished from the real heat source by analyzing the thermal radiation spectrum characteristics. Temperature difference is calculated only for pixels that conform to thermal characteristics, avoiding the resource waste of calculating every point in the entire image. Simultaneously, a temperature field sparse coding technique is used to compress redundant thermal data, reducing the amount of data processed in real time while preserving the edge gradients of key targets, achieving a balance between energy efficiency and accuracy.
[0032] In S2, a dynamic coupling model is used to adaptively allocate visible light and infrared modal weights, addressing the fusion misalignment problem caused by nonlinear changes in environmental parameters. Specifically, the coupling model takes pixel-level illumination intensity and temperature difference as inputs, and simulates their interaction under critical conditions using a nonlinear function: when the illumination intensity is above a critical threshold, the model prioritizes enhancing the visible light weight to utilize its rich texture details. If a significant temperature difference exists in the same area (such as a high-temperature target), the visible light weight is suppressed based on the temperature difference. Conversely, in low-light scenes, the model activates the infrared weight based on the temperature difference to supplement thermal features, and applies a smoothing constraint to the infrared weight through gradient changes in illumination intensity to avoid edge blurring of targets with low temperature differences. For example, in a hazy environment, by analyzing the coupling relationship between the degree of illumination attenuation and the thermal radiation penetration characteristics, the infrared weight is automatically increased in fog-covered areas, while residual illumination information is used to correct blurred edges in the thermal image, achieving cross-modal feature compensation.
[0033] Furthermore, the coupled model introduces a local interference suppression mechanism, distinguishing between real targets and transient noise through temporal stability analysis of temperature differences. When headlight glare is detected causing localized overbrightness in the visible light, the model dynamically reduces the visible light weight and activates the noise suppression coefficient of the infrared mode, based on the abrupt changes in temperature differences in that area (such as the spectral difference between transient thermal reflections and stable heat sources), preventing glare noise from being transmitted to the fused image. In complex thermal backgrounds, the model applies edge-guided constraints to the infrared weights based on the spatial distribution characteristics of temperature differences (such as the transition gradient between heat sources and the background), suppressing thermal diffusion noise while enhancing the target contour. For example, in snow scenes, by analyzing the spatial distribution of the temperature difference between the low-temperature snow background and the human body heat source, the visible light weight is reduced in the snow-reflective areas, while the infrared gradient weight is enhanced at the edges of the human body contour, preventing the target in the fused image from being overwhelmed by snow reflections or low-temperature noise.
[0034] In S3, a gradient-guided cross-modal feature competition mechanism is used to achieve complementarity of key target features. Based on the spatial gradient distribution of visible and infrared images, dynamic weights are used to nonlinearly modulate the two types of gradients: when the combined effect of the visible light gradient amplitude and its weight at a pixel is significantly higher than that of the infrared mode, visible light edge features are amplified while infrared noise is suppressed; conversely, when the infrared gradient exhibits a stronger target contour response under low illumination or complex thermal backgrounds, infrared features are preferentially preserved while visible light interference components are weakened. For example, in dense fog, the road marking gradient in the visible light image is attenuated due to scattering effects, while the thermal radiation gradient of the vehicle engine in the infrared image remains clear. In this case, the contribution of the infrared gradient to the fused features is automatically enhanced through weight comparison, achieving cross-modal feature complementarity.
[0035] Furthermore, a noise suppression strategy with gradient competition threshold adaptation is introduced to address the feature confusion problem caused by local interference. When an abnormal increase in gradient amplitude is detected in a visible light highlight region (such as headlight glare) and the corresponding infrared location exhibits abnormal temperature difference fluctuations, the fusion ratio of the visible light gradient in that region is reduced through dynamic weighting, and gradient features from the infrared image that have been verified for thermal stability are instead used. For thermal background clutter regions, the gradient response of isolated thermal noise points is suppressed through the spatial continuity constraint of the infrared weights. For example, in a snowy, highly reflective scene, the false visible light edges caused by snow reflection are filtered out by dynamic weight reduction, while the continuous temperature difference gradient of the human target in the infrared image is enhanced. The final fused image can both eliminate reflective artifacts and completely preserve the true edges of the human thermal features and the contact surface with the snow.
[0036] In S4, based on the fused target feature image, multi-level features are first extracted to construct cross-modal feature descriptors, which are then matched with a pre-stored database in multiple dimensions. For example, in a hazy scene, for vehicle targets affected by fog, by comparing the vehicle features retained in the fused features with the features of different vehicle models in the database, the vehicle type (truck or sedan) and its location can be effectively identified. Simultaneously, a spatiotemporal continuity verification mechanism is introduced to analyze the correlation between the target's motion trajectory and thermal feature changes in adjacent frames, eliminating false matches caused by instantaneous noise.
[0037] Furthermore, a confidence-driven dynamic decision-making mechanism is employed to optimize recognition accuracy. For the candidate target set generated by matching, a multi-dimensional confidence evaluation model is constructed by integrating local feature matching degree, environmental interference suppression coefficient, and target motion consistency: In strong backlight scenarios, for targets lacking visible light features due to glare, the infrared modal confidence weight is increased through thermal radiation stability analysis, and trajectory prediction is performed to complete the target's trajectory based on historical recognition results; for targets with low temperature differences in snow, the recognition threshold is dynamically adjusted through joint calculation of edge gradient continuity detection and background thermal noise suppression coefficient. For example, when recognizing a stationary human body in snow, the thermal conduction gradient characteristics of the contact surface between the target and the snow are analyzed to distinguish between a real human body and low-temperature interference objects such as rocks. Simultaneously, historical motion data is used to eliminate false triggers caused by snow slippage, ultimately outputting a recognition result that combines real-time performance and reliability.
[0038] In summary, the adaptive scene-switching target recognition method integrating visible light and infrared imaging firstly achieves adaptive allocation of cross-modal weights based on a nonlinear coupling model of pixel-level illumination intensity and temperature difference. In areas with drastic illumination fluctuations (such as strong backlighting or twilight), the visible light weights are dynamically suppressed by temperature difference to preserve texture details in high-light areas. In low-light or weather-disturbed scenarios (such as fog, haze, rain, and snow), the infrared weights are enhanced based on thermal radiation penetration characteristics. Simultaneously, residual illumination information is used to correct edge blurring in thermal images, significantly improving the signal-to-noise ratio and target contour clarity of the fused image. Furthermore, the combination of a gradient-guided feature competition mechanism and a local interference suppression strategy solves the noise superposition problem—for visible light interference such as headlight glare, the dominant mode is dynamically switched through thermal radiation spectrum stability analysis to eliminate false edges. In complex thermal backgrounds, the spatial gradient constraint of temperature difference is used to isolate thermal noise points and enhance the thermal conduction characteristics of the real target. Finally, the joint database analysis and confidence-driven decision-making mechanism further improve recognition accuracy.
[0039] like Figure 2 As shown, in one embodiment, obtaining the illumination intensity corresponding to the visible light image in S1 includes:
[0040] S11. Collect illumination information of the target area based on the photosensitive sensor array, and output an illumination intensity distribution map corresponding to the spatial position of the visible light image pixels based on the illumination information;
[0041] S12. Obtain the light intensity corresponding to the visible light image based on the light intensity distribution map.
[0042] In this embodiment, it should be noted that in S11, a photosensitive sensor array is used to perform refined sensing of the light intensity of the target area. To address the problem of drastic light fluctuations in complex scenes (such as global gradual changes during dawn / dusk or localized high dynamic range under strong backlight), the sensor employs a strategy combining sparse sampling and priority capture of key areas: firstly, the sampling density is dynamically divided based on scene features; dense sampling points are deployed in areas of abrupt light change (such as the center of vehicle headlight glare or the edge of building shadows), and key nodes of brightness change are extracted using local contrast extreme value detection; for areas with moderate light, sparse sampling is used to avoid redundant data acquisition. For example, in a hazy environment, the sensor prioritizes capturing residual light intensity extreme points penetrating the fog, and combines this with an atmospheric scattering model to infer the impact of fog thickness on light attenuation, constructing a sparse sampling network covering the entire area. This process is processed in real-time by an edge computing device, completing the initial light intensity feature extraction at the sensor end, reducing data transmission pressure.
[0043] In S12, pixel-level fast matching between the illumination intensity distribution map and the visible light image is performed, allowing each pixel of the visible light image to be directly associated with the illumination intensity value at the corresponding location.
[0044] like Figure 3 As shown, in one embodiment, obtaining the temperature difference corresponding to the infrared image in S1 includes:
[0045] S13. Collect radiation temperature information of the target area based on the infrared sensor array, and output a temperature distribution map corresponding to the spatial position of the infrared image pixels based on the radiation temperature information.
[0046] S14. Preset the base background temperature and obtain the temperature difference corresponding to the infrared image based on the base background temperature and temperature distribution map.
[0047] In this embodiment, it should be noted that in S13, the precise capture and spatial mapping of thermal radiation features are achieved through an infrared sensor array. For complex thermal backgrounds (such as snow reflection interference or noise from moving heat sources), the sensor array employs an adaptive scanning strategy: increasing the sampling frequency in areas with dense high-temperature targets (such as vehicle engine clusters) to capture subtle temperature fluctuations, while reducing the sampling density in low-temperature background areas (such as snow or water) to save energy. For example, in a headlight glare scenario, by analyzing the spectral characteristics of the thermal radiation signal (such as the transient nature of reflected heat pulses and the persistence of engine heat sources), pseudo-temperature difference signals caused by light reflection are dynamically filtered out, retaining only the true target temperature data that conforms to the laws of heat conduction. Simultaneously, the sensor fuses the texture features of the infrared image, performs edge-guided spatial correction on the temperature distribution map, eliminates temperature field distortion caused by sensor parallax, and ensures pixel-level alignment between the thermal radiation data and the infrared image.
[0048] In S14, accurate separation of target thermal features is achieved through dynamic background temperature modeling and anomalous temperature difference extraction techniques. First, based on the spatiotemporal continuity of the thermal scene, a sliding window is used to statistically analyze the temperature distribution patterns of historical frames to construct an adaptive unified background temperature baseline. Then, a temperature difference distribution map is obtained based on the temperature distribution map and the unified background temperature baseline. This temperature difference distribution map is then rapidly matched with the pixel-level data of the infrared image, allowing each pixel in the infrared image to be directly associated with the temperature difference at its corresponding location.
[0049] like Figure 4 As shown, in one embodiment, obtaining the illumination weight of the visible light image and the infrared weight of the infrared image based on the coupling coefficient in S2 includes:
[0050] S21. Directly map the coupling coefficient to the illumination weight of the visible light image;
[0051] S22. Generate the infrared weights of the infrared image based on the illumination weights and the preset weight constraint relationship.
[0052] In this embodiment, it should be noted that in S21, the dynamic response of the modal weights is achieved through a direct mapping between the coupling coefficient and the illumination weights. For example, the coupling coefficient is directly used as the illumination weight (coupling coefficient = illumination weight). This mapping mechanism preserves the coupling model's ability to analyze the nonlinear relationship between illumination intensity and temperature difference, ensuring that the weight allocation can reflect the microscopic changes in environmental parameters in real time. For instance, in a scene of alternating dawn and dusk, when the illumination intensity in a certain area is near the critical threshold and there is a high-temperature target (such as vehicle exhaust), the coupling coefficient will dynamically fine-tune the visible light weights according to the temperature difference—utilizing the residual texture of visible light under critical illumination while suppressing overexposure caused by reflection from metal surfaces through temperature difference suppression. This direct mapping method avoids the information attenuation caused by secondary transformation in existing methods, enabling the weight adjustment in high dynamic range scenes to have a sub-pixel-level response speed.
[0053] In S22, a global balance of cross-modal information is achieved through complementary weight constraints. For example, the infrared weight is calculated by subtracting the illumination weight from 1 and adding the minimum base weight (1 - illumination weight + min base weight = infrared weight). This ensures a strict inverse correlation between the infrared and visible light weights, forcing the fusion process to retain effective components of dual-modal information at any pixel. For instance, in a dense fog environment, when fog causes an overall decrease in the visible light weight, the infrared weight automatically becomes dominant. However, through residual illumination gradient constraints in the coupling model (such as the outline of a directional light source penetrating through the fog), the infrared weight retains the ability to correct visible light edges during its enhancement. This constraint mechanism not only prevents the complete failure of a single mode but also naturally suppresses noise superposition through weight competition. When the weight in a certain visible light region increases abnormally due to glare, the corresponding decrease in the infrared weight blocks the transmission of this noise to the fusion result, while compensating using the stable thermal characteristics of the infrared mode in that region.
[0054] like Figure 5 As shown, in one embodiment, obtaining the target recognition result based on the target feature image in step S4 includes:
[0055] S41. Match and compare the target feature image with the pre-stored target feature database, and determine multiple candidate target categories based on the matching similarity.
[0056] S42. Select the target category that meets the preset confidence conditions from multiple candidate target categories as the target recognition result, and output the target location and category information.
[0057] In this embodiment, it should be noted that in S41, accurate matching of target features is achieved through a cross-modal feature association model. Based on multi-level features extracted from fused images (such as edge topology and thermal radiation distribution patterns in existing technologies), a composite descriptor compatible with the database is constructed. An adaptive similarity measurement strategy from existing technologies is adopted to address environmental interference: in hazy scenes, the penetrability feature weight of the infrared mode is strengthened to compensate for the texture details lost due to visible light scattering; at the same time, a local occlusion analysis mechanism from existing technologies is introduced to dynamically partition the feature matching area, avoiding global similarity distortion caused by fog occlusion. For example, when identifying vehicles partially covered by snow, priority is given to matching the thermal conduction contour of the unoccluded area and the vehicle body geometric features remaining in the visible light. Combined with motion trajectory prediction, the feature vector of the snow-covered part is completed, effectively improving the matching accuracy of some visible targets. The spatiotemporal continuity verification mechanism further filters transient mismatches by analyzing the motion inertia and thermal feature change patterns of the target in multiple frames (such as the gradual change characteristics of engine temperature), eliminating false candidate categories caused by falling snowflakes or light and shadow swaying.
[0058] In S41, a confidence-based decision model from existing technologies is employed to address the ambiguity of target categories in complex environments. This model integrates feature matching degree, environmental interference suppression coefficient, and motion consistency index, dynamically adjusting the weights of each dimension: in strong backlight scenarios, the evaluation weight of the thermal radiation stability index is increased, and matching reliability is verified by analyzing the thermal inertia characteristics of high-temperature targets (such as the temperature change rate of vehicle exhaust pipes); in snowy low-temperature scenarios, the combined effect of edge gradient continuity and background noise suppression coefficient is enhanced, and the thermal diffusion characteristics of the target contour (such as the thermal conduction gradient between the human body and clothing) are analyzed to distinguish between real targets and static thermal artifacts. For example, when identifying stationary targets in snow, by combining thermal conduction pattern verification (such as the temperature difference attenuation between the human body and rocks) and motion history data (such as the temporal rationality of the target's location), even if the target's temperature difference is close to the environmental noise level, the multi-dimensional confidence assessment can still accurately determine the target category and output a probabilistically weighted recognition result and spatial positioning information.
[0059] In one implementation, the coupling coefficient obtained in S2 based on the coupling model, light intensity, and temperature difference is expressed as follows:
[0060] , ;in,
[0061] For pixels The corresponding coupling coefficient, The minimum coupling coefficient, pixels in a visible light image Corresponding light intensity This is the critical threshold for light intensity (light intensity at dusk, such as 50 lux). pixels in an infrared image The corresponding temperature difference This is the adjustment coefficient.
[0062] In this embodiment, it should be noted that throughout the entire expression, This is used to distinguish whether the illumination exceeds a critical threshold, resolving the weight abrupt change problem in critical light intensity scenarios such as dawn / dusk transitions. Specifically, when... hour, =1, allowing the temperature difference to adjust the visible light weight, entering the visible light-dominated mode, preserving texture details; when hour, Forced entry into low-light mode, triggering infrared enhancement, and forcing visible light weight to be no less than [a certain value]. Since the exponent term is always non-negative, Therefore, through the outer layer Constraints, ultimately take This ensures that visible light retains its basic contribution even under extremely low illumination conditions (such as metallic reflections under moonlight).
[0063] Furthermore, This is a fusion correction term; middle, The normalized offset used for light intensity relative to a critical value eliminates the influence of dimensions, and the absolute value is used to eliminate sign interference, ensuring that the change in light intensity always participates in the calculation in a positive direction. Used for regulating temperature difference, among which This is an adjustment coefficient used to ensure... and The sensitivity (e.g., k=0.05 increases the influence of temperature difference as a two-digit number).
[0064] Furthermore, when ,(like =1000 lux =50 lux, =10 degrees Celsius Take 0.05). =19, ,at this time Approaching 0.76, the coupling coefficient represents the extent to which visible light exceeds the influence of infrared light. When (like =50 lux, =50 lux), 0 , at this time equal In this case, the visible light weight is relatively low, and the temperature difference does not affect the coupling coefficient, thus avoiding texture loss. Slightly larger ,at this time 1, 1, Able to easily influence The value, for example, when =10 degrees Celsius If we take 0.05, then ,at this time Approaching 0.42, the coupling coefficient represents the extent to which infrared light exceeds visible light in terms of influence.
[0065] In one implementation, the target feature image obtained in S3 based on the illumination weight, infrared weight, visible light image, and infrared image is represented as follows:
[0066] ;in,
[0067] For the pixel points in the target feature image The corresponding eigenvalues, For illumination weight, pixels in a visible light image The corresponding gradient magnitude, For infrared weights, pixels in an infrared image The corresponding gradient magnitude.
[0068] In this embodiment, it should be noted that the gradient magnitude is used throughout the expression. and All of these are obtained using the Sobel operator in existing technologies. The Sobel operator is a classic image edge detection algorithm that identifies edges by calculating the spatial gradient of image gray values. Its core idea is to enhance the contrast of object boundaries in an image through directional differentiation. It is widely used in target recognition, image fusion and other fields.
[0069] Furthermore, and Both are products of gradient magnitude and weight, enhancing the effective features of the dominant mode. Higher weights result in a greater gradient contribution from the corresponding mode. This solves the feature weakening problem caused by fixed weights in existing methods. For example, in areas with strong illumination (… =0.9), even if the visible light gradient is small (e.g., =10), the product is 0.9 =9 still retains detail; in areas with low temperature differences ( =0.2), if the infrared gradient is significant (e.g., =30), and 0.2*30=6 cannot dominate the fusion result.
[0070] Furthermore, and All are correction terms, reflecting the competition mechanism; when At that time, within the visible light gradient amplitude, =1, the visible light coefficient is 1.5, while in the infrared gradient amplitude, the infrared coefficient is 0.5; when At this time, the sign is reversed, the infrared term is enhanced by 1.5 times, and the visible light term is weakened to 0.5 times. This suppresses the superposition of dual-mode noise; for example, the visible light weight in the headlight glare area is artificially high due to glare (…). =0.7), but its Anomalies indicate noise; infrared weighting =0.3, but the gradient is stable, reflecting the true edge; the noise term is filtered out by the max function, and finally the infrared dominates.
[0071] Furthermore, the maximum value function The choice is to retain the most prominent feature modes in the current environment and avoid detail blurring caused by weighted averaging. It solves the problem of blurred edges on targets with low temperature differences, such as a human body in snow: the visible light weight is low ( =0.3), the gradient is disturbed by snow reflection ( =15 pseudo-edge); high infrared weighting ( =0.7), gradient stable ( =10 true contour); infrared term equals 10.5, visible light term equals 2.25; the result is that infrared features are preserved and visible light noise is eliminated.
[0072] An adaptive scene-switching target recognition system integrating visible light and infrared imaging is also provided. The system includes: a data acquisition module for acquiring visible light and infrared images of the target area, acquiring the light intensity corresponding to the visible light image (each pixel corresponds to one light intensity), and acquiring the temperature difference corresponding to the infrared image (each pixel corresponds to one temperature difference); a data adaptive coupling module for acquiring a coupling coefficient based on a coupling model, light intensity, and temperature difference, and acquiring the illumination weight of the visible light image and the infrared weight of the infrared image based on the coupling coefficient; a data feature processing module for acquiring a target feature image based on the illumination weight, infrared weight, visible light image, and infrared image; and a comparison and recognition module for acquiring the target recognition result based on the target feature image.
[0073] In one embodiment, the data acquisition module is further configured to: collect illumination information of the target area based on the photosensitive sensor array, and output an illumination intensity distribution map corresponding to the spatial position of the visible light image pixels based on the illumination information; and obtain the illumination intensity corresponding to the visible light image based on the illumination intensity distribution map.
[0074] In one embodiment, the data acquisition module is further configured to: collect radiation temperature information of the target area based on the infrared sensor array, and output a temperature distribution map corresponding to the spatial position of the infrared image pixels based on the radiation temperature information; preset a base background temperature, and obtain the temperature difference corresponding to the infrared image based on the base background temperature and the temperature distribution map.
[0075] In this embodiment, it should be noted that the specific method of performing the operation of the above-mentioned adaptive scene switching target recognition system that integrates visible light and infrared imaging has been described in detail in the embodiments of the adaptive scene switching target recognition method that integrates visible light and infrared imaging, and will not be elaborated here.
[0076] The preferred embodiments of this disclosure have been described in detail above with reference to the accompanying drawings. However, this disclosure is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this disclosure, various simple modifications can be made to the technical solutions of this disclosure, and these simple modifications all fall within the protection scope of this disclosure.
[0077] It should also be noted that the various specific technical features described in the above embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, this disclosure will not describe the various possible combinations separately.
[0078] Furthermore, various different embodiments of this disclosure can be combined in any way, as long as they do not violate the spirit of this disclosure, they should also be regarded as the content disclosed in this disclosure.
[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the claims and specification of the present invention.
Claims
1. A method for adaptive scene switching target recognition fusing visible light and infrared imaging, characterized in that, The method comprises the following steps: obtaining a visible light image and an infrared image of a target region, and obtaining an illumination intensity corresponding to the visible light image and a temperature difference amount corresponding to the infrared image; obtaining a coupling coefficient based on a coupling model, the illumination intensity and the temperature difference amount, and obtaining an illumination weight of the visible light image and an infrared weight of the infrared image according to the coupling coefficient; wherein the coupling model is represented as: , ; wherein, is a pixel point corresponding coupling coefficient, is a minimum coupling coefficient, is a pixel point corresponding light intensity, is a light intensity critical threshold, is a pixel point corresponding temperature difference amount, is an adjustment coefficient; obtaining a target feature image according to the illumination weight, the infrared weight, the visible light image and the infrared image; obtaining a target recognition result according to the target feature image. 2.The method of claim 1, wherein, The method further comprises the following steps: obtaining illumination information of the target region according to a photosensitive sensor array, and outputting an illumination intensity distribution corresponding to a pixel spatial position of the visible light image according to the illumination information; obtaining the illumination intensity corresponding to the visible light image according to the illumination intensity distribution. 3.The method of claim 1, wherein, The method further comprises the following steps: obtaining radiation temperature information of the target region according to an infrared sensor array, and outputting a temperature distribution corresponding to a pixel spatial position of the infrared image according to the radiation temperature information; pre-setting a basic background temperature, and obtaining the temperature difference amount corresponding to the infrared image according to the basic background temperature and the temperature distribution. 4.The method of claim 1, wherein, The method further comprises the following steps: directly mapping the coupling coefficient to the illumination weight of the visible light image; generating the infrared weight of the infrared image according to a preset weight constraint relationship of the illumination weight. 5.The method of claim 1, wherein, The method further comprises the following steps: matching and comparing the target feature image with a pre-stored target feature database, determining a plurality of candidate target categories according to a matching similarity; selecting a target category meeting a preset confidence condition from the plurality of candidate target categories as the target recognition result, and outputting target position and category information. 6.The method of claim 1, wherein, The method is represented as: ; wherein, is a pixel point in the target feature image corresponding feature value, is an illumination weight, is a pixel point in the visible light image corresponding gradient amplitude, is an infrared weight, is a pixel point in the infrared image corresponding gradient amplitude.
7. An adaptive scene switching target recognition system fusing visible and infrared imaging, characterized in that, The system comprises: a data acquisition module, configured to obtain a visible light image and an infrared image of a target region, and obtain an illumination intensity corresponding to the visible light image and a temperature difference amount corresponding to the infrared image; a data adaptive coupling module, configured to obtain a coupling coefficient based on a coupling model, the illumination intensity and the temperature difference amount, and obtain an illumination weight of the visible light image and an infrared weight of the infrared image according to the coupling coefficient; wherein the coupling model is represented as: , ; wherein, is a pixel point corresponding coupling coefficient, is a minimum coupling coefficient, is a pixel point corresponding illumination intensity in the visible light image, is a light intensity critical threshold, is a pixel point corresponding temperature difference in the infrared image, is an adjustment coefficient; a data feature processing module, configured to obtain a target feature image according to the illumination weight, the infrared weight, the visible light image and the infrared image; a comparison and recognition module, configured to obtain a target recognition result according to the target feature image.
8. The adaptive scene cut target recognition system for fusing visible and infrared imaging according to claim 7, wherein, The data acquisition module is further configured to: obtain illumination information of the target region according to a photosensitive sensor array, and output an illumination intensity distribution corresponding to a pixel spatial position of the visible light image according to the illumination information; obtain the illumination intensity corresponding to the visible light image according to the illumination intensity distribution.
9. The adaptive scene cut target recognition system for fusing visible and infrared imaging according to claim 7, wherein, The data acquisition module is further configured to: obtain radiation temperature information of the target region according to an infrared sensor array, and output a temperature distribution corresponding to a pixel spatial position of the infrared image according to the radiation temperature information; A preset basic background temperature is set, and a temperature difference amount corresponding to the infrared image is obtained according to the basic background temperature and the temperature distribution map.
Citation Information
Patent Citations
Image fusion method and related assembly
CN117830121A
Night storage robot target detection method and system based on integral network
CN118097089A
Image digital processing method based on infrared polarized light imaging
CN119693599A