Fire scene image denoising enhancement and splicing method

CN122820488APending Publication Date: 2026-09-25JILIN AGRICULTURAL UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611294816.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-25
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]目前,国内多数消防头盔的研发仍聚焦于抗冲击、防火防热等基础防护功能,仅有少量研究机构将摄像、无线通讯功能集成于消防头盔,但配套的图像处理技术成熟度低,无法适配火灾现场的极端环境

Benefits of technology

本发明针对火场明火、高温自发光体的特点,建立“自发光与反射光”的双分量物理模型,结合暗通道先验与双目深度融合实现精准透射率与景深估计;采用时空域分区域去噪策略,动态估计混合噪声参数,同时预标记危险区域并衰减去噪强度,既有效抑制浓烟、粉尘带来的复杂噪声,又避免模糊边缘、滤除微弱危险信号。同时,本发明基于透射率划分轻、中、浓烟区域,采用非线性Beta变换实现差异化增强,同时引入景深、镜头污损、危险区域、自发光等多因子逐像素调整增强强度,既提升暗区被困人员、障碍物的可见度,又抑制亮区强光过曝;配合色度漂移补偿、空间/帧间亮度硬约束,还原色彩的同时避免视觉疲劳与眩晕,适配长时间佩戴需求。通过Harris与SIFT特征的融合,结合深度层级标注、静动态特征分类与静态先验库,解决浓烟低纹理环境下特征点失效、配准精度低的问题;采用多平面分层单应性配准与景深加权融合,突破单一平面假设的局限,配合动态遮挡检测、帧间防抖平滑,有效消除重影、拼接缝与画面抖动,实现180° 无缝全景拼接,满足无死角环境观测需求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820488A_ABST
    Figure CN122820488A_ABST
Patent Text Reader

Abstract

The application provides a fire scene-oriented image denoising enhancement and splicing method, and relates to the field of image data processing, comprising: image acquisition and preprocessing, adaptive denoising processing, adaptive enhancement and damage compensation, image feature point extraction and matching, and multi-path image splicing. The method realizes front-end integrated real-time processing of fire scene panoramic images through modules integrated in the front-end embedded platform of the helmet, provides near-eye real-time visual assistance for front-end firefighters, reduces wireless transmission bandwidth pressure and transmission delay, and is suitable for extreme working conditions such as high temperature, thick smoke and high dust in the fire scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image data processing technology, specifically to an image denoising enhancement and stitching method for fire scenes. Background Technology

[0002] Fire rescue scenes present extremely complex environments with dense smoke, high temperatures, low light, and high dust levels. In recent years, there have been frequent cases of firefighters being injured or killed during fire rescue operations. Among the core causes of these injuries and deaths are misjudgment of the situation at the fire scene, impaired vision due to dense smoke leading to an inability to accurately assess the environment, and disorientation. As a core piece of individual equipment for firefighters, the intelligent fire helmet, equipped with an image acquisition and processing system, assists firefighters in obtaining information about the fire scene environment, thereby improving rescue efficiency and personnel safety.

[0003] Currently, the research and development of most fire helmets in China still focuses on basic protective functions such as impact resistance, fire protection, and heat protection. Only a few research institutions have integrated camera and wireless communication functions into fire helmets, but the accompanying image processing technology is immature and cannot adapt to the extreme environment of fire scenes. At the same time, existing image processing for fire scenes mostly involves simply preprocessing the acquired images and then transmitting them directly to the back-end command platform via wireless network, without performing in-depth image processing of the fire scene environment in the front-end embedded devices, thus failing to provide real-time visual enhancement assistance to front-end firefighters.

[0004] In addition, existing image processing methods for fire scenes still have the following problems: 1. Existing image denoising algorithms cannot adapt to the complex noise environment of fire scenes: mean denoising easily causes blurring of image edges and details, resulting in the loss of key target information at the fire scene; median denoising significantly reduces its denoising effect when faced with large areas of mixed noise at the fire scene; wavelet denoising algorithms are highly complex and difficult to run in real time on the embedded low-power hardware platform of fire helmets, failing to meet the real-time requirements of fire scene rescue. 2. Existing image enhancement methods have poor adaptability to images with low contrast and uneven lighting at fire scenes: conventional linear enhancement methods easily cause local overexposure or underexposure of images, failing to simultaneously enhance key targets such as trapped personnel and fire sources in dark areas of the fire scene, while suppressing strong light interference in bright areas, making it difficult to improve firefighters' ability to identify the fire scene environment. 3. Existing image stitching methods cannot meet the requirements for real-time generation of panoramic fire scene images: Gray-scale registration-based methods are computationally intensive, have poor real-time performance, and are easily affected by image rotation, deformation, and occlusion; single-feature registration-based methods (such as Harris and SIFT algorithms) are prone to feature point extraction failure and insufficient registration accuracy in dense smoke and low-texture environments at fire scenes, making it impossible to achieve seamless stitching of images from multiple cameras and failing to meet the needs of firefighters for panoramic, blind-spot-free environmental observation. 4. Existing technologies have not formed an integrated image processing workflow adapted to the embedded platform of fire helmets: The existing "front-end acquisition - back-end processing" separation mode cannot provide real-time visual assistance to front-end firefighters, and also increases the data bandwidth pressure of wireless transmission, easily causing transmission delays and affecting the real-time performance of rescue command. Summary of the Invention

[0005] To address the problems existing in the prior art, the present invention aims to provide an image denoising enhancement and stitching method for fire scenes. This method achieves integrated real-time processing of panoramic fire scene images through a module of an embedded platform integrated in the front end of a helmet, adapting to extreme working conditions such as high temperature, dense smoke, and high dust in fire scenes.

[0006] The objective of this invention is achieved through the following technical solution: A method for image denoising, enhancement, and stitching in fire scenes includes: Step S1, Image Acquisition and Preprocessing: Images are acquired using three wide-angle high-definition cameras mounted on the left, center, and right sides of the helmet's front end, and preprocessed using an embedded ARM platform within the helmet. Step S2, Adaptive Denoising Processing: This includes separation of reflected light and self-emission dual components and transmittance estimation, mixed noise parameter estimation, hazard feature pre-detection and protected area marking, temporal motion saliency detection and region division, spatial region division, and regional denoising processing. Step S3, Adaptive Enhancement and Stain Compensation: This includes dual-region partitioning of reflected light and self-emission, construction of nonlinear enhancement transformation model, optimization of partition parameters, multi-factor pixel-by-pixel enhancement intensity adjustment, application of hard constraints, chromaticity drift partition compensation, and color image reconstruction. Step S4, Image Feature Point Extraction and Matching: This includes fusion feature point extraction, feature point description and normalization, feature point depth-level annotation, static-dynamic feature classification, hierarchical static prior coarse matching, and fine matching; Step S5, Multi-channel image stitching: Using multi-plane layered registration, dynamic occlusion detection, and layered weighted fusion mechanism, seamless stitching of panoramic fire scene images is achieved.

[0007] Based on further optimization of the above scheme, in step S1, the left, center, and right wide-angle high-definition cameras synchronously acquire images of the fire scene environment at a frame rate of 10 frames per second, with an original image resolution of 720p. The embedded ARM platform includes an ARM processor, a DDR4 memory module, a communication module, a display module, and a power management module. The ARM processor adopts an embedded processor with an ARM Cortex-A55 architecture, and has a built-in junction temperature sensor and performance counter. The left, center, and right wide-angle high-definition cameras are synchronously triggered through the GPIO interface and connected to the ARM processor through the MIPI CSI / DVP interface. The DDR4 memory module is connected to the ARM processor through the memory bus. The communication module adopts any one or more combinations of WIFI module, 4G module / 5G module, and is connected to the ARM processor through the USB interface. The display module is connected to the ARM processor through the MIPI interface. The power management module provides regulated power supply for all hardware units.

[0008] Based on further optimization of the above scheme, step S1, which utilizes the embedded ARM platform within the helmet for image preprocessing, specifically includes: First, the three acquired RGB raw images are converted to YC. b C r Color space, separating the luminance component Y and chrominance component C of an image. b C r ; Then, the lens distortion invalid areas at the edges of each image are cropped, retaining the effective imaging area of ​​1200×680 in the center; at the same time, a dedicated buffer area is opened to store the brightness component Y data and timing marks of 5 consecutive frames. Subsequently, based on local sharpness differences, areas of contamination such as smoke and water droplets on the lens surface are detected, and a contamination mask is generated. M smudge : ; In the formula: Clear()Indicates the local sharpness value. Laplacian() Represents the Laplace operator. gray(x,y) Representing coordinates (x,y) The single-channel grayscale value at that location (corresponding to the luminance component Y); When the local sharpness is less than 50% of the global average sharpness and the area of ​​a continuous region is greater than 30 pixels, it is marked as a dirty region, i.e., the mask value is 1; the rest are normal regions, i.e., the mask value is 0; the mask is updated every 30 frames.

[0009] Based on further optimization of the above scheme, step S2 uses the current frame luminance component Y and the original chrominance component C output in step S1. b With C r Contamination mask M smudge For input; The separation of reflected light and spontaneous emission components and the estimation of transmittance are specifically as follows: Addressing the issue of traditional scattering models failing due to spontaneous emission sources in fire scenes, the imaging is decomposed into a spontaneous emission component and a scene reflected light component, establishing a dual-component imaging physical model. ; In the formula: This represents the self-luminous component (i.e., the light emitted by an open flame or a high-temperature object). Indicates the components of light reflected from the scene. t(x,y) Indicates smoke transmittance. A f Indicates the global ambient light intensity; Then, a coarse detection of the self-illuminating region is performed: first, the redness ratio is calculated: ; Preset percentage threshold T cr ,like R ratio > T cr If the value is 0, it is marked as a chroma candidate pixel; otherwise, it is not marked. Next, calculate the maximum brightness within a 15×15 neighborhood window. If the current pixel brightness is equal to the local maximum and is higher than 1.5 times the global average brightness, then mark it as a brightness candidate pixel. If a pixel is both a chroma candidate pixel and a luminance candidate pixel, it is marked as a self-illuminating candidate; and a 3×3 morphological dilation is performed on the candidate region to obtain a coarse mask. M emissive-c ; Then, for a consecutive 5-frame luminance sequence, the temporal mean and temporal variance of each pixel are calculated: ; And the pixel fluctuation frequency is calculated using the luminance zero-crossing rate: ; In the formula: Indicates the frame interval time. sign() Represents a symbolic function; like Z(x,y) If the frequency is ∈[8Hz,15Hz], then it is determined to have flame flicker characteristics; If the coarse mask simultaneously satisfies the time-domain variance > 20 and Z(x,y) If the Hz range is ∈[8Hz, 15Hz], it is retained as a self-emissive region; static high-temperature self-emissive bodies have no flicker characteristics, and their chromaticity and brightness are directly preserved to obtain the final self-emissive mask. M emissive ; Next, estimate the transmittance of the non-emissive region: for the luminance component of the non-emissive region, calculate the minimum luminance value within the local window: ; In the formula: Indicates ( x,y A local window centered on ); Take the top 0.1% of pixels in the dark channel brightness and calculate the average of their corresponding original brightness as the global ambient light intensity. A f ; The initial transmittance is: , Indicates the smoke retention coefficient; Using the original brightness map as a guide map, the initial transmittance is optimized by guided filtering: ; In the formula: Represents the linear coefficients within the window; W k This represents the k-th filtering window; Perform time-domain weighted smoothing of the transmittance of the current frame and historical frames: ; In the formula: Represents the smoothing coefficient. This represents the optimized transmittance of the current frame; The scattering model is not applicable to the self-emissive region, so the transmittance is directly set to 1; For non-self-emissive areas (i.e.) M emissive =0), directly inversely solve for the reflected light component: ; For self-emissive areas (i.e.) M emissive =1): ; Finally, according to the Beer-Lambert law, the model depth of field is obtained through inversion: ; In the formula: Indicates the smoke scattering coefficient; For close-up areas within 3 meters, depth is calculated using the binocular parallax of the left and right cameras: ; In the formula: b d Indicates the binocular baseline distance. f d Indicates the camera's pixel focal length. This represents the disparity between corresponding matching pixels in the left and right images; Near-field area with binocular depth For reference, the mid-to-long-range views are based on the model's depth of field. To determine the optimal depth-of-field map, an S-shaped weighted smooth transition is applied to the transition area, resulting in the final depth-of-field map. d(x,y) .

[0010] Based on further optimization of the above scheme, in step S2, the estimation of the mixed noise parameters specifically involves: fitting a Poisson-Gaussian mixed noise model using maximum likelihood estimation and dynamically updating the noise parameters. ; In the formula: Poisson() Indicates the Poisson distribution. z Indicates the grayscale value of the observed pixel. Indicates Poisson noise intensity; Gauss() Indicates a Gaussian distribution. This represents the mean of Gaussian noise. Indicates the standard deviation of Gaussian noise; This represents the convolution operation; Noise parameters are updated once per frame to achieve noise reduction adaptation across the entire temperature range; The specific steps of hazard feature pre-detection and protected area marking are as follows: Before noise reduction, a hazard area mask is generated based on color features, flicker frequency features, and local contrast features. M danger : Color characteristics: Pixels with a red channel ratio (R / (R+G+B)) greater than 0.6 and a brightness higher than the global average are marked as open flame candidates; Flicker frequency characteristics: Pixels whose brightness falls within the 8-15Hz range for 5 consecutive frames are marked as candidates for flame flicker; Local contrast characteristics: If the average brightness difference between the candidate region and its neighbors is greater than 30 gray levels, it is judged as a high-risk target; If all three characteristics mentioned above are met, it is considered a hazardous area mask. Mdanger And perform 3×3 morphological expansion on the mask; Temporal motion saliency detection and region segmentation specifically involve calculating the motion amplitude and fluctuation variance between the current frame and the previous frame. ; In the formula: Var() This represents variance calculation. k - N 5: k Indicates from k - N 5 to the k A sequence of consecutive frames; Preset motion threshold Tm With noise threshold Td ,like M(x,y) > Tm If so, it is determined to be a dynamic target area; if M(x,y) ≤ Tm and D(x,y) > Td If the pixel is 0, it is considered a candidate transient noise region; the remaining pixels are static background regions. For candidate transient noise regions, further verification is performed using the brightness fluctuation period of 5 consecutive frames. If the fluctuation period is within (0.067s, 0.125s), it is officially determined to be a transient noise region; otherwise, it is classified as a dynamic target region. The spatial domain region division is as follows: For the dynamic target region, for the current frame's luminance component Y, a 3×3 neighborhood window is used to calculate the local variance of each pixel. At the same time, the noise threshold is obtained. Th1 , Th2 : ; In the formula: This represents the standard deviation of Gaussian noise at room temperature. like If , then it is a flat subregion; if If , then it is a texture sub-region; if If so, then it is an edge sub-region; The specific steps of region-based noise reduction are as follows: For static background areas: using the current frame as the center, an exponentially decaying weight function is applied: ; In the formula: i Indicates the historical frame number. Indicates the time smoothing coefficient; The temporal filtering result is obtained by normalizing and weighting the 5-frame luminance values ​​for each pixel: ; The time-domain output is smoothed with a 3×3 Gaussian kernel to further suppress spatial residual noise and avoid blurring weak details in the background. For transient noise regions: for pixels ( x,y Extract continuous n The brightness values ​​of the frames form a one-dimensional time series: The sequence is sorted by grayscale value, and the median value is taken as the output. If it is a normal transient noise area, then n=5; if it is a dangerous area, then n=3. For dynamic target areas: If it is a flat sub-region, first calculate the side length of the adaptive filtering window. And obtain the neighboring pixel weights. , ( u,v ) represents the coordinates of neighboring pixels within the window, ( x 0 ,y 0) represents the center pixel coordinates of the window, u,v () represents the coordinates of neighboring pixels within the window. The minimum value is obtained; the final denoising result is: ; If it is a textured sub-region: the initial window size is 3×3, and the maximum window size is... , Indicates the reference Poisson noise intensity; First layer (noise determination): If Z min < Z med < Z max If the window size is positive, it indicates that there is no strong impulse noise within the window, and the process proceeds to the next layer; otherwise, the window size is increased, and the first layer judgment is repeated until the window reaches the specified size. W max ; Second layer (center pixel determination): If Z min < Z xy < Z max This indicates that the center pixel is not a noise point, and its original value is retained. Z xy Otherwise, if the center pixel is noise, replace it with... Z med ; If the first-level condition is still not met even when the window reaches its maximum size, the value in the window will be output directly. For edge sub-regions: use the sym4 wavelet basis to perform brightness mapping on the edge region. J Layer decomposition yields one low-frequency approximation coefficient and J The high-frequency detail coefficients of the group, and the threshold for the high-frequency coefficients of the j-th layer are: ; In the formula: N t Indicates the total number of pixels in the image; Perform soft threshold shrinkage on all high-frequency coefficients: ; In the formula: Represents the original high-frequency coefficients; The low-frequency coefficients are retained unchanged, and the processed high-frequency coefficients are combined with the sym4 wavelet inverse transform to reconstruct the image of the denoised edge region. The base strength of all the above filters is the depth-guided baseline denoising strength. Decide: ; In the formula: Indicates the baseline denoising intensity. k d This represents the depth-of-field gain coefficient. d max This represents the global maximum depth of field estimate; If a pixel is in a dangerous area, the denoising intensity is reduced to 30% of the baseline denoising intensity, and the filter window is simultaneously reduced and the threshold intensity is lowered.

[0011] Based on further optimization of the above scheme, step S3 uses the denoised luminance component Y and the original chrominance component... C b and C r Transmittance diagram t(x,y) Depth of field map d(x,y) Dirt mask M smudge Danger mask M danger Self-illuminating mask M emissive For input; The dual-region division of reflected light and self-emitting light is specifically as follows: the image is divided into two major categories based on a self-emitting light mask: a self-emitting light region and a reflected light region. Furthermore, based on transmittance, the reflected light region is divided into three sub-regions: high transmittance, medium transmittance, and low transmittance. Statistical analysis of the transmittance values ​​of effective pixels: ; The total number of pixels in the effective area is N valid Sort all effective transmittance in ascending order to obtain an ordered sequence. ; First, obtain the initial threshold for the partition: Then obtain the final threshold: ; like If so, it is a low-transmittance sub-region; if If , then it is the mid-transmitter region; if This is the high-transmittance sub-region.

[0012] Based on further optimization of the above scheme, in step S3, the construction of the nonlinear enhancement transformation model is specifically as follows: ; In the formula: I in (x,y) This represents the normalized input grayscale value. These represent the shape control parameters of the transformation function; Represents a fully Beta function: ; In the formula: Represents the Gamma function; Indicates an incomplete Beta function: ; The specific optimization of the partition parameters involves using an immune genetic algorithm to optimize the three sub-regions of the reflected light region. parameter: ; ; In the formula: These represent the corresponding weight coefficients; E pt Represents image information entropy. L Represents the number of gray levels. p i Indicates the first i The probability of a gray level appearing in an image. p i = n i / (M×N), n i Indicates grayscale value i The total number of pixels, M×N represents the total number of pixels in the image (i.e., width × height). C jb Indicates local contrast. Indicates the local standard deviation of a single pixel. This represents a local neighborhood window centered at (x, y). W n Indicates the side length of the neighborhood window. This represents the average grayscale value within the neighborhood window; G avg Represents the average gradient. I x (x,y) , I y (x,y) These represent the horizontal gradient and the vertical gradient, respectively. D yx Indicates the effective detail ratio, This represents the gradient magnitude at pixel (x, y). L over Indicates the percentage of overexposed pixels. n over Indicates the total number of overexposed pixels; The multi-factor pixel-wise enhancement intensity adjustment specifically involves: combining the depth-of-field guiding coefficient. Pollution compensation coefficient Dangerous area enhancement coefficient With self-luminescence enhancement intensity coefficient Pixel-by-pixel fine-tuning to enhance intensity: ; Depth of field guiding coefficient : ; In the formula: Indicates the reference reinforcement strength; Pollution compensation factor :like M smudge (x,y) =1, then Conversely, no pollution compensation coefficient is introduced; Dangerous area enhancement coefficient If it is a dangerous area, then Conversely, no enhancement coefficient for dangerous areas is introduced; Self-luminescence enhancement intensity coefficient :like M emissive (x,y) =1, then Conversely, no self-luminescence enhancement strength coefficient is introduced; Final realization Adjust parameters pixel by pixel: ; In the formula: These represent the baseline optimization parameters for the corresponding regions; The hard constraints are specifically applied by applying two hard constraints to the enhancement results: spatial brightness abrupt change constraint and inter-frame brightness fluctuation constraint, taking into account the visual comfort characteristics of the human eye. Spatial brightness abrupt change constraint: The local brightness gradient at the splicing boundary and strong edge is limited to no more than 20 gray levels / pixel. If it exceeds this, a gradient smoothing is performed. Inter-frame brightness fluctuation constraints: , This represents the global average brightness of the current k frames; if it exceeds this range, the brightness range is linearly compressed, i.e.: ; Indicates the average brightness of the target: ; Chromaticity drift regional compensation specifically involves: adaptively compensating for chromaticity components based on transmittance, taking into account the selective scattering characteristics of smoke wavelengths, and restoring true colors in different regions. For the area of ​​reflected light: C b The gain coefficient is 1+k cb (1- t(x,y) ), C r The gain coefficient is 1+k cr (1- t(x,y) ), k cb k cr They are the gain compensation coefficients and k cb >k cr To compensate for the decrease in color saturation caused by smoke scattering; For self-illuminating areas: Increase the weight of the red channel to restore the warm color tone of the flame, C r The component gain was adjusted to 0.8, C b The component gain was adjusted to 0.3; Color image reconstruction specifically involves mapping the enhanced luminance component back to the [0, 255] interval and matching it with the compensated chrominance component C. b C r The images are merged and then transformed inversely from YCbCr to RGB to obtain the enhanced color image.

[0013] Based on further optimization of the above scheme, step S4 uses the three-channel enhanced color image and depth map output in step S3. d(x,y) Self-illuminating mask M emissive For input; The fusion feature point extraction process is as follows: for each image path, extract Harris corner features and SIFT scale-invariant features respectively, and fuse them to form a feature point set; Feature point description and normalization include SIFT feature point descriptor generation, Harris corner point descriptor adaptation and descriptor normalization. After normalization, the vector magnitude is equal to 1. The specific depth level labeling of feature points is as follows: if the depth value of the location corresponding to the feature point coordinates is greater than 10m, it is labeled as a distant static background; if the depth value of the location corresponding to the feature point coordinates is greater than 3m and not greater than 10m, it is labeled as a medium-distance object; if the depth value of the location corresponding to the feature point coordinates is not greater than 3m, it is labeled as a close-range target. The static-dynamic feature classification is as follows: first, the static and dynamic characteristics of the features are determined, and then the static feature library is updated frame by frame. The hierarchical static prior coarse matching specifically involves: performing feature matching at different depth levels for adjacent image paths (left-center, center-right); and using the nearest neighbor / second nearest neighbor distance ratio (NNDR) method for coarse matching. ; In the formula: d frist Indicates the nearest neighbor distance. d second Indicates the distance to the next nearest neighbor; Preset distance ratio threshold T NNDR ,like NNDR < T NNDR If the match is positive, it is considered a valid coarse matching point pair. The Euclidean distance of the second most similar points; static feature matching weight is higher than dynamic feature matching (prioritizing the registration accuracy of stable structures). The fine matching process involves first solving the homography matrix according to the depth level. H far , H mid , H near Construct a homography optimization objective function with static prior weights: ; In the formula: H Represents the homography matrix. p s,i , q s,i These represent the first and second characters in the static feature library, respectively. i Homogeneous coordinates of each feature point, and homogeneous coordinates of the corresponding matching points in adjacent path images. p d,i , q d,i These represent the dynamic feature points extracted in the current frame, i.e., the i-th and j-th points respectively. i The homogeneous coordinates of each point, and the homogeneous coordinates of the corresponding matching points in the adjacent path images. n s , n d These represent the total number of static matching points and the total number of dynamic matching points, respectively. Indicates the static term weight coefficient; Additional binocular baseline parallax constraint is added to the near-field layer: ,in, d avg This represents the average depth of field for that layer. When the actual disparity obtained from the solution deviates from the theoretical disparity by more than 20%, the corresponding point pair is discarded. After solving the initial homography for each layer, the Random Sample Consensus Algorithm (RANSAC) is used for fine matching to eliminate mismatched point pairs that do not conform to the geometric transformation relationship, and the accurate homography matrix of each layer is output. H far , H mid , H near .

[0014] Based on further optimization of the above scheme, step S5 specifically includes: First, global bundle adjustment optimization is performed on the left-center and center-right hierarchical homography matrices. Using the center image as the reference coordinate system, the global consistency of the left and right transformations is constrained. The objective function for bundle adjustment is: ; In the formula: This indicates the homography of the l-th layer from left to middle. Indicates the homography of the lth layer from the middle to the right. This indicates a direct left-to-right homography match. Indicates the consistency constraint weight; Optimization goal: That is, to obtain all hierarchical homography matrices with the minimum total error; Then, based on the parallax consistency test, the pixels in the left image are mapped to the corresponding positions in the right image through homography, the depth values ​​of both sides are extracted, and the relative deviation is calculated: ; In the formula: for( x,y Map this to the corresponding coordinates in the right figure; Preset depth of field relative deviation threshold T ed ,like If the mapped location exceeds the effective range of the image, or if there are no effective depth-of-field / matching pixels at the corresponding location, it is marked as an occlusion area, and an occlusion mask is generated. M occlusion ; Next, inter-frame first-order low-pass filtering is performed on the homography matrices of each layer: ; In the formula: Indicates the firstk Homography after frame smoothing Indicates the first k Original frame homography, representing the smoothing coefficient; Then, using the central image as a reference, the left and right images are mapped to a unified panoramic coordinate system through the corresponding hierarchical homography matrix, thus completing the coordinate alignment of the three images; Next, for the non-occluded areas: based on the actual depth value of the pixel, calculate its membership weights in the far, middle, and near depth layers. And weighted fusion of the three-layer registered images: ; ; In the formula: Indicates the perspective decomposition threshold. Indicates the near-field decomposition threshold. k se This represents the steepness coefficient of the S-shaped function. These represent the images whose far, middle, and near views are mapped to the panoramic coordinate system using homography. For occluded areas: If it is unilateral occlusion (i.e., only one image stream can see the area), the pixels of the same depth on the non-occluded side are directly used as the output; if it is bilateral occlusion (i.e., neither image stream has a valid pixel at this location, such as the edge of the field of view or the foreground is completely occluded), first determine whether the location has been a static background in the past 3 frames. If so, the background pixels of the corresponding location in the previous frame are reused for completion. If it is a dynamic area, bilinear interpolation is performed with reference to the static background pixels of the same depth in the neighborhood. At the same time, for near-field obstacles, the original unilateral view image is retained and no stitching or fusion is performed. Finally, after completing the left-center-right three-way fusion, the invalid black borders around the edges are trimmed, and the seams are smoothed with a brightness gradient to output a standardized panoramic image. ; In the formula: x seam Indicates the x-coordinate of the seam. W gdd Indicates the width of the transition band. These represent the left and right fusion weights, respectively. Traverse the four boundaries of the panoramic image, shrink inward until all positions are valid pixels, and output a standardized 180° panoramic image with a resolution of 3200×680 after cropping.

[0015] The following are the technical effects of the present invention: This invention addresses the characteristics of open flames and high-temperature self-luminous bodies in fire scenes by establishing a dual-component physical model of "self-luminous and reflected light." It combines dark channel priors with binocular depth fusion to achieve accurate transmittance and depth-of-field estimation. A spatiotemporal regional denoising strategy is employed to dynamically estimate mixed noise parameters, while pre-marking hazardous areas and attenuating denoising intensity. This effectively suppresses complex noise from dense smoke and dust while avoiding blurred edges and filtering out weak hazard signals. Furthermore, based on transmittance, the invention divides smoke areas into light, medium, and dense smoke zones, using nonlinear Beta transformation for differentiated enhancement. It also incorporates multiple factors such as depth of field, lens contamination, hazardous areas, and self-luminescence to adjust enhancement intensity pixel-by-pixel, improving visibility of trapped personnel and obstacles in dark areas while suppressing overexposure in bright areas. Combined with chromatic aberration compensation and spatial / inter-frame brightness hard constraints, it restores color while avoiding visual fatigue and dizziness, making it suitable for extended wear. By fusing Harris and SIFT features, combined with deep hierarchical annotation, static and dynamic feature classification, and a static prior library, the problem of feature point failure and low registration accuracy in dense smoke and low texture environments is solved. Multi-plane hierarchical homography registration and depth-weighted fusion are adopted to overcome the limitations of the single-plane assumption. With the help of dynamic occlusion detection and inter-frame anti-shake smoothing, ghosting, stitching seams and image jitter are effectively eliminated, achieving 180° seamless panoramic stitching and meeting the needs of observation in environments without blind spots.

[0016] This invention employs a strategy of processing only the luminance component in the YCbCr color space to significantly reduce computational load. Combined with regional differentiated processing (lightweight filtering in flat areas and precise noise reduction in edge areas) and static feature library reuse, it controls algorithm complexity while ensuring performance and supports low-power real-time operation on embedded devices. The entire process of this invention can be completed locally on the helmet without relying on backend computation. It provides real-time near-eye visual assistance for front-end firefighters while reducing wireless transmission bandwidth pressure and transmission latency, effectively adapting to the high real-time and weak communication environment requirements of fire rescue. Attached Figure Description

[0017] Figure 1 This is a flowchart of the image denoising and stitching method in an embodiment of the present invention.

[0018] Figure 2 This is a hardware structure diagram of the image denoising and stitching method in an embodiment of the present invention.

[0019] Figure 3 This is the final output image of the image denoising and stitching method in this embodiment of the invention. Detailed Implementation

[0020] The technical solutions in the embodiments of the present invention will be clearly and completely described below. In the following description, specific details such as specific system structures and technologies are presented for illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present invention.

[0021] Example 1: A method for image denoising, enhancement, and stitching in fire scenes includes: Step S1, Image Acquisition and Preprocessing: Image acquisition is performed using three wide-angle high-definition cameras (left, center, and right) mounted on the front of the helmet. These cameras simultaneously acquire images of the fire scene at a frame rate of 10 frames per second, with an original image resolution of 720p (1280×720). The embedded ARM platform includes an ARM processor, DDR4 memory module, communication module, display module, and power management module. The ARM processor uses an ARM Cortex-A55 architecture and incorporates a junction temperature sensor and performance counter. The three wide-angle high-definition cameras are synchronously triggered via GPIO interfaces and through MIPI... The CSI / DVP interface connects to the ARM processor (to achieve synchronous image acquisition). The DDR4 memory module connects to the ARM processor via the memory bus (to provide main memory space for multi-frame image caching and algorithm operation). The communication module uses one or more combinations of WIFI, 4G / 5G modules and connects to the ARM processor via the USB interface (to complete the wireless transmission of encoded images). The display module connects to the ARM processor via the MIPI interface (to achieve real-time near-eye display of processed panoramic images, such as AR near-eye display modules). The power management module provides regulated power supply for all hardware units (supporting power consumption monitoring and dynamic voltage adjustment).

[0022] Image preprocessing is performed using an embedded ARM platform within the helmet, specifically including: First, the three acquired RGB raw images are converted to YC. b C r Color space, separating the luminance component Y and chrominance component C of an image. b C r (Only the luminance component Y is processed, while the chrominance component C is retained) b C r (Used for final image reconstruction, significantly reducing computational load) ; In the formula: R, G, and B represent the three-channel grayscale values ​​of the original RGB image (range [0, 255]); the Y component ranges from [16, 235]; C b C r The component values ​​range [16, 240]; Then, the lens distortion invalid areas at the edges of each image are cropped, retaining the effective imaging area of ​​1200×680 in the center; at the same time, a dedicated buffer area is opened to store the luminance component Y data and timing mark of 5 consecutive frames (to provide data support for subsequent spatiotemporal denoising and frame rate transformation; the chrominance component is only cached for the current frame to save memory usage). Subsequently, based on local sharpness differences, areas of contamination such as smoke and water droplets on the lens surface are detected, and a contamination mask is generated. M smudge (To provide a basis for subsequent enhanced compensation): ; In the formula: Clear() Indicates the local sharpness value. Laplacian() Represents the Laplace operator. gray(x,y) Representing coordinates (x,y) The single-channel grayscale value at that location (corresponding to the luminance component Y); When the local sharpness is less than 50% of the global average sharpness and the area of ​​a continuous region is greater than 30 pixels, it is marked as a dirty region, i.e., the mask value is 1; the rest are normal regions, i.e., the mask value is 0; the mask is updated every 30 frames (to avoid false detection caused by occasional noise).

[0023] Step S2, Adaptive denoising processing: Using the current frame's luminance component Y and original chrominance component C output in step S1... b With C r Contamination mask M smudge For input; specifically including: The separation of reflected light and spontaneous emission components and transmittance estimation are specifically addressed by: To resolve the issue of traditional scattering models failing due to spontaneous emission sources in fire scenes, the imaging is decomposed into a spontaneous emission component and a scene reflected light component, establishing a dual-component imaging physical model. ; In the formula: This represents the self-luminous component (i.e., the light emitted by an open flame or a high-temperature object). Indicates the components of light reflected from the scene. t(x,y) This represents the smoke transmittance (typically (0,1)). A f Indicates the global ambient light intensity; Then, a rough detection of the self-luminous area was performed (fire / high-temperature self-luminous bodies are mainly red-orange in tone, with the red chromaticity component C). r Significantly high, blue chromaticity component C b (Too low): First calculate the percentage of redness: ; Preset percentage threshold T cr(Usually 0.65), if R ratio > T cr If the value is 0, it is marked as a chroma candidate pixel; otherwise, it is not marked. Next, calculate the maximum brightness within a 15×15 neighborhood window. If the current pixel brightness is equal to the local maximum and is more than 1.5 times the global average brightness, then mark it as a brightness candidate pixel. If a pixel is both a chroma candidate pixel and a luminance candidate pixel, it is marked as a self-illuminating candidate; and a 3×3 morphological dilation is performed on the candidate region (to fill the voids inside the target) to obtain a coarse mask. M emissive-c ; Then, for a consecutive 5-frame luminance sequence, the temporal mean and temporal variance of each pixel are calculated: ; And the pixel fluctuation frequency is calculated using the luminance zero-crossing rate: ; In the formula: This indicates the frame interval time (usually 0.1s). sign() Indicates the sign function (takes 1 when the value is positive, and -1 when it is negative); like Z(x,y) If the frequency is ∈[8Hz,15Hz], then it is determined to have flame flicker characteristics; If the coarse mask simultaneously satisfies the time-domain variance > 20 and Z(x,y) If the Hz range is ∈[8Hz, 15Hz], it is retained as a self-emissive region; static high-temperature self-emissive bodies (such as red-hot metal) have no flicker characteristics, and their chromaticity and brightness are directly retained to obtain the final self-emissive mask. M emissive ; Next, estimate the transmittance of non-emissive regions: For the luminance component of non-emissive regions, calculate the minimum luminance value within the local window (i.e., the dark channel value; self-emissive regions are not included in the calculation and are filled by interpolation using neighborhood values): ; In the formula: Indicates ( x,y A local window centered on ) (size 15×15); Take the top 0.1% of pixels in the dark channel brightness and calculate the average of their corresponding original brightness as the global ambient light intensity. A f ; The initial transmittance is: , This indicates the smoke retention factor (typically 0.95). Using the original luminance map as a guide map, guided filtering optimization is performed on the initial transmittance (smoothing the transmittance map while preserving depth edges and eliminating block effects): ; In the formula: The linear coefficients within the window are represented (calculated using the covariance and mean of the guide plot and the input plot). W k This represents the k-th filtering window; Perform time-domain weighted smoothing of the transmittance of the current frame and historical frames: ; In the formula: This represents the smoothing coefficient (usually 0.7). This represents the optimized transmittance of the current frame; The scattering model is not applicable to the self-luminous region, and the transmittance is directly set to 1 (to avoid model distortion). For non-self-emissive areas (i.e.) M emissive =0), directly inversely solve for the reflected light component: ; For self-emissive areas (i.e.) M emissive =1): ; Finally, according to the Beer-Lambert law, the model depth of field is obtained through inversion: ; In the formula: This represents the smoke scattering coefficient (an initial value is preset by the calibrated scene and subsequently adjusted adaptively according to the environment). For near-field areas within 3 meters, depth is calculated using the binocular parallax of the left and right cameras (to correct for errors in the scattering model): ; In the formula: b d This indicates the binocular baseline distance (i.e., the distance between the left and right cameras on the helmet, typically 0.06m). f d Indicates the camera's pixel focal length. This represents the disparity between corresponding matching pixels in the left and right images; Near-field area with binocular depth For reference, the mid-to-long-range views are based on the model's depth of field. To determine the optimal depth-of-field map, an S-shaped weighted smooth transition is applied to the transition area, resulting in the final depth-of-field map. d(x,y) .

[0024] The mixed noise parameter estimation specifically involves fitting a Poisson-Gaussian mixed noise model using maximum likelihood estimation and dynamically updating the noise parameters. ; In the formula: Poisson() Indicates the Poisson distribution. z Indicates the grayscale value of the observed pixel. Indicates Poisson noise intensity; Gauss() Indicates a Gaussian distribution. This represents the mean of Gaussian noise. Indicates the standard deviation of Gaussian noise; This represents the convolution operation; Noise parameters are updated once per frame to achieve noise reduction adaptation across the entire temperature range.

[0025] Hazard feature pre-detection and protected area marking specifically involves generating a hazardous area mask based on color features, flicker frequency features, and local contrast features before noise reduction processing. M danger (To avoid smoothing out weak danger signals during subsequent noise reduction): Color characteristics: Pixels with a red channel ratio (R / (R+G+B)) greater than 0.6 and a brightness higher than the global average are marked as open flame candidates; Flicker frequency characteristics: Pixels whose brightness falls within the 8-15Hz range for 5 consecutive frames are marked as candidates for flame flicker; Local contrast characteristics: If the average brightness difference between the candidate region and its neighbors is greater than 30 gray levels, it is judged as a high-risk target; If all three characteristics mentioned above are met, it is considered a hazardous area mask. M danger And perform 3×3 morphological expansion on the mask (to completely preserve the target area).

[0026] Temporal motion saliency detection and region segmentation specifically involve calculating the motion amplitude and fluctuation variance between the current frame and the previous frame. ; In the formula: Var() This represents variance calculation. k - N 5: k Indicates from k - N 5 to the k A sequence of consecutive frames; Preset motion threshold Tm (Typically 20) and noise threshold Td (Usually 15), if M(x,y) > Tm If so, it is determined to be a dynamic target area (continuous movement, corresponding to real targets such as moving people or objects); ifM(x,y) ≤ Tm and D(x,y) > Td If the pixel is 0, it is considered a candidate transient noise region; the remaining pixels are static background regions. For candidate transient noise regions, further verification is performed through the brightness fluctuation period of 5 consecutive frames. If the fluctuation period is within (0.067s, 0.125s] (corresponding to 8-15Hz flame flicker), it is officially determined to be a transient noise region; otherwise, it is classified as a dynamic target region.

[0027] Spatial domain region division is as follows: For dynamic target regions, the local variance of each pixel is calculated using a 3×3 neighborhood window for the current frame's luminance component Y. At the same time, the noise threshold is obtained. Th1 , Th2 : ; In the formula: This represents the standard deviation of Gaussian noise at room temperature (typically 10). like If , then it is a flat subregion; if If , then it is a texture sub-region; if Then it is an edge sub-region.

[0028] The noise reduction process is divided into regions, specifically: For static background areas: using the current frame as the center (the farther the historical frame is from the current frame, the lower its weight), an exponentially decaying weight function is used (to ensure that recent frames contribute more): ; In the formula: i Indicates the historical frame number (a total of 5 frames are involved in the calculation, i.e., i∈{ k -4, k -3, k -2, k -1, k}), This represents the time smoothing coefficient (usually 2). The temporal filtering result is obtained by normalizing and weighting the 5-frame luminance values ​​for each pixel: ; The time-domain output is smoothed with a 3×3 Gaussian kernel (with a Gaussian kernel standard deviation of 0.8) to further suppress spatial residual noise and avoid blurring weak details in the background. For transient noise regions: for pixels ( x,y Extract continuous n The brightness values ​​of the frames form a one-dimensional time series: The sequence is sorted by grayscale value, and the median value is taken as the output. If it is a normal transient noise area, then n=5; if it is a dangerous area, then n=3. For dynamic target areas: If it is a flat sub-region, first calculate the side length of the adaptive filtering window. And obtain the neighboring pixel weights. , ( u,v ) represents the coordinates of neighboring pixels within the window, ( x 0 ,y 0) represents the center pixel coordinates of the window, u,v () represents the coordinates of neighboring pixels within the window. It is the minimum value (generally 1). 10 -6 The final denoising result is as follows: ; If it is a textured sub-region: the initial window size is 3×3, and the maximum window size is... , Indicates the reference Poisson noise intensity; First layer (noise determination): If Z min (Minimum grayscale value) Z med (Window grayscale median) Z max (Maximum grayscale value) indicates that there is no strong impulse noise within the window, proceed to the next layer; otherwise, increase the window size (increase the side length by 2 each time), repeat the first layer judgment, until the window reaches the maximum grayscale value. W max ; Second layer (center pixel determination): If Z min < Z xy (Center pixel value) Z max This indicates that the center pixel is not a noise point, and its original value is retained. Z xy Otherwise, if the center pixel is noise, replace it with... Z med ; If the first-level condition is still not met even when the window reaches its maximum size, the value in the window will be output directly. For edge sub-regions: use the sym4 wavelet basis to perform brightness mapping on the edge region. J Layer (usually 2 layers) decomposition yields one low-frequency approximation coefficient and J The high-frequency detail coefficients of the group (horizontal, vertical, and diagonal directions), and the threshold for the high-frequency coefficients of the j-th layer are: ; In the formula: Nt Indicates the total number of pixels in the image; Perform soft threshold shrinkage on all high-frequency coefficients: ; In the formula: Represents the original high-frequency coefficients; The low-frequency coefficients are retained unchanged, and the processed high-frequency coefficients are combined with the sym4 wavelet inverse transform to reconstruct the image of the denoised edge region. The base strength of all the above filters is the depth-guided baseline denoising strength. The factors that determine (the smoothing coefficient of time-domain filtering, the maximum window size of adaptive median filtering, and the threshold coefficient of wavelet denoising) are all related to... Positive correlation; the greater the depth of field, the stronger the noise caused by smoke scattering, and the stronger the basic noise reduction intensity is simultaneously. ; In the formula: This represents the baseline denoising strength (typically 1). k d This represents the depth-of-field gain factor (typically 0.5). d max This represents the global maximum depth of field estimate; If a pixel is in a dangerous area, the denoising intensity is reduced to 30% of the baseline denoising intensity, and the filter window is reduced and the threshold intensity is lowered simultaneously (to ensure that dangerous signals are not over-smoothed).

[0029] Step S3, Adaptive Enhancement and Smudge Compensation: Using the denoised luminance component Y and the original chrominance component... C b and C r Transmittance diagram t(x,y) Depth of field map d(x,y) Dirt mask M smudge Danger mask M danger Self-illuminating mask M emissive For input; specifically including: The dual-region division of reflected light and self-emitting light is as follows: Based on the self-emitting light mask, the image is divided into two major categories: self-emitting light region (actively emitting light region such as open flame and high temperature object) and reflected light region (scene reflection imaging region); Furthermore, based on transmittance, the reflected light region is divided into three sub-regions: high transmittance (light smoke), medium transmittance (medium smoke), and low transmittance (dense smoke). Statistical analysis of the transmittance values ​​of effective pixels: ; The total number of pixels in the effective area isN valid Sort all effective transmittance in ascending order to obtain an ordered sequence. ; First, obtain the initial threshold for the partition: Then obtain the final threshold: ; like If so, it is a low-transmittance sub-region; if If , then it is the mid-transmitter region; if This is the high-transmittance sub-region.

[0030] The nonlinear enhancement transformation model is constructed as follows: ; In the formula: I in (x,y) This represents the normalized input grayscale value (usually [0,1]). These represent the shape control parameters of the transformation function; Represents a fully Beta function: ; In the formula: Represents the Gamma function; Indicates an incomplete Beta function: .

[0031] The optimization of partitioning parameters involves using an immune genetic algorithm to optimize the three sub-regions of the reflected light region separately. parameter: ; ; In the formula: These represent the corresponding weighting coefficients (generally) ); E pt Represents image information entropy. L This indicates the number of gray levels (typically 256 in an 8-bit grayscale image). p i Indicates the first i The probability of a gray level appearing in an image. p i = n i / (M×N), n i Indicates grayscale value i The total number of pixels, M×N represents the total number of pixels in the image (i.e., width × height).C jb Indicates local contrast. Indicates the local standard deviation of a single pixel. This represents a local neighborhood window centered at (x,y) (typically a 3×3 window). W n Indicates the side length of the neighborhood window. This represents the average grayscale value within the neighborhood window; G avg Represents the average gradient. I x (x,y) , I y (x,y) These represent the horizontal gradient and the vertical gradient (obtained by convolving the image with the Sobel operator), respectively. D yx Indicates the effective detail ratio, This represents the gradient magnitude at pixel (x,y) (i.e. ); L over Indicates the percentage of overexposed pixels. n over This indicates the total number of overexposed pixels.

[0032] Multi-factor pixel-wise enhancement intensity adjustment, specifically: combining depth-of-field guiding coefficients. Pollution compensation coefficient Dangerous area enhancement coefficient With self-luminescence enhancement intensity coefficient (Non-staining area) Non-dangerous areas Non-self-illuminating areas ), Pixel-by-pixel fine-tuning to enhance intensity: ; Depth of field guiding coefficient : ; In the formula: Indicates the baseline reinforcement strength (typically 1); Pollution compensation factor :like M smudge (x,y) =1, then Conversely, no pollution compensation coefficient is introduced; Dangerous area enhancement coefficient If it is a dangerous area, then Conversely, no enhancement coefficient for dangerous areas is introduced; Self-luminescence enhancement intensity coefficient :likeM emissive (x,y) =1, then Conversely, no self-luminescence enhancement strength coefficient is introduced; Final realization Adjust parameters pixel by pixel: ; In the formula: These represent the baseline optimization parameters for the corresponding regions.

[0033] The hard constraints are applied specifically as follows: taking into account the visual comfort characteristics of the human eye, two hard constraints are applied to the enhancement results: spatial brightness abrupt change constraint and inter-frame brightness fluctuation constraint (to avoid visual fatigue and dizziness caused by prolonged wear): Spatial brightness abrupt change constraint: The local brightness gradient at the splicing boundary and strong edge is limited to no more than 20 gray levels / pixel. If it exceeds this, a gradient smoothing is performed. Inter-frame brightness fluctuation constraints: , This represents the global average brightness of the current k frames; if it exceeds this range, the brightness range is linearly compressed (to eliminate flicker), i.e.: ; Indicates the average brightness of the target: .

[0034] Chromaticity drift regional compensation specifically involves: adaptively compensating for chromaticity components based on transmittance, taking into account the selective scattering characteristics of smoke wavelengths, and restoring true colors in different regions. For areas of reflected light (lower transmittance, denser smoke, and more significant chromatic aberration): C b The gain coefficient is 1+k cb (1- t(x,y) ), C r The gain coefficient is 1+k cr (1- t(x,y) ), k cb k cr They are the gain compensation coefficients and k cb >k cr To compensate for the decrease in color saturation caused by smoke scattering; For self-illuminating areas: Increase the weight of the red channel to restore the warm color tone of the flame, C r The component gain was adjusted to 0.8, C b The component gain was adjusted to 0.3 (to suppress color cast while preserving the visual characteristics of the flame).

[0035] Color image reconstruction specifically involves mapping the enhanced luminance component back to the [0,255] interval and matching it with the compensated chrominance component C. b Cr The images are merged and then transformed inversely from YCbCr to RGB to obtain the enhanced color image.

[0036] Step S4, Image Feature Point Extraction and Matching: Using the three-channel enhanced color image and depth map output in step S3... d (x,y) Self-illuminating mask M emissive For input; specifically including: The feature point extraction process involves extracting Harris corner features and SIFT scale-invariant features for each image path and fusing them to form a feature point set. Harris corner feature extraction: First, calculate the grayscale gradient of the input grayscale image in the horizontal and vertical directions. I x , I y And obtain the product term of the gradient. Then, construct the second-order moment matrix for each pixel. Mix(x,y) : ; In the formula: Indicates the Gaussian smoothing kernel (general) (1.5) Finally, calculate the corner response function R: ; In the formula: det(Mix) Represents the determinant of a matrix. tr(Mix) Represents the trace of a matrix. k R This represents an empirical coefficient (typically 0.04 to 0.06). Pixels with R>0 correspond to corner points, pixels with R=0 correspond to flat areas, and pixels with R<0 correspond to edges. SIFT Scale Invariant Feature Extraction: First, perform different steps on the original image... Gaussian blurring, layer-by-layer downsampling, forming multiple octaves, each with multiple layers, creates a pyramid structure: ; Subtracting Gaussian images of adjacent scales within the same group pixel by pixel yields the DoG response map (used for extremum detection): ; In the formula: k L This represents the scale factor (usually 2). 1 / s (where s is the number of layers in each group). Local extrema points are selected within the 3×3×3 neighborhood of the DoG pyramid (3×3 neighborhood at the same level + 3×3 neighborhood at each adjacent scale above and below, for a total of 27 pixels) as candidate feature points; the sub-pixel positions and scales of the feature points are refined through Taylor expansion, low-contrast points and edge response points are removed, and the final stable SIFT feature points are output. For falling on a self-illuminating mask M emissive All feature points within the =1 region have their weights reduced in subsequent matching and homography solving stages (to avoid registration drift caused by dynamic flame textures). Feature point description and normalization are as follows: SIFT feature point descriptor generation: Calculate the gradient direction histogram of all pixels in the neighborhood of the feature point, take the peak of the histogram as the main direction of the feature point to achieve rotation invariance; rotate the neighborhood coordinate system to the main direction and divide it into 4×4 sub-regions. Calculate the gradient weighted histogram of 8 directions in each sub-region to finally form a 128-dimensional feature vector. Harris corner descriptor adaptation: The baseline Gaussian layer corresponding to the current image resolution is used as a fixed scale; at this baseline scale, the gradient direction histogram of the corner neighborhood is calculated, and the peak value is taken as the main direction; a 128-dimensional descriptor is generated according to the SIFT standard 4×4 sub-region method to achieve a unified description space with SIFT features (similarity matching can be performed directly). Descriptor normalization: Assuming 128-dimensional descriptor vectors v =[ v 1 ,v 2 ,…,v 128 ], then the normalized vector is: ; After normalization, the vector magnitude is equal to 1.

[0037] Feature point depth-level labeling is as follows: If the depth value of the location corresponding to the feature point coordinates is greater than 10m, it is labeled as a distant static background (such as building walls, ceilings, etc., with minimal parallax and stable structure); if the depth value of the location corresponding to the feature point coordinates is greater than 3m and not greater than 10m, it is labeled as a medium-distance object (such as tables, chairs, door frames, equipment, etc.); if the depth value of the location corresponding to the feature point coordinates is not greater than 3m, it is labeled as a close-range target (such as nearby obstacles, trapped people, etc., with large parallax and prone to splicing misalignment). Static-dynamic feature classification, specifically: First, determine the static and dynamic characteristics of the features: for each non-self-illuminating region of the current frame... f c Within the same depth level of the static feature library (initially the historical static feature library), retrieve the feature point with the smallest Euclidean distance from the descriptor.f s : ; In the formula: d c,i Represents the first feature point. i dimensional descriptor vector, d s,i Represents the first candidate feature point in the static feature library. i 3D descriptor vector; And calculate the current feature point ( x c ,y c ) and candidate matching points ( x s ,y s Pixel position deviation between (verifying static position stability): ; Preset similarity threshold T desc (Typically 0.25–0.3) and the position deviation threshold T pos (Usually 2 pixels), if d desc < T desc and d pos < T pos If a feature is marked as a static feature, the number of consecutive stable frames for the corresponding feature point in the static library is incremented by 1. If no condition is met, the feature is marked as a dynamic feature, which only participates in the matching calculation of the current frame and is not included in the statistics and updates of the static library. Then perform frame-by-frame sliding updates of the static feature library: Initial Stage: Extract feature points from all non-self-illuminating regions in frame 1, label the depth level, and use them as the initial candidate set. The number of consecutive stable frames for all candidate points is initialized to 1. For each feature point in frames 2 to 5, perform inter-frame matching at the same depth level with the candidate set of the previous frame: if the match is successful and the inter-frame pixel position deviation is less than 1 pixel, the number of consecutive stable frames for that point is incremented by 1; if the match fails or the inter-frame pixel position deviation is not less than 1 pixel, the number of consecutive stable frames for that point is reset to 1, and it is used as a new candidate point. Select feature points that meet the following conditions from the candidate set: if the number of consecutive stable frames is not less than 5 (i.e., it exists stably for 5 consecutive frames) and the inter-frame pixel position deviation is less than 1 pixel, it is officially included in the initial static feature library. The library inclusion constraint rule is: the total size of the library for each image is controlled within 200 points, and the priority is: far-field features > mid-field features > near-field features. When the capacity exceeds the limit, features are removed from low to high priority. Operational Phase: For new feature points in the current frame that are not matched in the static library, they are first added to the candidate tracking queue, and the number of consecutive stable frames is counted. When all of the following conditions are met, they are officially added to the static feature library: a) Stable matching in 3 consecutive frames: inter-frame positional deviation is less than 2 pixels and descriptor similarity meets the standard; b) The feature point is located in a non-self-illuminating area (excluding dynamic targets such as flames); c) The library capacity of the depth level has not reached the upper limit. When adding data, initialize the number of consecutive stable frames for that point to 3, and synchronously record its reference coordinates, reference descriptor and depth level. Feature points already existing in the static library will be removed from the library immediately if any of the following conditions occur: a) No corresponding feature point is matched in the current frame for two consecutive frames; b) The coordinates of the feature point exceed the effective imaging area of ​​the current image (e.g., 1200×680); c) The feature point falls into a self-illuminating area (obscured by flames or thick smoke, losing its static reference value). To ensure matching efficiency on the embedded end, the total size of the static image library for each channel is fixed at 200 points. When the capacity exceeds the limit and cropping is triggered, feature points are retained in the following order of priority from high to low: First priority: static features of the distant layer (the scene structure is the most stable and the registration reference value is the highest); Second priority: feature points with a larger number of consecutive stable frames within the same layer (the longer the consecutive stability time, the higher the confidence). Pruning order: Prioritize removing feature points with the smallest consecutive stable frame count and the closest depth to achieve dynamic balance of library size.

[0038] The hierarchical static prior coarse matching is as follows: for two adjacent images (left-center, center-right), feature matching is performed according to the depth level; the nearest neighbor / second nearest neighbor distance ratio (NNDR) method is used for coarse matching. ; In the formula: d frist It represents the nearest neighbor distance (i.e., the Euclidean distance between the points in the graph to be matched that are most similar to the current feature point descriptor). d second This represents the second nearest neighbor distance (the Euclidean distance between the points with the second highest similarity). Preset distance ratio threshold T NNDR (Usually 0.6), if NNDR < T NNDR If the match is positive, it is considered a valid coarse matching point pair. The Euclidean distance of the second most similar points is used, and the matching weight of static features is higher than that of dynamic features (prioritizing the registration accuracy of stable structures).

[0039] Fine matching specifically involves: first, solving the homography matrix according to the depth level. Hfar , H mid , H near Construct a homography optimization objective function with static prior weights: ; In the formula: H This represents the homography matrix (i.e., the 3×3 homogeneous projection transformation matrix). p s,i , q s,i These represent the first and second characters in the static feature library, respectively. i Homogeneous coordinates of each feature point, and homogeneous coordinates of the corresponding matching points in adjacent path images. p d,i , q d,i These represent the dynamic feature points extracted in the current frame, i.e., the i-th and j-th points respectively. i The homogeneous coordinates of each point, and the homogeneous coordinates of the corresponding matching points in the adjacent path images. n s , n d These represent the total number of static matching points and the total number of dynamic matching points, respectively. This represents the static term weight coefficient (typically 0.3). For the near-field layer, additional binocular baseline disparity constraints are added (disparity constraints are added to the left-center and center-right groups respectively: the left-center group uses the distance between the left and center cameras as the baseline, and the center-right group uses the distance between the center and right cameras as the baseline, and the theoretical disparity is calculated and verified respectively): ,in, d avg This represents the average depth of field for that layer. When the actual disparity obtained from the solution deviates from the theoretical disparity by more than 20%, the corresponding point pair is discarded. After solving the initial homography for each layer, the Random Sample Consensus Algorithm (RANSAC) is used for fine matching to eliminate mismatched point pairs that do not conform to the geometric transformation relationship, and the accurate homography matrix of each layer is output. H far , H mid , H near .

[0040] Step S5, Multi-channel Image Stitching: Employing multi-planar layered registration, dynamic occlusion detection, and layered weighted fusion mechanisms, seamless stitching of panoramic fire scene images is achieved (completely eliminating ghosting and stitching seams); specifically: First, global bundle adjustment optimization is performed on the left-middle and middle-right layered homography matrices. Using the middle image as the reference coordinate system, the global consistency of the transformations in the left and right paths is constrained (to avoid overall deformation and interlayer misalignment after stitching). The objective function for bundle adjustment is: ; In the formula: This indicates the homography of the l-th layer from left to middle. Indicates the homography of the lth layer from the middle to the right. This indicates a direct left-to-right homography match. This represents the consistency constraint weight (typically 0.5). Optimization goal: That is, to obtain all hierarchical homography matrices with the minimum total error; Then, based on the parallax consistency test, the pixels in the left image are mapped to the corresponding positions in the right image through homography, the depth values ​​of both sides are extracted, and the relative deviation is calculated: ; In the formula: for( x,y Map this to the corresponding coordinates in the right figure; Preset depth of field relative deviation threshold T ed (Usually 0.2), if If the mapped location exceeds the effective range of the image, or if there are no effective depth-of-field / matching pixels at the corresponding location, it is marked as an occluded area (i.e., M occlusion (x,y) =1), generate an occlusion mask. M occlusion ; Next, inter-frame first-order low-pass filtering is applied to the homography matrices of each layer (to smooth out image jitter caused by firefighters walking and turning their heads): ; In the formula: Indicates the first k Homography after frame smoothing Indicates the first k The original homography of the frame is represented by the smoothing coefficient (typically 0.6). Then, using the central image as a reference, the left and right images are mapped to a unified panoramic coordinate system through the corresponding hierarchical homography matrix, thus completing the coordinate alignment of the three images; Next, for the non-occluded areas: based on the actual depth value of the pixel, calculate its membership weights in the far, middle, and near depth layers. The three-layer registered images are then weighted and fused (completely eliminating boundary abrupt changes and stitching marks caused by hard layering): ; ; In the formula: This indicates the threshold for distant scene decomposition (typically 10m). This indicates the near-field resolution threshold (typically 3m). k se This represents the steepness coefficient of the S-shaped function (typically 1.5). These represent the images whose far, middle, and near views are mapped to the panoramic coordinate system using homography. For occluded areas: If it is unilateral occlusion (i.e., only one image stream can see the area), the pixels of the same depth on the non-occluded side are directly used as the output (preserving the original sharpness); if it is bilateral occlusion (i.e., neither image stream has a valid pixel at this location, such as the edge of the field of view or the foreground is completely occluded), first determine whether the location has been a static background in the past 3 frames. If so, the background pixels of the corresponding location in the previous frame are reused for completion (avoiding black blocks). If it is a dynamic area, bilinear interpolation is performed with reference to the static background pixels of the same depth in the neighborhood. At the same time, for near-field obstacles (people, equipment, etc., i.e., depth of field no greater than 3m), the original unilateral view image is preserved, and stitching and fusion are not performed (completely eliminating ghosting). Finally, after completing the left-center-right three-channel fusion, the invalid black borders around the edges are cropped, and the seams are smoothed with a brightness gradient (eliminating visible stitching marks caused by brightness differences between the two images), outputting a standardized panoramic image: ; In the formula: x seam Indicates the x-coordinate of the seam. W gdd This indicates the width of the transition band (usually 30 pixels). These represent the left and right fusion weights, respectively. Traverse the four boundaries of the panoramic image, shrink inward until all positions are valid pixels (not black borders), and output a standardized 180° panoramic image with a resolution of 3200×680 after cropping.

[0041] Example 2: As another preferred embodiment of the present invention, based on the above-described embodiment 1, in order to further reduce the consumption of computing resources, a four-dimensional dynamic resource assessment is also performed in step S1, specifically including: Construct a four-dimensional evaluation model of "temperature-computing power-memory-power consumption" to obtain the global resource reserve coefficient. ; in, Represents the remaining computing power: ; In the formula:U v This indicates the real-time CPU utilization rate (the percentage of computing power currently used by the processor, which is directly collected by the system performance counter and is generally [0,1]). f(T j ) This represents the temperature-induced frequency reduction factor function (describing the frequency reduction loss caused by the increase in processor junction temperature; it is generally [0,1] and equal to 1 at room temperature, obtained through experimental calibration); Indicates the remaining memory: ; In the formula: B u This represents the real-time memory bandwidth utilization rate (i.e., the proportion of memory bandwidth currently used by the system to the total bandwidth, collected by the memory performance counter, typically [0,1]). Indicates power consumption margin: ; ; In the formula: P rated Indicates the system's rated power consumption; P est This indicates that the system estimates the total power consumption in real time. P cpu This represents the CPU's full-load reference power consumption (i.e., the dynamic power consumption reference value of the CPU at 100% utilization and room temperature, obtained from the device datasheet or actual measurement calibration). P mem This indicates the reference power consumption at full bandwidth (i.e., the dynamic power consumption reference value of the memory at 100% bandwidth utilization and room temperature, which is calibrated by the device datasheet or actual measurement). P leak0 This represents the room temperature reference leakage current power consumption (which is an inherent static loss of CMOS devices). k P This indicates the temperature coefficient of leakage current (typically 0.01–0.02℃). -1 ), T j This indicates the processor junction temperature (i.e., the operating temperature of the chip core). T 0 indicates the reference room temperature (typically 25°C).

[0042] In step S2, dynamic pruning and adaptation of the corresponding computing power are performed: if This indicates sufficient resources, employing wavelet 3-level decomposition, 7×7 maximum median window, 5-frame temporal filtering, and full-frame danger detection; if This indicates that the resources are generally adequate. A two-level wavelet decomposition, a 5×5 maximum median window, a 3-frame temporal filtering, and downsampling risk detection are employed. This indicates insufficient resources. Wavelet denoising is canceled, and the edges are replaced with 3×3 edge protection mean filtering, 2-frame time-domain averaging, and danger fine detection is turned off.

[0043] In step S3, dynamic pruning and adaptation of the corresponding computing power are performed: if It employs independent optimization of the three reflection light areas, pixel-by-pixel intensity adjustment, perfect human factor constraint, and fine color compensation; if It employs a two-zone optimization approach, regional intensity adjustment, core human factor constraints, and basic colorimetric compensation; if Fixed enhancement parameters, canceled optimization, turned off human factors fine constraints, and only basic chromaticity compensation.

[0044] In step S4, dynamic pruning and adaptation of the corresponding computing power are performed: if The method employs Harris+SIFT full feature extraction, three-layer independent solution, RANSAC iteration 100 times, and a complete static feature library; if Only SIFT features are extracted, a two-layer solution is used, RANSAC is iterated 50 times, and the static feature library is simplified; if The homography matrix of the previous frame is reused, and only 3 core static feature points are extracted for verification and fine-tuning.

[0045] In step S5, dynamic pruning and adaptation of the corresponding computing power are performed: if It employs three-layer fine fusion, occlusion detection and filling, inter-frame smoothing, and edge brightness smoothing; if It employs two-layer fusion, simplified occlusion detection, and basic inter-frame smoothing; if It employs single-plane global homography, no layering, and occlusion detection disabled.

Claims

1. A method for image denoising, enhancement, and stitching in fire scenes, characterized in that: include: Step S1, Image Acquisition and Preprocessing: Images are acquired using three wide-angle high-definition cameras mounted on the left, center, and right sides of the helmet's front end, and preprocessed using an embedded ARM platform within the helmet. Step S2, Adaptive Denoising Processing: This includes separation of reflected light and self-emission dual components and transmittance estimation, mixed noise parameter estimation, hazard feature pre-detection and protected area marking, temporal motion saliency detection and region division, spatial region division, and regional denoising processing. Step S3, Adaptive Enhancement and Stain Compensation: This includes dual-region partitioning of reflected light and self-emission, construction of nonlinear enhancement transformation model, optimization of partition parameters, multi-factor pixel-by-pixel enhancement intensity adjustment, application of hard constraints, chromaticity drift partition compensation, and color image reconstruction. Step S4, Image Feature Point Extraction and Matching: This includes fusion feature point extraction, feature point description and normalization, feature point depth-level annotation, static-dynamic feature classification, hierarchical static prior coarse matching, and fine matching; Step S5, Multi-channel image stitching: Using multi-plane layered registration, dynamic occlusion detection, and layered weighted fusion mechanism, seamless stitching of panoramic fire scene images is achieved.

2. The image denoising enhancement and stitching method for fire scenes according to claim 1, characterized in that: In step S1, the left, center, and right wide-angle high-definition cameras synchronously acquire images of the fire scene environment at a frame rate of 10 frames per second, with an original image resolution of 720p. The embedded ARM platform includes an ARM processor, a DDR4 memory module, a communication module, a display module, and a power management module. The ARM processor is an embedded processor based on the ARM Cortex-A55 architecture, with a built-in junction temperature sensor and performance counter. The left, center, and right wide-angle high-definition cameras are synchronously triggered through the GPIO interface and connected to the ARM processor through the MIPI interface. The DDR4 memory module is connected to the ARM processor through the memory bus. The communication module uses one or more combinations of WIFI, 4G / 5G modules and is connected to the ARM processor through the USB interface. The display module is connected to the ARM processor through the MIPI interface. The power management module provides regulated power to all hardware units.

3. The image denoising enhancement and stitching method for fire scenes according to claim 1 or 2, characterized in that: In step S1, image preprocessing using the embedded ARM platform within the helmet specifically includes: First, the three acquired RGB raw images are converted to YC. b C r Color space, separating the luminance component Y and chrominance component C of an image. b C r ; Then, the lens distortion invalid areas at the edges of each image are cropped, retaining the effective imaging area of ​​1200×680 in the center; at the same time, a dedicated buffer area is opened to store the brightness component Y data and timing marks of 5 consecutive frames. Subsequently, based on local sharpness differences, areas of contamination such as smoke and water droplets on the lens surface are detected, and a contamination mask is generated. M smudge : ; In the formula: Clear() Indicates the local sharpness value. Laplacian() Represents the Laplace operator. gray(x,y) Representing coordinates (x,y) The single-channel grayscale value at that location; When the local sharpness is less than 50% of the global average sharpness and the area of ​​a continuous region is greater than 30 pixels, it is marked as a dirty region, i.e., the mask value is 1; the rest are normal regions, i.e., the mask value is 0; the mask is updated every 30 frames.

4. The image denoising enhancement and stitching method for fire scenes according to claim 3, characterized in that: Step S2 uses the current frame luminance component Y and the original chrominance component C output in step S1. b With C r Contamination mask M smudge For input; The separation of reflected light and spontaneous emission components and the estimation of transmittance are specifically as follows: Addressing the issue of traditional scattering models failing due to spontaneous emission sources in fire scenes, the imaging is decomposed into a spontaneous emission component and a scene reflected light component, establishing a dual-component imaging physical model. ; In the formula: Indicates the self-luminous component. Indicates the components of light reflected from the scene. t(x,y) Indicates smoke transmittance. A f Indicates the global ambient light intensity; Then, a coarse detection of the self-illuminating region is performed: first, the redness ratio is calculated: ; Preset percentage threshold T cr ,like R ratio > T cr If the value is 0, it is marked as a chroma candidate pixel; otherwise, it is not marked. Next, calculate the maximum brightness within a 15×15 neighborhood window. If the current pixel brightness is equal to the local maximum and is more than 1.5 times the global average brightness, then mark it as a brightness candidate pixel. If a pixel is both a chroma candidate pixel and a luminance candidate pixel, it is marked as a self-illuminating candidate; and a 3×3 morphological dilation is performed on the candidate region to obtain a coarse mask. M emissive-c ; Then, for a consecutive 5-frame luminance sequence, the temporal mean and temporal variance of each pixel are calculated: ; And the pixel fluctuation frequency is calculated using the luminance zero-crossing rate: ; In the formula: Indicates the frame interval time. sign() Represents a symbolic function; like Z(x,y) If the frequency is ∈[8Hz,15Hz], then it is determined to have flame flicker characteristics; If the coarse mask simultaneously satisfies the time-domain variance > 20 and Z(x,y) If the Hz range is ∈[8Hz, 15Hz], it is retained as a self-emissive region; static high-temperature self-emissive bodies have no flicker characteristics, and their chromaticity and brightness are directly preserved to obtain the final self-emissive mask. M emissive ; Next, estimate the transmittance of the non-emissive region: for the luminance component of the non-emissive region, calculate the minimum luminance value within the local window: ; In the formula: Indicates ( x,y A local window centered on ); Take the top 0.1% of pixels in the dark channel brightness and calculate the average of their corresponding original brightness as the global ambient light intensity. A f ; The initial transmittance is: , Indicates the smoke retention coefficient; Using the original brightness map as a guide map, the initial transmittance is optimized by guided filtering: ; In the formula: Represents the linear coefficients within the window; W k This represents the k-th filtering window; Perform time-domain weighted smoothing of the transmittance of the current frame and historical frames: ; In the formula: Represents the smoothing coefficient. This represents the optimized transmittance of the current frame; The scattering model is not applicable to the self-emissive region, so the transmittance is directly set to 1; For non-self-emissive regions, the reflected light component is directly inverted: ; For self-emissive areas: ; Finally, according to the Beer-Lambert law, the model depth of field is obtained through inversion: ; In the formula: Indicates the smoke scattering coefficient; For close-up areas within 3 meters, depth is calculated using the binocular parallax of the left and right cameras: ; In the formula: b d Indicates the binocular baseline distance. f d Indicates the camera's pixel focal length. This represents the disparity between corresponding matching pixels in the left and right images; Near-field area with binocular depth For reference, the mid-to-long-range views are based on the model's depth of field. To determine the optimal depth-of-field map, an S-shaped weighted smooth transition is applied to the transition area, resulting in the final depth-of-field map. d(x,y) .

5. The image denoising enhancement and stitching method for fire scenes according to claim 4, characterized in that: In step S2, the mixed noise parameter estimation specifically involves: fitting a Poisson-Gaussian mixed noise model using maximum likelihood estimation and dynamically updating the noise parameters. ; In the formula: Poisson() Indicates the Poisson distribution. z Indicates the grayscale value of the observed pixel. Indicates Poisson noise intensity; Gauss() Indicates a Gaussian distribution. This represents the mean of Gaussian noise. Indicates the standard deviation of Gaussian noise; This represents the convolution operation; Noise parameters are updated once per frame to achieve noise reduction adaptation across the entire temperature range; The specific steps of hazard feature pre-detection and protected area marking are as follows: Before noise reduction, a hazard area mask is generated based on color features, flicker frequency features, and local contrast features. M danger And perform 3×3 morphological expansion on the mask; Temporal motion saliency detection and region segmentation specifically involve calculating the motion amplitude and fluctuation variance between the current frame and the previous frame. ; In the formula: Var() This represents variance calculation. k - N 5: k Indicates from k - N 5 to the k A sequence of consecutive frames; Preset motion threshold Tm With noise threshold Td ,like M(x,y) > Tm If so, it is determined to be a dynamic target area; if M(x,y) ≤ Tm and D (x,y) > Td If so, it is determined to be a candidate transient noise region; The remaining pixels are static background areas; For candidate transient noise regions, further verification is performed using the brightness fluctuation period of 5 consecutive frames. If the fluctuation period is within (0.067s, 0.125s), it is officially determined to be a transient noise region; otherwise, it is classified as a dynamic target region. The spatial domain region division is as follows: For the dynamic target region, for the current frame's luminance component Y, a 3×3 neighborhood window is used to calculate the local variance of each pixel. At the same time, the noise threshold is obtained. Th1 , Th2 : ; In the formula: This represents the standard deviation of Gaussian noise at room temperature. like If , then it is a flat subregion; if If , then it is a texture sub-region; if Then it is an edge sub-region.

6. The image denoising enhancement and stitching method for fire scenes according to claim 5, characterized in that: In step S2, the region-based noise reduction process specifically includes: For static background areas: using the current frame as the center, an exponentially decaying weight function is applied: ; In the formula: i Indicates the historical frame number. Indicates the time smoothing coefficient; The temporal filtering result is obtained by normalizing and weighting the 5-frame luminance values ​​for each pixel: ; The time-domain output is smoothed with a 3×3 Gaussian kernel to further suppress spatial residual noise and avoid blurring weak details in the background. For transient noise regions: for pixels ( x,y Extract continuous n The brightness values ​​of the frames form a one-dimensional time series: The sequence is sorted by grayscale value, and the median value is taken as the output. If it is a normal transient noise area, then n=5; if it is a dangerous area, then n=3. For dynamic target areas: If it is a flat sub-region, first calculate the side length of the adaptive filtering window. And obtain the neighboring pixel weights. , ( u,v ) represents the coordinates of neighboring pixels within the window, ( x 0 ,y 0) represents the pixel coordinates of the window center. The minimum value is obtained; the final denoising result is: ; If it is a textured sub-region: the initial window size is 3×3, and the maximum window size is... , Indicates the reference Poisson noise intensity; First layer: If Z min < Z med < Z max If the window size is positive, it indicates that there is no strong impulse noise within the window, and the process proceeds to the next layer; otherwise, the window size is increased, and the first layer judgment is repeated until the window reaches the specified size. W max ; Second layer: If Z min < Z xy < Z max This indicates that the center pixel is not a noise point, and its original value is retained. Z xy Otherwise, if the center pixel is noise, replace it with... Z med ; If the first-level condition is still not met even when the window reaches its maximum size, the value in the window will be output directly. For edge sub-regions: use the sym4 wavelet basis to perform brightness mapping on the edge region. J Layer decomposition yields one low-frequency approximation coefficient and J The high-frequency detail coefficients of the group, and the threshold for the high-frequency coefficients of the j-th layer are: ; In the formula: N t Indicates the total number of pixels in the image; Perform soft threshold shrinkage on all high-frequency coefficients: ; In the formula: Represents the original high-frequency coefficients; The low-frequency coefficients are retained unchanged, and the processed high-frequency coefficients are combined with the sym4 wavelet inverse transform to reconstruct the image of the denoised edge region. The base strength of all the above filters is the depth-guided baseline denoising strength. Decide: ; In the formula: Indicates the baseline denoising intensity. k d This represents the depth-of-field gain coefficient. d max This represents the global maximum depth of field estimate; If a pixel is in a dangerous area, the denoising intensity is reduced to 30% of the baseline denoising intensity, and the filter window is simultaneously reduced and the threshold intensity is lowered.

7. The image denoising enhancement and stitching method for fire scenes according to claim 6, characterized in that: Step S3 uses the denoised luminance component Y and the original chrominance component. C b and C r Transmittance diagram t(x,y) Depth of field map d(x,y) Dirt mask M smudge Danger mask M danger Self-illuminating mask M emissive For input; The dual-region division of reflected light and self-emitting light is specifically as follows: the image is divided into two major categories based on a self-emitting light mask: a self-emitting light region and a reflected light region. Furthermore, based on transmittance, the reflected light region is divided into three sub-regions: high transmittance, medium transmittance, and low transmittance. Statistical analysis of the transmittance values ​​of effective pixels: ; The total number of pixels in the effective area is N valid Sort all effective transmittance in ascending order to obtain an ordered sequence. ; First, obtain the initial threshold for the partition: Then obtain the final threshold: ; like If so, it is a low-transmittance sub-region; if If , then it is the mid-transmitter region; if This is the high-transmittance sub-region.

8. The image denoising enhancement and stitching method for fire scenes according to claim 7, characterized in that: In step S3, the construction of the nonlinear enhancement transformation model is specifically as follows: ; In the formula: I in (x,y) This represents the normalized input grayscale value. These represent the shape control parameters of the transformation function; Represents a fully Beta function: ; In the formula: Represents the Gamma function; Indicates an incomplete Beta function: 。 9. The image denoising enhancement and stitching method for fire scenes according to claim 8, characterized in that: In step S3, the optimization of the partitioning parameters specifically involves using an immune genetic algorithm to optimize the three sub-regions of the reflected light region. parameter: ; ; In the formula: These represent the corresponding weight coefficients; E pt Represents image information entropy. L Represents the number of gray levels. p i Indicates the first i The probability of a gray level appearing in an image. p i = n i / (M×N), n i Indicates grayscale value i The total number of pixels, M×N represents the total number of pixels in the image; C jb Indicates local contrast. Indicates the local standard deviation of a single pixel. This represents a local neighborhood window centered at (x, y). W n Indicates the side length of the neighborhood window. This represents the average grayscale value within the neighborhood window; G avg Represents the average gradient. I x (x,y) , I y (x,y) These represent the horizontal gradient and the vertical gradient, respectively. D yx Indicates the effective detail ratio, This represents the gradient magnitude at pixel (x, y). L over Indicates the percentage of overexposed pixels. n over This indicates the total number of overexposed pixels.

10. The image denoising enhancement and stitching method for fire scenes according to claim 9, characterized in that: In step S3, the multi-factor pixel-wise enhancement intensity adjustment specifically involves: combining the depth-of-field guiding coefficient. Pollution compensation coefficient Dangerous area enhancement coefficient With self-luminescence enhancement intensity coefficient Pixel-by-pixel fine-tuning to enhance intensity: ; Depth of field guiding coefficient : ; In the formula: Indicates the reference reinforcement strength; Pollution compensation factor :like M smudge (x,y) =1, then Conversely, no pollution compensation coefficient is introduced; Dangerous area enhancement coefficient If it is a dangerous area, then Conversely, no enhancement coefficient for dangerous areas is introduced; Self-luminescence enhancement intensity coefficient :like M emissive (x,y) =1, then Conversely, no self-luminescence enhancement strength coefficient is introduced; Final realization Adjust parameters pixel by pixel: ; In the formula: These represent the baseline optimization parameters for the corresponding regions.