Flame detection method based on multi-modal fusion image
By using multimodal image fusion technology, visible light, near-infrared and thermal infrared images are acquired and aligned simultaneously to extract flame region features, which solves the problem of the limited applicability of single-modal detection and achieves high-precision flame detection and enhanced environmental adaptability.
Patent Information
- Application Number
- CN202511573492.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-31
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-10-31
AI Technical Summary
Existing flame detection methods that rely on single-modal images have limited applicability in complex environments, high false alarm rates, and poor adaptability to complex environments, failing to effectively improve the accuracy and response speed of flame detection.
By simultaneously acquiring visible light, near-infrared, and thermal infrared images, and using interpolation techniques for precise alignment, the features of the multimodal images are extracted. Combined with tile boundary tracking and temperature change analysis, the flame region is identified.
It improves the accuracy and environmental adaptability of flame detection, reduces false alarms and missed alarms, and ensures that the system has high stability and robustness in dynamic and interference environments.
Smart Images

Figure CN121053503A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a flame detection method based on multimodal fusion images. Background Technology
[0002] Image processing technology encompasses techniques for acquiring, analyzing, and understanding visible or non-visual image data to achieve purposes such as image information recognition, extraction, and discrimination. Its core content covers multiple aspects including image acquisition, image preprocessing, feature extraction, image recognition, and image fusion. Image acquisition involves collecting image data from different types of sensors, such as visible light and infrared; image preprocessing is used for noise reduction, contrast enhancement, or adjusting image structure to optimize subsequent processing performance; feature extraction includes modeling target regions in terms of shape, texture, and spectrum; image recognition uses extracted features to identify targets; and image fusion technology establishes connections between multi-source image information, improving the robustness and accuracy of recognition. It is widely used in scenarios such as video surveillance, intelligent security, traffic management, and disaster early warning. Especially in fire detection missions, image processing has become an important supplement to overcome the limitations of traditional early warning methods.
[0003] One flame detection method based on multimodal image fusion refers to an image processing technique that identifies the presence of flames by fusing information from different types of images. It primarily addresses the limitations of traditional smoke detectors and optical flame detectors, such as limited applicability, high false alarm rates, and weak adaptability to complex environments. The method proposes fusing multimodal images to enhance the stability of flame recognition. The technical aspects include the simultaneous acquisition and collaborative processing of visible light, near-infrared, and thermal infrared images; spatial alignment of multi-source images using image registration techniques; construction of feature representations for each modality through image feature extraction methods; and joint analysis of multimodal image information and flame region identification based on a feature fusion strategy. By leveraging the complementary characteristics of different image modalities in spectral response, spatial resolution, and temperature sensitivity, the method enhances the perception of flame image features, thereby enabling the discrimination and detection of flame targets in complex environments.
[0004] Existing technologies primarily rely on smoke detectors and optical flame detectors, which suffer from limitations in applicable scenarios, high false alarm rates, and poor adaptability to complex environments. Current methods fail to fully utilize the collaborative characteristics of multiple image modalities, resulting in the system's inability to accurately determine flame status and significant positioning errors in environments with high interference and rapid changes. Single-modal images cannot meet the flame detection needs of variable environments, and the lack of cross-modal information fusion processing fails to effectively improve the accuracy and response speed of flame detection, impacting the real-time performance and reliability of fire early warning systems. Summary of the Invention
[0005] To address the technical problems existing in the prior art, this invention provides a flame detection method based on multimodal image fusion. The technical solution is as follows: A flame detection method based on multimodal fusion images includes the following steps: S1: Acquire visible light, near-infrared and thermal infrared images of the same time point in the fire video monitoring area, extract pixel position indexes, establish a unified coordinate group based on the reference image, and generate an image position matching map group by matching the pixel position relationship between different images through coordinate mapping and interpolation. S2: Based on the position index in the image position matching group, extract the red-green difference of the visible light image, the grayscale shift of the near-infrared image, and the temperature distribution of the thermal infrared image, determine whether the corresponding pixel features of the three are shifted at the same time, extract the synchronous shift position, and generate a channel synchronous feature position set. S3: Call the channel synchronization feature location set, combine closed regions according to the adjacency rule, extract edge direction changes in the visible light image, extract brightness trends in the near-infrared image, filter out patches with stable structural orientation and continuous brightness trends, and generate a fusionable boundary tracking map group. S4: Extract multi-frame data of thermal infrared images from the fusionable boundary tracking map group, analyze the temperature change trend of each pixel, identify continuous heating points, determine whether they are spatially clustered to form a region, extract the concentrated heating structure, and generate thermal change clustering graphic blocks.
[0006] As a further aspect of the present invention, the image position matching map group includes time tags, size parameters, and pixel position indexes of visible light images, near-infrared images, and thermal infrared images; the channel synchronization feature position set includes red-green channel difference, near-infrared image grayscale shift, thermal infrared image temperature value, and synchronization deviation position set; the fusionable boundary tracking map group includes patch edge grayscale direction sequence, brightness trend curve, direction stability, and brightness consistency region; the thermal change clustering graphic block includes continuous inter-frame temperature change trend, clustering region, and continuously heating pixels.
[0007] As a further aspect of the present invention, the step of obtaining the image location matching map group is as follows: S101: Acquire visible light, near-infrared and thermal infrared images of the same time in the fire monitoring area, read the time label and size parameters of each type of image, extract the horizontal and vertical indexes of the image pixels, construct a set of coordinate indexes for the visible light images based on the time label matching, and establish a reference coordinate system based on the width and height in its size parameters to generate a visible light coordinate reference group. S102: Based on the visible light coordinate reference group, call the width and height parameters of the near-infrared and thermal infrared images, calculate the aspect ratio with the visible light image, perform coordinate scaling and offset processing, adjust the pixel position index of the corresponding image, establish the mapping relationship between coordinates, and obtain the coordinate mapping relationship group; S103: Call the coordinate mapping relationship group and the pixel intensity value of each image, select the surrounding pixels of the corresponding position according to each position under the reference coordinates, perform grayscale value interpolation and fill, integrate the pixel information of all channel images under the same coordinates, and obtain the image position matching map group.
[0008] As a further aspect of the present invention, the step of obtaining the channel synchronization feature location set is as follows: S201: Based on the position index in the image position matching map group, extract the red and green channel values of the corresponding positions in the visible light image, calculate the numerical difference between the two channels, compare the difference with the channel difference threshold, retain the position index that exceeds the threshold, and obtain the red and green channel difference position set. S202: Call the position index in the red-green channel difference position set, extract the gray value of the near-infrared image, calculate the difference magnitude of the difference relative to the overall gray value of the image, determine whether it exceeds the gray value offset threshold, filter out the position index that meets the condition, and obtain the gray value offset position set. S203: Based on the index in the grayscale offset position set, extract the temperature value of the thermal infrared image, compare the temperature with the high-order interval formed by the average temperature and the upper limit of the image temperature, filter out the position index that simultaneously meets the three-channel offset conditions, and obtain the channel synchronization feature position set.
[0009] As a further aspect of the present invention, the channel difference threshold is a value dynamically determined based on the statistical distribution of the overall red and green channels of the visible light image; The grayscale offset threshold is a value dynamically set based on the statistical dispersion of grayscale values in the near-infrared image. The process of comparing the temperature with the high-order interval formed by the average temperature of the image and the upper limit of the temperature specifically involves determining whether the temperature value of the thermal infrared image is greater than the average temperature of the image and less than the upper limit of the temperature.
[0010] As a further aspect of the present invention, the step of obtaining the fusionable boundary tracking map group is as follows: S301: Call all position indices in the channel synchronization feature position set, determine the continuous connection relationship according to the horizontal and vertical adjacency rules, combine the closed regions into a closed tile structure in the image, and generate a closed tile number set after excluding incomplete edge regions and isolated pixels. S302: Based on the closed patch number set, extract the gray values of edge pixels of each patch in the visible light image, arrange them according to the edge direction to generate a gray direction sequence, make a coherence judgment on the adjacent gray change directions in the sequence, and filter out the patch numbers whose direction change amplitude is less than the direction stability threshold to obtain a gray direction stable patch set. S303: Based on the patch number in the gray-scale direction stable patch set, extract the brightness value distribution of the corresponding region in the near-infrared image, generate a brightness trend curve in the horizontal direction, determine whether the continuous pixel segments in the curve meet the increasing trend condition, and perform intersection filtering with the result of the previous sub-step to obtain a fusionable boundary tracking patch group.
[0011] As a further aspect of the present invention, the step of obtaining the thermal change clustering graphic block is as follows: S401: Within a specified area in the fusionable boundary tracking map group, extract multi-frame sequence data corresponding to the thermal infrared image, record the temperature value of each frame in sequence according to the pixel position, calculate the temperature increment trend between consecutive frames, determine whether the increment is continuous and positive, retain the pixel position of continuous temperature rise, and obtain the pixel distribution value of continuous temperature rise. S402: Based on the continuously heated pixel distribution value, detect its horizontal and vertical adjacent distribution relationship in the image space, count the cluster density between adjacent heated pixels, judge the number and distribution range of pixels in the continuous connected area, filter out the area number that meets the cluster density threshold, and obtain the heated cluster area index set. S403: Call the tile number in the index set of the heating accumulation area, extract the closed boundary information of each area in the thermal infrared image, mark the position of the corresponding spatial area in the image, integrate the boundary frame and position coordinates of all areas, establish a spatial identification layer of the continuous heating area, and generate thermal change accumulation graphic blocks.
[0012] As a further aspect of the present invention, the process of determining whether the increment is continuously positive specifically requires that within a consecutive preset number of frames, the pixel temperature value of each frame image is higher than the pixel temperature value of the previous frame image at the same position. The process of judging the number and distribution range of pixels within a continuous connected region further includes calculating the perimeter and area of the continuous connected region, and filtering the continuous connected region based on the ratio of the perimeter to the area. The process of generating the thermal change cluster graphic block specifically involves merging all the selected continuous connected regions in the image into a single pixel set, and using the smallest convex hull that can completely cover all the merged pixels as the thermal change cluster graphic block.
[0013] As a further aspect of the present invention, the method further includes: S5: Extract the red and green color shift trends in the visible light images of all blocks in the thermal change clustering graphic block, analyze the color shift of the visible light image, the grayscale response of the near-infrared image, and the distribution of the heating position in the thermal infrared image, determine whether flame features are present at the same time, mark the areas that meet the conditions, and generate flame detection and localization results under the fused image. The flame detection and localization results include the location of regions that match flame characteristics and the flame localization in the fused image.
[0014] As a further aspect of the present invention, the step of obtaining the flame detection and positioning result is as follows: S501: Extract the visible light image regions corresponding to all blocks in the thermal change clustering graphic block, call the red and green channel values to calculate the pixel difference change trend, determine whether there is a continuous red and green channel difference enhancement phenomenon in the pixels inside the block, filter out the region numbers with a continuous offset trend, and obtain the red and green offset trend value. S502: Based on the region number in the red-green offset trend value, extract the gray value of the corresponding region in the near-infrared image, calculate the region gray mean, and compare it with the gray high response reference threshold to confirm whether the high response condition is met. Remove the numbers with gray mean values lower than the threshold to obtain the high response patch index set. S503: Extract the distribution range of heated pixels in the thermal infrared image based on the number in the high-response patch index set, determine whether the heated pixels are concentrated in a single area, filter the patch numbers that simultaneously meet the requirements of red-green channel offset, near-infrared high response and thermal infrared concentrated heating, establish corresponding coordinate markers, and generate flame detection and positioning results.
[0015] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: This invention simultaneously acquires visible light, near-infrared, and thermal infrared images, and uses interpolation technology to precisely align images of different modalities, optimizing cross-modal image position matching and improving image fusion performance. Based on the characteristics of image pixel positions deviating from conventional areas, this scheme accurately extracts the flame region, reducing false alarms and false negatives. Through patch boundary tracking and temperature change analysis, the accuracy of flame localization is further improved, especially in complex environments, effectively identifying flame targets. Utilizing the complementarity of multimodal images enhances the flame detection system's adaptability to environmental changes, ensuring high stability and robustness even in dynamic and interference-prone environments. Attached Figure Description
[0016] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a flowchart of the process for obtaining the image location matching map group according to the present invention; Figure 3This is a flowchart illustrating the process of obtaining the channel synchronization feature location set according to the present invention. Figure 4 This is a flowchart illustrating the process of obtaining boundary tracing graph groups that can be fused according to the present invention. Figure 5 This is a flowchart illustrating the process of obtaining the thermal change aggregation graphic block in this invention. Figure 6 This is a flowchart illustrating the process of obtaining the flame detection and positioning results of the present invention. Detailed Implementation
[0017] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0018] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0019] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the difference between them, their intended meanings are consistent. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the difference between them, their intended meanings are consistent.
[0020] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.
[0021] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0022] Please see Figure 1 This invention provides a technical solution: a flame detection method based on multimodal fusion images, comprising the following steps: S1: Acquire visible light, near-infrared and thermal infrared images of the same time in the fire video monitoring area, read the time label and size parameters of each image, extract the pixel position index and construct a reference coordinate group based on the reference image, call the aspect ratio rule to perform position mapping on the coordinates of non-reference images, complete cross-channel matching through interpolation, and generate an image position matching map group. S2: Based on the position index in the image position matching group, extract the difference between the red and green channels of the visible light image, calculate the grayscale shift amplitude of the near-infrared image, determine whether the temperature value of the thermal infrared image is in the high-level region, and determine whether the features are synchronously deviating from the normal region distribution based on the corresponding pixel positions of the three images. Extract the synchronous deviation positions to form a set and generate the channel synchronous feature position set. S3: Call the channel synchronization feature location set, combine it into a closed patch according to the horizontal and vertical continuous adjacency rules, extract the gray-scale direction sequence of the patch edge in the visible light image and determine whether the direction change is stable, extract the brightness trend curve in the near-infrared image and determine whether it is continuously increasing, filter the patch areas that meet the direction continuity and brightness trend, and generate a fusionable boundary tracking patch group. S4: Extract multi-frame data of thermal infrared image from the specified area of the fusion boundary tracking map group, calculate the temperature change trend of each pixel between consecutive frames, identify pixels with continuous heating, determine whether the pixels form a cluster area in space, filter out the map area with concentrated heating characteristics, and generate thermal change cluster graphic blocks. S5: Extract the red and green color shift trend of all blocks in the visible light image of the thermal change clustering graphic block, determine whether the grayscale of the near-infrared image is in a high response state, confirm whether the heating position in the thermal infrared image is concentrated in a single area, and if the information in the three types of images meets the requirements of flame characteristics, then delineate the area location and generate the flame detection and positioning result under the fused image.
[0023] The image location matching map set includes time labels, size parameters, and pixel location indices for visible light images, near-infrared images, and thermal infrared images; the channel synchronization feature location set includes red-green channel differences, near-infrared image grayscale shift, thermal infrared image temperature values, and synchronization deviation location sets; the fusionable boundary tracking map set includes patch edge grayscale direction sequences, brightness trend curves, direction stability, and brightness consistency regions; the thermal change clustering graphic blocks include continuous inter-frame temperature change trends, clustering regions, and continuously heating pixels; the flame detection and localization results include the location of regions that match flame characteristics and the flame localization in the fused image.
[0024] Please see Figure 2 The steps for obtaining the image location matching map group are as follows: S101: Acquire visible light, near-infrared and thermal infrared images of the same time in the fire monitoring area, read the time label and size parameters of each type of image, extract the horizontal and vertical indexes of the image pixels, construct a set of coordinate indexes for the visible light images based on the time label matching, and establish a reference coordinate system based on the width and height in its size parameters to generate a visible light coordinate reference group. A coordinate index set for visible light images is constructed based on time tag matching. Specifically, the visible light, near-infrared, and thermal infrared image acquisition devices at the system's front end are all set to the same clock source and synchronously trigger acquisition at a certain time T0, acquiring visible light image A, near-infrared image B, and thermal infrared image C respectively. At this time, the time tag of each of the three images is a unique identifier for T0. After the system reads this tag to confirm its consistency, it then reads the size parameters of each image. Visible light image A has a width of 1920 pixels and a height of 1080 pixels; near-infrared image B has a width of 768 pixels and a height of 576 pixels; and thermal infrared image C... The image has a width of 384 pixels and a height of 288 pixels. The system extracts the position indices of all pixels in the visible light image A. The horizontal index range is [0, 1919] and the vertical index range is [0, 1079]. A two-dimensional Cartesian coordinate system is established based on the width of 1920 and the height of 1080 of the visible light image. The origin (0, 0) of this coordinate system is located at the upper left corner of the image. The positive X-axis is horizontal to the right and the positive Y-axis is vertical to the bottom. Each integer coordinate point (x, y) in this coordinate system uniquely corresponds to a pixel in the visible light image A. Thus, a reference coordinate system is established, and a visible light coordinate reference group is generated.
[0025] S102: Based on the visible light coordinate reference group, call the width and height parameters of the near-infrared and thermal infrared images, calculate the aspect ratio between the near-infrared and thermal infrared images, perform coordinate scaling and offset processing, adjust the pixel position index of the corresponding image, establish the mapping relationship between coordinates, and obtain the coordinate mapping relationship group; Based on the visible light coordinate reference set, the width (768) and height (576) parameters of near-infrared image B are used to calculate its width ratio with that of visible light image A. and height ratio The width (384) and height (288) parameters of the thermal infrared image C are used to calculate the width ratio between it and the visible light image A. and height ratio When performing coordinate scaling and offset processing, for any point in the visible light coordinate reference set... Its corresponding mapped coordinates in near-infrared image B The calculation process is as follows and For example, the pixel with coordinates (1000, 800) in the reference group has the mapped coordinates (1000 / 2.5, 800 / 1.875) = (400, 426.67) in the near-infrared image. Similarly, its corresponding mapped coordinates in the thermal infrared image C are... The calculation process is as follows and The mapped coordinates of the point (1000, 800) in the thermal infrared image are (1000 / 5.0, 800 / 3.75) = (200, 213.33). By traversing all the coordinate points in the visible light coordinate reference group, the pixel position index of the near-infrared and thermal infrared images is adjusted and mapped to obtain the coordinate mapping relationship group.
[0026] S103: Call the coordinate mapping relationship group and the pixel intensity value of each image, select the surrounding pixels of the corresponding position according to each position under the reference coordinate, perform gray value interpolation and fill, integrate the pixel information of all channel images under the same coordinate, and obtain the image position matching map group. The coordinate mapping relationship group is called, along with the pixel intensity values of each image. Taking the aforementioned visible light reference coordinates (1000, 800) as an example, its mapped coordinates in the near-infrared image are (400, 426.67). These coordinates are not integers, so the pixel values cannot be directly extracted. Therefore, grayscale interpolation is performed. Specifically, four integer pixels around it are selected, namely the top left corner... =(400,426), top right corner =(401,426), bottom left corner =(400,427), bottom right corner =(401,427), read the grayscale values of these four pixels, assuming they are respectively... , , , Then, linear interpolation is performed in the x-direction to calculate the grayscale values of the projection points of point (400, 426.67) onto the two horizontal lines y=426 and y=427. Since its x-coordinate is 400 and there is no horizontal offset, the two intermediate values are respectively... and Linear interpolation is performed on these two intermediate values in the y-direction. The decimal part of the target point's y-coordinate is 0.67. Therefore, the interpolated grayscale value is... , The grayscale value 156.7 is assigned to the near-infrared channel at visible light coordinates (1000, 800). The same interpolation and filling operation is performed on the thermal infrared image and all other reference coordinate points. The pixel information of all channel images is integrated under a unified 1920x1080 coordinate system to obtain an image position matching map group.
[0027] Please see Figure 3 The steps for obtaining the channel synchronization feature location set are as follows: S201: Based on the position index in the image position matching map group, extract the red and green channel values of the corresponding positions in the visible light image, calculate the numerical difference between the two channels, compare the difference with the channel difference threshold, retain the position index that exceeds the threshold, and obtain the red and green channel difference position set; The channel difference threshold is a value dynamically determined based on the statistical distribution of the overall red and green channels in the visible light image; Based on the position index in the image position matching map group, a channel difference threshold is set. The threshold is determined as follows: Under normal monitoring conditions before the fire, 1000 consecutive frames of visible light images are acquired, and the red channel value of all pixels in each frame is calculated. With green channel value The difference Calculate the global average of all differences across these 1000 frames. and global standard deviation Experiments have verified that, in an office environment, the statistics show... , Set threshold Substituting the data yields This setting ensures that only pixels exceeding 99% of the normal fluctuation range are initially screened out. During detection, the pixel at coordinates (950, 520) in the visible light image is extracted, with a corresponding red channel value of 188 and a green channel value of 135. The difference between the two channel values is calculated. The difference of 53 was compared with the channel difference threshold of 37.3. The position index (950, 520) is preserved. Traverse all position indices in the image position matching group, perform the same calculation and comparison, and obtain the set of red and green channel difference positions.
[0028] S202: Call the position index in the red-green channel difference position set, extract the gray value of the near-infrared image, calculate the difference magnitude of the difference relative to the overall gray value of the image, determine whether it exceeds the gray value offset threshold, filter out the position index that meets the condition, and obtain the gray value offset position set. The grayscale offset threshold is a dynamically set value based on the statistical dispersion of grayscale values in the near-infrared image. Call the position index from the red-green channel difference position set and set the grayscale offset threshold. This threshold is dynamically set based on the statistical dispersion of the grayscale values of the current frame of the near-infrared image. Specifically, the average grayscale value of all 768x576 pixels in the current frame of the near-infrared image is calculated. With gray standard deviation Assuming that for the current frame image, the calculation is as follows: , Then set the grayscale offset threshold. The coefficient of 2.5 is the optimal value calibrated in 500 tests under different lighting conditions in non-flame scenarios. It is used to balance the false alarm rate and the missed alarm rate. A location index is extracted from the set of differences between the red and green channels, such as (950, 520) mentioned earlier. The near-infrared image grayscale value corresponding to this location is found to be 185.2 through image location matching. The difference between this value and the overall grayscale mean of the image is calculated. To determine whether the difference exceeds the grayscale offset threshold of 64.0, because... The location index (950, 520) meets the condition and is selected. This process is repeated for all location indices in the red-green channel difference location set to obtain the grayscale offset location set.
[0029] S203: Based on the index in the grayscale offset position set, extract the temperature value of the thermal infrared image, compare the temperature with the high-order interval formed by the image temperature mean and temperature upper limit, filter out the position index that simultaneously meets the three-channel offset conditions, and obtain the channel synchronization feature position set. The process of comparing the temperature with the high-order interval formed by the average temperature and the upper limit of the image temperature is specifically to determine whether the temperature value of the thermal infrared image is greater than the average temperature of the image and less than the upper limit of the temperature. Based on the index of the grayscale offset position set, the temperature value of the thermal infrared image is extracted, and this temperature value is compared with a high-order interval, which is constructed based on the temperature mean of the current frame of the thermal infrared image. With the preset upper temperature limit Among them, the upper temperature limit Based on the effective range and safety specifications of the thermal infrared sensor, the sensor used in this embodiment has a range of -20°C to 400°C. Therefore, it is set... For the current frame of the thermal infrared image, the system calculates the average temperature of all 384x288 pixels. °C, therefore, the criterion for determining the high-level range is the temperature value. Need to meet and Now, the index (950, 520) of the grayscale offset position set is extracted. A query of the image position matching group yields the corresponding thermal infrared image temperature value of 85.6°C. This temperature value is compared with the higher-order intervals because... and Since the condition is met, the position index (950, 520) is confirmed to satisfy the three-channel offset condition. The position index is retained, and the same temperature comparison and filtering is performed on all indices in the grayscale offset position set to obtain the channel synchronization feature position set.
[0030] Please see Figure 4The steps for obtaining a fusion boundary tracing map set are as follows: S301: Call all position indices in the channel synchronization feature position set, determine the continuous connection relationship according to the horizontal and vertical adjacency rules, combine the closed regions into a closed tile structure within the image, and generate a closed tile number set after excluding incomplete edge regions and isolated pixels. The system calls all position indices in the channel synchronization feature position set. For example, if the set contains the point set {(950,520),(951,520),(950,521),(966,530)}, the system uses the four-adjacency rule to determine continuous connection relationships. Specifically, it checks whether there are other indices in the set in the four directions (up, down, left, and right) for each position index. For (950,520), there are (951,520) to its right and (950,521) below it, so these three points are considered connected. However, the four adjacent positions of (966,530) are not in the set; if it is an isolated point, it is excluded. Traverse all points in the set, and group all interconnected points into the same connected component to form tiles. Assuming two tiles are formed, tile 1 consists of 15 consecutive pixels, and tile 2 consists only of the point (966, 530), the system will exclude tile 2 as an isolated pixel. At the same time, check whether tile 1 touches the four boundaries of the image. Assuming that tile 1 is located in the center of the image and does not touch the boundaries, it is a complete closed tile structure, and the system will assign it the number "Block_01". Perform the same combination and filtering on all connected components composed of channel synchronization feature position sets to generate a closed tile number set.
[0031] S302: Based on the closed patch number set, extract the gray value of the edge pixels of each patch in the visible light image, arrange them according to the edge direction to generate a gray value direction sequence, judge the continuity of adjacent gray value change directions in the sequence, and filter out the patch numbers whose direction change amplitude is less than the direction stability threshold to obtain a gray value direction stable patch set. Based on the closed patch number set, the patch numbered "Block_01" is extracted, and its edge pixels are located in the visible light image. Edge pixels are determined as follows: for any pixel within a patch, if at least one of its four neighboring pixels does not belong to that patch, then that pixel is an edge pixel. After extracting all edge pixels of "Block_01", they are arranged clockwise to obtain an ordered sequence of edge pixels, for example, the sequence is... The gray-level gradient direction at each edge pixel location is calculated by applying a 3x3 Sobel operator to the gray-level values of the pixel's neighborhood, thus calculating its horizontal gradient. and vertical gradient Then the gradient direction angle .
[0032] Table 1: Calculation Table of Edge Pixel Gradient Direction
[0033] Table 1 lists examples of gradient calculation for three consecutive pixels on an edge sequence, and the orientation stability threshold. The settings were based on statistical analysis of edge grayscale direction changes in 100 different flame video samples, taking the 95th percentile of the rate of change and setting it as [value missing]. The continuity of the direction of adjacent gray-level changes in the sequence is determined, i.e., calculation is performed. ,as well as ,because and This indicates that the edge direction change is stable. If the direction change amplitude of the entire patch edge sequence is less than this threshold, then its number "Block_01" is retained to obtain a set of grayscale direction stable patches.
[0034] S303: Based on the patch number in the gray-scale stable patch set, extract the brightness value distribution of the corresponding region in the near-infrared image, generate a brightness trend curve in the horizontal direction, determine whether the continuous pixel segments in the curve meet the increasing trend condition, and perform intersection filtering with the results of the previous sub-step to obtain a fusionable boundary tracking patch group. Based on the patch number "Block_01" in the gray-scale oriented stable patch set, the corresponding pixel region is extracted from the near-infrared image. This region is then scanned horizontally, row by row, to calculate the average brightness value of each row of pixels, generating a brightness trend curve. It is assumed that the "Block_01" region spans from the vertical position... arrive The range is calculated, and the resulting average brightness value sequence (i.e., brightness trend curve) is: To determine whether the curve satisfies the condition of a continuous increasing trend, the specific process is as follows: calculate the difference between adjacent elements in the sequence to obtain the difference sequence. The increasing trend condition requires that the proportion of positive elements in the difference sequence exceeds a preset ratio, such as 80%, and there are no two consecutive negative or zero differences. In this example, all differences are positive, satisfying the condition of 100% positive. Therefore, the brightness trend of the block is determined to be continuously increasing. The block number is then intersected with the result of the previous sub-step for filtering. Since "Block_01" satisfies both grayscale stability and increasing brightness trend, it is filtered out, thus obtaining a fusionable boundary tracking map group.
[0035] Please see Figure 5 The steps for obtaining the thermal change clustering graphic blocks are as follows: S401: Within a specified area in the fusionable boundary tracking map group, extract multi-frame sequence data corresponding to the thermal infrared image, record the temperature value of each frame in sequence according to the pixel position, calculate the temperature increment trend between consecutive frames, determine whether the increment is continuous and positive, retain the pixel position of continuous temperature rise, and obtain the pixel distribution value of continuous temperature rise. The process of determining whether the increment is continuously positive is specifically required that within a consecutive preset number of frames, the pixel temperature value of each frame is higher than the pixel temperature value at the same position in the previous frame. In the fusionable boundary tracking map group, specify the "Block_01" region, and extract 5 consecutive frames of data from the thermal infrared image at the corresponding position in this region, with timestamps of T0, T0+100ms, T0+200ms, T0+300ms, and T0+400ms. Record the temperature value of each frame sequentially according to the pixel position, and calculate the temperature increment between consecutive frames. Determine whether the increment is continuous and positive. Specifically, this determination process requires that within a preset number of consecutive frames, such as 4 frames, the pixel temperature value of each frame is higher than the pixel temperature value at the same position in the previous frame. As shown in Table 2 below, three pixels P1, P2, and P3 within the "Block_01" region are selected for display: Table 2: Multi-frame pixel temperature monitoring table
[0036] As shown in Table 2, for pixel P1, its four consecutive increments are +5.5, +6.7, +6.4, and +6.5, all of which are positive values. Therefore, P1 is determined to be a pixel with continuously rising temperatures. For pixel P2, its temperature of 77.9 in the third frame is lower than that of 78.2 in the second frame, and the increment is -0.3, which does not meet the condition of being continuously positive. Therefore, P2 is excluded. For pixel P3, its increments are also all positive values, and it is determined to be a pixel with continuously rising temperatures. This judgment is performed on all pixels in the "Block_01" area, and the positions of all pixels determined to be continuously rising temperatures are retained to obtain the distribution value of continuously rising temperature pixels.
[0037] S402: Based on the distribution value of continuously heated pixels, detect their horizontal and vertical adjacent distribution relationship in the image space, count the cluster density between adjacent heated pixels, judge the number and distribution range of pixels in continuous connected areas, filter out the area numbers that meet the cluster density threshold, and obtain the heated cluster area index set. The process of determining the number and distribution range of pixels within a continuous connected region further includes calculating the perimeter and area of the continuous connected region, and filtering the continuous connected region based on the ratio of the perimeter to the area. Based on the distribution values of pixels undergoing continuous heating, their horizontal and vertical adjacency relationships in the image space are detected. All continuously heated pixels are treated as a set, and four-neighbor connected component analysis is applied again to combine them into multiple continuous connected regions. The number of pixels and their distribution range in each region are then determined. This process further includes calculating the perimeter of the contour of each continuous connected region. With area ,area That is, the total number of pixels in the region, and the perimeter of the outline. The number of pixels at the region edge, and according to and The ratio screening region specifically uses shape factor as the screening criterion. The value is 1 for circular areas and less than 1 for other shapes. The aggregation density threshold is set on this shape factor. Statistical analysis was performed on the shape factors of 50 real flame samples and 50 samples of interfering heat sources (such as lights and reflections).
[0038] Table 3: Statistical Analysis of Shape Factor
[0039] Table 3 lists the experimental statistics used to determine the threshold. It was found that the shape factor of the flame region was generally greater than 0.4, while the shape factor of the interference source was mostly below 0.3. To achieve a better segmentation effect between the two, the aggregation density threshold was set to [value missing]. Assuming a continuous region of 28 pixels is composed of continuously heated pixels, and its perimeter is calculated to be 20 pixels, then its shape factor is... ,because The region meets the cluster density requirement, and its number is retained to obtain the index set of warming cluster regions.
[0040] S403: Call the tile number in the index set of heating accumulation area, extract the closed boundary information of each area in the thermal infrared image, mark the position of the corresponding spatial area in the image, integrate the boundary frame and position coordinates of all areas, establish a spatial identification layer of continuous heating area, and generate thermal change accumulation graphic block. The process of generating thermal change clustering graphic blocks involves merging the pixels of all selected continuous connected regions in the image and using the smallest convex hull that can completely cover all merged pixels as the thermal change clustering graphic block. The system retrieves all tile numbers from the heated cluster region index set, extracts the closed boundary information of each region from the thermal infrared image, and merges the pixel sets of all selected continuous connected regions in the image. Assuming that the heated cluster region index set contains two regions "Region_A" and "Region_B" that meet the conditions, the system merges the pixel coordinates of these two regions into a large point set. Then, it calculates a minimum convex hull that can completely cover this merged point set. This process uses a convex hull algorithm, such as Graham's scan or Monoton's Chain, to find a minimum polygon whose vertices are all taken from the merged point set, and all points are inside or on the edges of the polygon. The region defined by this convex hull polygon is the generated thermal change cluster graphic block.
[0041] Please see Figure 6 The steps for obtaining flame detection and location results are as follows: S501: Extract the visible light image regions corresponding to all patches in the thermal change clustering graphic block, call the red and green channel values to calculate the pixel difference change trend, determine whether there is a continuous red and green channel difference enhancement phenomenon in the pixels inside the patch, filter out the region numbers with a continuous offset trend, and obtain the red and green offset trend value. Extract the visible light image regions corresponding to all tiles in the thermal change clustering graphic block, and calculate their pixel differences using the red and green channel values. This yields a difference distribution map, which is used to determine whether there is a continuous direction of red-green channel difference enhancement within the pixels of an image patch. Specifically, the gradient of each pixel is calculated on this difference distribution map. It is a vector containing magnitude and direction, representing the direction and rate of the fastest change in the difference. The system then analyzes the directional distribution of all pixel gradient vectors within that region. If more than 70% of the pixels in a region have gradient directions pointing to a concentrated small area, for example, if the direction angle distribution is within a certain range... Within the interval, it is assumed that there is an enhancement phenomenon in a continuous direction. Assuming that for the current block, the difference between the red and green channels of most pixels gradually increases from 50 at the bottom to 80 at the top, forming a gradient field with a relatively consistent direction, then the region is considered to have a continuous offset trend, and its corresponding region number is retained to obtain the red-green offset trend value.
[0042] S502: Based on the region number in the red-green offset trend value, extract the gray value of the corresponding region in the near-infrared image, calculate the region gray mean, and compare it with the gray high response reference threshold to confirm whether the high response condition is met. Remove the numbers with gray mean values lower than the threshold to obtain the high response patch index set. Based on the region number in the red-green offset trend value, the gray value of the corresponding region in the near-infrared image is extracted, and the gray mean of the region is calculated. Gray-scale high response reference threshold The setting method is to take the 98th percentile of the grayscale values of all pixels in the current near-infrared image. This method ensures that the threshold can adapt to changes in the overall brightness of the image under different lighting conditions. Assuming that the 98th percentile of the grayscale value distribution of the current near-infrared image is 195.0, therefore... Now, calculate the mean gray value of the regions corresponding to the region numbers selected in the previous steps, and obtain... The mean is compared with the threshold because Once the region is confirmed to meet the high response condition, its number is retained. If the grayscale average of another region is 180.4, which is lower than the threshold, its number will be filtered out. All candidate regions are processed in this way to obtain the high response patch index set.
[0043] S503: Extract the distribution range of heated pixels in the thermal infrared image based on the number in the high-response patch index set, determine whether the heated pixels are concentrated in a single area, filter the patch numbers that simultaneously meet the requirements of red-green channel offset, near-infrared high response and thermal infrared concentrated heating, establish corresponding coordinate markers, and generate flame detection and positioning results. Based on the index number of the high-response patch set, the distribution range of heated pixels in the thermal infrared image is extracted, and it is determined whether these heated pixels are concentrated in a single region. This determination is achieved by calculating the standard deviation of the heated pixel coordinates, and the horizontal coordinates of all continuously heated pixels (derived from S401) within that region are calculated respectively. and vertical coordinates Standard deviation and And calculate its spatial discreteness. Simultaneously calculate the dimensions of the region itself, such as the diagonal length of its bounding rectangle. ,like If the concentration coefficient is less than a preset value, such as 0.3, the heating locations are considered to be concentrated. This coefficient of 0.3 is obtained by statistically analyzing 100 flame samples and taking the 90th percentile of the ratio of dispersion to size. For the currently indexed region, if its red-green channel offset trend meets the requirements, the near-infrared grayscale shows a high response state, and its heating pixel distribution is... If the calculation result is 0.21, which is less than 0.3, then the area is identified as a flame. The system establishes the coordinate markers of its bounding rectangle, for example (x:945,y:515,width:30,height:25), and generates the flame detection and localization result under the fused image.
[0044] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A flame detection method based on multimodal fusion images, characterized in that, Includes the following steps: S1: Acquire visible light, near-infrared and thermal infrared images of the same time point in the fire video monitoring area, extract pixel position indexes, establish a unified coordinate group based on the reference image, and generate an image position matching map group by matching the pixel position relationship between different images through coordinate mapping and interpolation. S2: Based on the position index in the image position matching group, extract the red-green difference of the visible light image, the grayscale shift of the near-infrared image, and the temperature distribution of the thermal infrared image, determine whether the corresponding pixel features of the three are shifted at the same time, extract the synchronous shift position, and generate a channel synchronous feature position set. S3: Call the channel synchronization feature location set, combine closed regions according to the adjacency rule, extract edge direction changes in the visible light image, extract brightness trends in the near-infrared image, filter out patches with stable structural orientation and continuous brightness trends, and generate a fusionable boundary tracking map group. S4: Extract multi-frame data of thermal infrared images from the fusionable boundary tracking map group, analyze the temperature change trend of each pixel, identify continuous heating points, determine whether they are spatially clustered to form a region, extract the concentrated heating structure, and generate thermal change clustering graphic blocks.
2. The flame detection method based on multimodal fusion image according to claim 1, characterized in that: The image location matching map group includes time tags, size parameters, and pixel location indexes of visible light images, near-infrared images, and thermal infrared images; the channel synchronization feature location set includes red-green channel difference, near-infrared image grayscale shift, thermal infrared image temperature value, and synchronization deviation location set; the fusionable boundary tracking map group includes patch edge grayscale direction sequence, brightness trend curve, direction stability, and brightness consistency region. The thermal change clustering graphic block includes the continuous inter-frame temperature change trend, clustering region, and continuously heating pixels.
3. The flame detection method based on multimodal fusion image according to claim 1, characterized in that, The steps for obtaining the image location matching map group are as follows: S101: Acquire visible light, near-infrared and thermal infrared images of the same time in the fire monitoring area, read the time label and size parameters of each type of image, extract the horizontal and vertical indexes of the image pixels, construct a set of coordinate indexes for the visible light images based on the time label matching, and establish a reference coordinate system based on the width and height in its size parameters to generate a visible light coordinate reference group. S102: Based on the visible light coordinate reference group, call the width and height parameters of the near-infrared and thermal infrared images, calculate the aspect ratio with the visible light image, perform coordinate scaling and offset processing, adjust the pixel position index of the corresponding image, establish the mapping relationship between coordinates, and obtain the coordinate mapping relationship group; S103: Call the coordinate mapping relationship group and the pixel intensity value of each image, select the surrounding pixels of the corresponding position according to each position under the reference coordinates, perform grayscale value interpolation and fill, integrate the pixel information of all channel images under the same coordinates, and obtain the image position matching map group.
4. The flame detection method based on multimodal fusion image according to claim 1, characterized in that, The steps for obtaining the channel synchronization feature location set are as follows: S201: Based on the position index in the image position matching map group, extract the red and green channel values of the corresponding positions in the visible light image, calculate the numerical difference between the two channels, compare the difference with the channel difference threshold, retain the position index that exceeds the threshold, and obtain the red and green channel difference position set. S202: Call the position index in the red-green channel difference position set, extract the gray value of the near-infrared image, calculate the difference magnitude of the difference relative to the overall gray value of the image, determine whether it exceeds the gray value offset threshold, filter out the position index that meets the condition, and obtain the gray value offset position set. S203: Based on the index in the grayscale offset position set, extract the temperature value of the thermal infrared image, compare the temperature with the high-order interval formed by the average temperature and the upper limit of the image temperature, filter out the position index that simultaneously meets the three-channel offset conditions, and obtain the channel synchronization feature position set.
5. The flame detection method based on multimodal fusion image according to claim 4, characterized in that, The channel difference threshold is a value dynamically determined based on the statistical distribution of the overall red and green channels of the visible light image; The grayscale offset threshold is a value dynamically set based on the statistical dispersion of grayscale values in the near-infrared image. The process of comparing the temperature with the high-order interval formed by the average temperature of the image and the upper limit of the temperature specifically involves determining whether the temperature value of the thermal infrared image is greater than the average temperature of the image and less than the upper limit of the temperature.
6. The flame detection method based on multimodal fusion image according to claim 1, characterized in that, The steps for obtaining the fusionable boundary tracing map set are as follows: S301: Call all position indices in the channel synchronization feature position set, determine the continuous connection relationship according to the horizontal and vertical adjacency rules, combine the closed regions into a closed tile structure in the image, and generate a closed tile number set after excluding incomplete edge regions and isolated pixels. S302: Based on the closed patch number set, extract the gray values of edge pixels of each patch in the visible light image, arrange them according to the edge direction to generate a gray direction sequence, make a coherence judgment on the adjacent gray change directions in the sequence, and filter out the patch numbers whose direction change amplitude is less than the direction stability threshold to obtain a gray direction stable patch set. S303: Based on the patch number in the gray-scale direction stable patch set, extract the brightness value distribution of the corresponding region in the near-infrared image, generate a brightness trend curve in the horizontal direction, determine whether the continuous pixel segments in the curve meet the increasing trend condition, and perform intersection filtering with the result of the previous sub-step to obtain a fusionable boundary tracking patch group.
7. The flame detection method based on multimodal fusion image according to claim 1, characterized in that, The steps for obtaining the thermal change clustered graphic block are as follows: S401: Within a specified area in the fusionable boundary tracking map group, extract multi-frame sequence data corresponding to the thermal infrared image, record the temperature value of each frame in sequence according to the pixel position, calculate the temperature increment trend between consecutive frames, determine whether the increment is continuous and positive, retain the pixel position of continuous temperature rise, and obtain the pixel distribution value of continuous temperature rise. S402: Based on the continuously heated pixel distribution value, detect its horizontal and vertical adjacent distribution relationship in the image space, count the cluster density between adjacent heated pixels, judge the number and distribution range of pixels in the continuous connected area, filter out the area number that meets the cluster density threshold, and obtain the heated cluster area index set. S403: Call the tile number in the index set of the heating accumulation area, extract the closed boundary information of each area in the thermal infrared image, mark the position of the corresponding spatial area in the image, integrate the boundary frame and position coordinates of all areas, establish a spatial identification layer of the continuous heating area, and generate thermal change accumulation graphic blocks.
8. The flame detection method based on multimodal fusion image according to claim 7, characterized in that, The process of determining whether the increment is continuously positive specifically requires that within a consecutive preset number of frames, the pixel temperature value of each frame image is higher than the pixel temperature value at the same position in the previous frame image. The process of judging the number and distribution range of pixels within a continuous connected region further includes calculating the perimeter and area of the continuous connected region, and filtering the continuous connected region based on the ratio of the perimeter to the area. The process of generating the thermal change cluster graphic block specifically involves merging all the selected continuous connected regions in the image into a single pixel set, and using the smallest convex hull that can completely cover all the merged pixels as the thermal change cluster graphic block.
9. The flame detection method based on multimodal fusion image according to claim 1, characterized in that, The method further includes: S5: Extract the red and green color shift trends in the visible light images of all blocks in the thermal change clustering graphic block, analyze the color shift of the visible light image, the grayscale response of the near-infrared image, and the distribution of the heating position in the thermal infrared image, determine whether flame features are present at the same time, mark the areas that meet the conditions, and generate flame detection and localization results under the fused image. The flame detection and localization results include the location of regions that match flame characteristics and the flame localization in the fused image.
10. The flame detection method based on multimodal fusion image according to claim 9, characterized in that, The steps for obtaining the flame detection and positioning results are as follows: S501: Extract the visible light image regions corresponding to all blocks in the thermal change clustering graphic block, call the red and green channel values to calculate the pixel difference change trend, determine whether there is a continuous red and green channel difference enhancement phenomenon in the pixels inside the block, filter out the region numbers with a continuous offset trend, and obtain the red and green offset trend value. S502: Based on the region number in the red-green offset trend value, extract the gray value of the corresponding region in the near-infrared image, calculate the region gray mean, and compare it with the gray high response reference threshold to confirm whether the high response condition is met. Remove the numbers with gray mean values lower than the threshold to obtain the high response patch index set. S503: Extract the distribution range of heated pixels in the thermal infrared image based on the number in the high-response patch index set, determine whether the heated pixels are concentrated in a single area, filter the patch numbers that simultaneously meet the requirements of red-green channel offset, near-infrared high response and thermal infrared concentrated heating, establish corresponding coordinate markers, and generate flame detection and positioning results.
Citation Information
Patent Citations
Fire alarm early warning system based on temperature monitoring
CN115862259A
Forest fire detection method based on near-infrared and thermal imaging image fusion
CN116403160A
Flame detection method and device based on visible light image and near-infrared image
CN118691932A
Multi-mode indoor fire identification method, device and equipment and storage medium
CN119380156A
Multi-modal data fusion human body target identification method based on automatic control
CN120635949A
Cited By
Wetland information extraction method and system based on multi-source remote sensing data fusion
CN121365369A