A flame detection method based on multi-modal fusion image

By using multimodal image fusion technology, the problem of insufficient accuracy and adaptability of existing flame detection methods in complex environments has been solved, achieving high-precision flame positioning and reducing false alarms and missed alarms, thus improving the reliability of fire early warning.

CN121053503BActive Publication Date: 2026-02-06SHANXI AIWEISEN AVIATION ENG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511573492.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2026-02-06
Estimated Expiration
2045-10-31

AI Technical Summary

Technical Problem

Existing flame detection methods rely on single-modal images, which have limitations in applicable scenarios, high false alarm rates, and poor adaptability to complex environments. They cannot accurately determine the flame status in environments with significant interference and rapid changes, thus affecting the real-time performance and reliability of fire early warning.

Method used

By simultaneously acquiring visible light, near-infrared, and thermal infrared images, and using interpolation techniques for precise alignment, the system extracts image pixel position deviation features. Combined with patch boundary tracking and temperature change analysis, it identifies flame regions and generates flame detection and localization results.

Benefits of technology

It improves the accuracy and response speed of flame detection, enhances the system's adaptability in complex environments, reduces false alarms and missed alarms, and ensures high stability and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121053503B_ABST
    Figure CN121053503B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image processing, in particular to a flame detection method based on multi-modal fusion image, which comprises the steps of acquiring visible light, near-infrared and thermal infrared images and matching positions, extracting a feature set generated by multi-channel difference, combining adjacent blocks to determine direction and brightness trend, identifying a temperature rising and gathering area, confirming flame features and positioning detection results by comprehensively analyzing three types of images. The present application collects visible light, near-infrared and thermal infrared images, realizes accurate alignment across modalities by using interpolation, and optimizes the fusion effect. The flame area is accurately extracted according to the pixel deviation feature, and the false positives and false negatives are reduced. The positioning accuracy is improved by combining block boundary tracking and temperature change analysis, and the flame target is effectively identified in a complex environment. The multi-modal complementation enhances the adaptability of the system, ensuring high stability and robustness in a dynamic interference environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, and in particular to a flame detection method based on multi-modal fusion images. BACKGROUND

[0002] The technical field of image processing includes techniques for collecting, analyzing, and understanding visual or non-visual image data to achieve image information recognition, extraction, and discrimination. The core content covers image acquisition, image preprocessing, feature extraction, image recognition, and image fusion. Image acquisition involves the collection of image data by different types of sensors, such as visible light and infrared. Image preprocessing is used to remove noise, enhance contrast, or adjust image structure to optimize subsequent processing performance. Feature extraction includes modeling target regions in terms of shape, texture, and spectrum. Image recognition relies on extracted features for target discrimination. Image fusion technology establishes connections between multiple sources of image information, improving the robustness and accuracy of recognition. It is widely used in video surveillance, intelligent security, traffic management, disaster warning, and other scenarios. In particular, in fire detection tasks, image processing has become an important supplement to break through the limitations of traditional warning methods.

[0003] Among them, a flame detection method based on multi-modal fusion images refers to an image processing technology scheme that fuses different types of image information to identify the presence of flames. It mainly addresses the problems of limited application scenarios, high false alarm rates, and weak adaptability to complex environments of traditional smoke alarms and optical flame detectors. The idea of fusing multi-modal images to enhance the stability of flame recognition is proposed. The technology involved includes the synchronous acquisition and collaborative processing of visible light images, near-infrared images, and thermal infrared images, the use of image registration technology for multi-source image spatial alignment, the construction of feature expression for each modality through image feature extraction methods, and the joint analysis of multi-modal image information and flame region recognition based on feature fusion strategies. By utilizing the complementary characteristics of different image modalities in spectral response, spatial resolution, and temperature sensitivity, the perception of flame image features is improved, thereby completing the discrimination and detection of flame targets in complex environments.

[0004] Existing technologies mainly rely on smoke alarms and optical flame detectors, which have limitations in application scenarios, high false alarm rates, and poor adaptability to complex environments. Existing methods do not fully utilize the collaborative characteristics of multiple image modalities, resulting in inaccurate flame state determination and large positioning errors in environments with significant interference and rapid changes. Single modal images cannot meet the flame detection requirements in variable environments, and lack of cross-modal information fusion processing cannot effectively improve the accuracy and response speed of flame detection, affecting the real-time and reliability of fire warning. SUMMARY

[0005] In order to solve the technical problems existing in the prior art, the embodiment of the present application provides a flame detection method based on multi-modal fusion images. The technical scheme is as follows:

[0006] A flame detection method based on multi-modal fusion images comprises the following steps:

[0007] S1: Obtain visible light, near-infrared and thermal infrared images at the same time point in a fire video monitoring area, extract pixel position indexes, establish a unified coordinate group according to a reference image, match pixel position relationships among different images through coordinate mapping and interpolation, and generate an image position matching image group;

[0008] S2: Based on the position indexes in the image position matching image group, extract the red-green color difference value of the visible light image, the gray offset of the near-infrared image and the temperature distribution of the thermal infrared image, judge whether the corresponding pixel features of the three are simultaneously offset, extract the synchronous offset position, and generate a channel synchronous feature position set;

[0009] S3: Call the channel synchronous feature position set, combine the closed area according to the adjacency rule, extract the edge direction change in the visible light image and the brightness trend in the near-infrared image, screen the image blocks with stable structure direction and continuous brightness trend, and generate a fusible boundary tracking image group;

[0010] S4: Extract multiple frames of data of the thermal infrared image in the fusible boundary tracking image group, analyze the temperature change trend of each pixel, identify the continuous warming point position, judge whether it is gathered to form an area in space, extract the concentrated warming structure, and generate a thermal change aggregation graphic block.

[0011] As a further scheme of the present application, the image position matching image group comprises time labels, size parameters and pixel position indexes of the visible light image, the near-infrared image and the thermal infrared image; the channel synchronous feature position set comprises red-green channel difference, near-infrared image gray offset, thermal infrared image temperature value and synchronous offset position set; the fusible boundary tracking image group comprises image block edge gray direction sequence, brightness trend curve, direction stability and brightness consistency area; and the thermal change aggregation graphic block comprises continuous inter-frame temperature change trend, aggregation area and continuous warming pixel point.

[0012] As a further scheme of the present application, the acquisition step of the image position matching image group is as follows:

[0013] S101: Obtain visible light, near-infrared and thermal infrared images at the same time in a fire monitoring area, read the time labels and size parameters of each type of image, extract the horizontal and vertical indexes of the image pixels, construct a coordinate index set of the visible light image according to the time label matching, and establish a reference coordinate system according to the width and height in the size parameters, to generate a visible light coordinate reference group;

[0014] S102: based on the visible light coordinate reference group, call the width and height parameters of the near-infrared and thermal infrared image, calculate the width-height ratio between the visible light image, perform coordinate scaling and offset processing, adjust the pixel position index of the corresponding image, establish the mapping relationship between the coordinates, obtain the coordinate mapping relationship group;

[0015] S103: call the coordinate mapping relationship group and the pixel intensity value of each image, according to each position under the reference coordinate, select the corresponding position peripheral pixel point, perform gray value interpolation filling, integrate the pixel information of all channel images under the unified coordinate, obtain the image position matching graph group.

[0016] As a further scheme of the application, the channel synchronization feature position set acquisition step is:

[0017] S201: based on the position index in the image position matching graph group, extract the red and green channel values of the corresponding position of the visible light image, calculate the numerical difference value between the two channels, compare the difference value with the channel difference threshold, retain the position index exceeding the threshold, and obtain the red-green channel difference position set;

[0018] S202: call the position index in the red-green channel difference position set, extract the gray value of the near-infrared image, calculate the difference amplitude of the gray value relative to the overall gray mean value of the image, judge whether it exceeds the gray offset threshold, select the position index that meets the condition, and obtain the gray offset position set;

[0019] S203: according to the index in the gray offset position set, extract the temperature value of the thermal infrared image, compare the temperature with the high bit interval composed of the image temperature mean value and the temperature upper limit, select the position index that meets the three-channel offset condition at the same time, and obtain the channel synchronization feature position set.

[0020] As a further scheme of the application, the channel difference threshold is a value dynamically determined according to the statistical distribution of the overall red channel and green channel of the visible light image;

[0021] The gray offset threshold is a value dynamically set according to the statistical dispersion degree of the gray value in the near-infrared image;

[0022] The process of comparing the temperature with the high bit interval composed of the image temperature mean value and the temperature upper limit is specifically to judge whether the temperature value of the thermal infrared image is greater than the image temperature mean value and less than the temperature upper limit.

[0023] As a further scheme of the application, the acquisition step of the fusible boundary tracking graph group is:

[0024] S301: Call all position indexes in the channel synchronization feature position set, judge the continuous connection relationship according to the horizontal and vertical adjacency rules, combine the connected closed area into the closed block structure in the image, after excluding the incomplete edge area and isolated pixel points, generate the closed block number set;

[0025] S302: Based on the closed block number set, extract the edge pixel gray value of each block in the visible light image, generate a gray direction sequence according to the edge direction, judge the continuity of the adjacent gray direction in the sequence, select the block number whose direction change amplitude is less than the direction stability threshold, and obtain the gray direction stable block set;

[0026] S303: According to the block number in the gray direction stable block set, extract the brightness value distribution of the corresponding area in the near-infrared image, generate a brightness trend curve according to the horizontal direction, judge whether the continuous pixel segment in the curve satisfies the increasing trend condition, and perform intersection filtering with the result of the previous sub-step, obtain the fusible boundary tracking graph group.

[0027] As a further scheme of the present application, the acquisition step of the thermal change aggregation pattern block is:

[0028] S401: In the specified area range in the fusible boundary tracking graph group, extract the corresponding multi-frame sequence data of the thermal infrared image, record the temperature value of each frame in turn according to the pixel position, calculate the temperature increment trend between the continuous frames, judge whether the increment is continuously positive, retain the pixel position with continuous temperature rise, and obtain the continuous temperature rise pixel distribution value;

[0029] S402: Based on the continuous temperature rise pixel distribution value, detect its horizontal and vertical adjacent distribution relationship in the image space, count the aggregation density between the adjacent temperature rise pixels, judge the pixel number and distribution range in the continuous connection area, select the area number that satisfies the aggregation density threshold, and obtain the temperature rise aggregation area index set;

[0030] S403: Call the block number in the temperature rise aggregation area index set, extract the closed boundary information of each area in the thermal infrared image, mark the position of the corresponding space area in the image, integrate the boundary framework and position coordinates of all areas, establish the space identification layer of the continuous temperature rise area, and generate the thermal change aggregation pattern block.

[0031] As a further scheme of the present application, the process of judging whether the increment is continuously positive is specifically that the temperature value of each frame image is required to be higher than the temperature value of the pixel at the same position in the previous frame image within a continuous preset frame number;

[0032] The process of judging the number and distribution range of pixels in the continuous connection region further comprises calculating the contour perimeter and area of the continuous connection region, and screening the continuous connection region according to the ratio of the contour perimeter to the area.

[0033] The process of generating the heat change aggregation graphic block is specifically that all the screened continuous connection regions are combined in the image, and a minimum convex hull capable of completely covering all the combined pixels is taken as the heat change aggregation graphic block.

[0034] As a further scheme of the present application, the method further comprises:

[0035] S5: Extracting a red-green color offset trend in all tile visible light images in the heat change aggregation graphic block, analyzing the visible light image color offset, the near-infrared image gray response, and the heat infrared image temperature rise position distribution, judging whether the flame features are simultaneously present, marking the region meeting the conditions, and generating a flame detection positioning result under the fusion image.

[0036] The flame detection positioning result comprises a region position meeting the flame features and a flame positioning under the fusion image.

[0037] As a further scheme of the present application, the acquisition step of the flame detection positioning result is:

[0038] S501: Extracting a visible light image region corresponding to all tiles in the heat change aggregation graphic block, calling the red and green channel values to calculate the pixel difference change trend, judging whether the red-green channel difference in the continuous direction is enhanced in the tile internal pixels, screening out the region number with the continuous offset trend, and obtaining the red-green offset trend value.

[0039] S502: Based on the region number in the red-green offset trend value, extracting the gray value of the corresponding region in the near-infrared image, calculating the region gray mean value, and comparing the region gray mean value with the gray high response reference threshold value to confirm whether the high response condition is met, screening out the number whose gray mean value is lower than the threshold value, and obtaining a high response tile index set.

[0040] S503: Extracting the temperature rise pixel distribution range in the heat infrared image according to the number in the high response tile index set, judging whether the temperature rise pixels are concentrated in a single region, screening the tile number meeting the red-green channel offset, the near-infrared high response, and the heat infrared concentrated temperature rise at the same time, establishing a corresponding coordinate mark, and generating a flame detection positioning result.

[0041] The technical scheme provided by the embodiment of the present application has at least the following beneficial effects:

[0042] In the present application, the visible light, near-infrared and thermal infrared images are synchronously collected, and the interpolation technology is used to accurately align the different modal images, thereby optimizing the cross-modal image position matching and improving the image fusion effect. According to the characteristics of the image pixel position deviating from the conventional region, the flame region is accurately extracted, and the false positives and false negatives are reduced. Through the tile boundary tracking and temperature change analysis, the accuracy of flame positioning is further improved, especially in complex environments, the flame target can be effectively identified. By using the complementarity of multi-modal images, the adaptability of the flame detection system to environmental changes is enhanced, and the system still has high stability and high robustness in dynamic and disturbed environments. BRIEF DESCRIPTION OF DRAWINGS

[0043] Figure 1 is a flowchart of the method of the present application;

[0044] Figure 2 is a flowchart of the image position matching group acquisition of the present application;

[0045] Figure 3 is a flowchart of the acquisition of the channel synchronization feature position set of the present application;

[0046] Figure 4 is a flowchart of the acquisition of the fusible boundary tracking group of the present application;

[0047] Figure 5 is a flowchart of the acquisition of the thermal change aggregation image block of the present application;

[0048] Figure 6 is a flowchart of the acquisition of the flame detection positioning result of the present application. DETAILED DESCRIPTION

[0049] The technical solutions in the present application will be described below with reference to the drawings.

[0050] In the embodiments of the present application, the words such as "example", "for example" are used to represent as an example, illustration or description. Any embodiment or design scheme described as "example" in the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the word "example" is intended to present the concept in a specific manner. In addition, in the embodiments of the present application, the meaning expressed by "and / or" can be both, or can be one of the two.

[0051] In the embodiments of the present application, "image" and "picture" can be used interchangeably at times, and it should be pointed out that when the distinction is not emphasized, the meanings expressed are consistent. "Of", "corresponding" and "corresponding" can be used interchangeably at times, and it should be pointed out that when the distinction is not emphasized, the meanings expressed are consistent.

[0052] In the embodiments of the present application, sometimes the subscript such as W1 may be written in the form of non-subscript such as W1, and the meanings expressed thereby are consistent when the difference is not emphasized.

[0053] In order to make the technical problems, technical solutions and advantages to be solved by the present application more clear, the following will be described in detail in combination with the drawings and specific embodiments.

[0054] Please refer to Figure 1 The present application provides a technical solution: a flame detection method based on multi-modal fusion image, comprising the following steps:

[0055] S1: acquiring visible light image, near-infrared image and thermal infrared image in the same time in the fire video monitoring area, reading the time label and size parameter of each image, extracting the pixel position index and constructing the reference coordinate group according to the reference image, calling the width-height ratio rule to perform position mapping on the non-reference image coordinates, completing cross-channel matching through interpolation method, and generating image position matching image group;

[0056] S2: based on the position index in the image position matching image group, corresponding to extract the red-green color channel difference value of the visible light image, calculating the gray scale offset amplitude of the near-infrared image, judging whether the temperature value of the thermal infrared image is in the high position area, judging whether the features are synchronously deviated from the conventional area distribution according to the corresponding pixel positions of the three images, extracting the synchronously deviated position group to generate the channel synchronous feature position set;

[0057] S3: calling the channel synchronous feature position set, combining into a closed tile according to the horizontal and vertical continuous adjacency rule, extracting the edge gray direction sequence in the visible light image and judging whether the direction change is stable, extracting the brightness trend curve in the near-infrared image and judging whether it is continuously increasing, screening the tile area that satisfies the direction continuity and brightness trend consistency, and generating the fusion boundary tracking image group;

[0058] S4: extracting the thermal infrared image multi-frame data in the specified area of the fusion boundary tracking image group, calculating the temperature change trend of each pixel between consecutive frames, identifying the pixel points with continuous temperature rise, judging whether the pixel points form an aggregation area in space, screening out the tile area with the characteristics of concentrated temperature rise, and generating the heat change aggregation graphic block;

[0059] S5: extracting the red-green color offset trend in the visible light image of all tiles in the heat change aggregation graphic block, judging whether the near-infrared image gray scale is in a high response state, confirming whether the temperature rise position in the thermal infrared image is concentrated in a single area, and if the information in the three types of images meets the flame feature requirements, then the region position is circled, and the flame detection positioning result under the fusion image is generated.

[0060] The image position matching graph set comprises time tags, size parameters, and pixel position indexes of visible light images, near-infrared images, and thermal infrared images; the channel synchronization feature position set comprises a red-green channel difference, a near-infrared image gray offset, a thermal infrared image temperature value, and a synchronization deviation position set; the fusible boundary tracking graph set comprises a graph block edge gray direction sequence, a brightness trend curve, a direction stability, and a brightness consistency area; the thermal change aggregation graph block comprises a continuous frame temperature change trend, an aggregation area, and a sustained temperature rise pixel point; and the flame detection positioning result comprises a region position meeting the flame feature and a flame positioning under a fusion image.

[0061] Referring to Figure 2 , the acquisition step of the image position matching graph set is as follows:

[0062] S101: Visible light, near-infrared, and thermal infrared images at the same time in the fire monitoring area are acquired, time tags and size parameters of each type of image are read, horizontal and vertical indexes of image pixels are extracted, a coordinate index set of the visible light image is matched and constructed according to the time tags, a reference coordinate system is established according to the width and height in the size parameters, and a visible light coordinate reference group is generated;

[0063] According to the time tag matching, the coordinate index set of the visible light image is constructed. Specifically, the visible light, near-infrared, and thermal infrared image acquisition devices at the front end of the system are all set to the same clock source, and are synchronously triggered to acquire the visible light image A, the near-infrared image B, and the thermal infrared image C at a time T0. At this time, the time tags of the three images are all unique identifiers of T0. After the system reads the tags and confirms their consistency, the size parameters of each image are read. The width of the visible light image A is 1920 pixels, the height is 1080 pixels, the width of the near-infrared image B is 768 pixels, the height is 576 pixels, the width of the thermal infrared image C is 384 pixels, and the height is 288 pixels. The system extracts the full pixel position index of the visible light image A, the horizontal direction index range is [0, 1919], and the vertical direction index range is [0, 1079]. A two-dimensional Cartesian coordinate system is established based on the width 1920 and the height 1080 of the visible light image. The origin (0, 0) of the coordinate system is located at the upper left corner of the image, the positive direction of the X-axis is horizontal to the right, and the positive direction of the Y-axis is vertical downward. Each integer coordinate point (x, y) in the coordinate system corresponds to a unique pixel of the visible light image A. Thus, the reference coordinate system is established, and the visible light coordinate reference group is generated.

[0064] S102: Based on the visible light coordinate reference group, the width and height parameters of the near-infrared and thermal infrared images are called, the width-height ratio between the visible light image and the corresponding image is calculated, coordinate scaling and offset processing is performed, the pixel position index of the corresponding image is adjusted, the mapping relationship between the coordinates is established, and the coordinate mapping relationship group is acquired.

[0065] Based on the visible light coordinate reference group, the width 768 and height 576 parameters of the near-infrared image B are called to calculate the width ratio between the near-infrared image B and the visible light image A , and the height ratio The width 384 and height 288 parameters of the thermal infrared image C are called to calculate the width ratio between the thermal infrared image C and the visible light image A , and the height ratio When performing coordinate scaling and offset processing, for any point in the visible light coordinate reference group , the calculation process of the corresponding mapping coordinates in the near-infrared image B is and For example, the pixel point with coordinates (1000, 800) in the reference group has mapping coordinates (1000 / 2.5, 800 / 1.875) = (400, 426.67) in the near-infrared image. Similarly, the calculation process of the corresponding mapping coordinates in the thermal infrared image C is and The mapping coordinates of the point (1000, 800) in the thermal infrared image are (1000 / 5.0, 800 / 3.75) = (200, 213.33). By traversing all the coordinate points in the visible light coordinate reference group, the pixel position index adjustment and mapping of the near-infrared and thermal infrared images are completed, and the coordinate mapping relationship group is obtained.

[0066] S103: Call the coordinate mapping relationship group and the pixel intensity values of each image, and select the corresponding position surrounding pixel points according to each position under the reference coordinates to perform gray value interpolation filling, integrate the pixel information of all channel images under the unified coordinates, and obtain the image position matching image group.

[0067] Take the visible light reference coordinates (1000, 800) as an example. The mapping coordinates of the near-infrared image are (400, 426.67), which are not integers and cannot be directly extracted as pixel values. Therefore, gray value interpolation filling is performed. Specifically, four surrounding integer pixel points are selected, which are the upper left corner (400, 426), the upper right corner (401, 426), the lower left corner (400, 427), and the lower right corner (401, 427). The gray values of these four pixel points are assumed to be , , , Then, linear interpolation is performed in the x-direction to calculate the grayscale values ​​of the projection points of point (400, 426.67) onto the two horizontal lines y=426 and y=427. Since its x-coordinate is 400 and there is no horizontal offset, the two intermediate values ​​are respectively... and Linear interpolation is performed on these two intermediate values ​​in the y-direction. The decimal part of the target point's y-coordinate is 0.67. Therefore, the interpolated grayscale value is... The gray value 156.7 is assigned to the near-infrared channel at visible light coordinates (1000, 800). The same interpolation and filling operation is performed on the thermal infrared image and all other reference coordinate points. The pixel information of all channel images is integrated under a unified 1920x1080 coordinate system to obtain an image position matching map group.

[0068] Please see Figure 3 The steps for obtaining the channel synchronization feature location set are as follows:

[0069] S201: Based on the position index in the image position matching map group, extract the red and green channel values ​​of the corresponding positions in the visible light image, calculate the numerical difference between the two channels, compare the difference with the channel difference threshold, retain the position index that exceeds the threshold, and obtain the red and green channel difference position set;

[0070] The channel difference threshold is a value dynamically determined based on the statistical distribution of the overall red and green channels in the visible light image;

[0071] Based on the position index in the image position matching map group, a channel difference threshold is set. The threshold is determined as follows: under normal monitoring conditions before the fire, 1000 consecutive frames of visible light images are acquired, and the red channel value of all pixels in each frame is calculated. With green channel value The difference Calculate the global average of all differences across these 1000 frames. and global standard deviation Experiments have verified that, in an office environment, the statistics show... , Set threshold Substituting the data yields This setting ensures that only pixels exceeding 99% of the normal fluctuation range are initially screened out. During detection, the pixel at coordinates (950, 520) in the visible light image is extracted, with a corresponding red channel value of 188 and a green channel value of 135. The difference between the two channel values ​​is calculated. The difference of 53 was compared with the channel difference threshold of 37.3. The position index (950, 520) is reserved, and the same calculation and comparison are performed on all position indexes in the image position matching graph set to obtain the red-green channel difference position set.

[0072] S202: The position index in the red-green channel difference position set is called to extract the gray value of the near-infrared image, the difference amplitude of the gray value relative to the overall gray mean value of the image is calculated, and it is judged whether the difference amplitude exceeds the gray offset threshold value, and the position index meeting the condition is screened to obtain the gray offset position set.

[0073] The gray offset threshold value is a value dynamically set according to the statistical dispersion degree of the gray value in the near-infrared image;

[0074] The position index in the red-green channel difference position set is called, and the gray offset threshold value is set The threshold value is dynamically set according to the statistical dispersion degree of the gray value of the current frame of near-infrared image, and the specific process is that the gray mean value of all 768x576 pixels of the current frame of near-infrared image is calculated and the gray standard deviation , it is assumed that for the current frame of image, the gray mean value is calculated as , , the gray offset threshold value is set as , the setting coefficient 2.5 is the optimal value calibrated in 500 different light conditions of non-flame scene test, which is used to balance the false alarm rate and the false alarm rate, and a position index in the red-green channel difference position set is extracted, for example, the aforementioned (950, 520), the corresponding near-infrared image gray value of the position is found through the image position matching graph set, which is 185.2, and the difference amplitude of the gray value relative to the overall gray mean value of the image is calculated , it is judged whether the difference amplitude exceeds the gray offset threshold value 64.0, because , the position index (950, 520) meets the condition and is screened out, and the process is repeated for all position indexes in the red-green channel difference position set to obtain the gray offset position set.

[0075] S203: According to the index in the gray offset position set, the temperature value of the thermal infrared image is extracted, the temperature is compared with the high position interval composed of the image temperature mean value and the temperature upper limit, and the position index meeting the three-channel offset condition is screened to obtain the channel synchronization feature position set.

[0076] The process of comparing the temperature with the high position interval composed of the image temperature mean value and the temperature upper limit is specifically judging whether the temperature value of the thermal infrared image is greater than the image temperature mean value and less than the temperature upper limit.

[0077] According to the index of the set of gray offset positions, the temperature value of the thermal infrared image is extracted, and the temperature value is compared with a high bit interval, and the construction of the high bit interval depends on the temperature mean value of the current frame thermal infrared image and the preset temperature upper limit , wherein the temperature upper limit According to the effective range and safety specification of the thermal infrared sensor, the sensor range used in this embodiment is-20°C to 400°C, so the high bit interval is set to For the current frame thermal infrared image, the system calculates the temperature mean value of all 384x288 pixels as °C, therefore, the judgment condition of the high bit interval is the temperature value needs to meet and , and the index (950, 520) in the set of gray offset positions is extracted, and the corresponding thermal infrared image temperature value of the image position matching group is 85.6°C, which is compared with the high bit interval, because and , the condition is established, so the position index (950, 520) is confirmed to meet the three-channel offset conditions at the same time, and the position index is retained, and the same temperature comparison screening is performed on all indexes in the set of gray offset positions to obtain the set of channel synchronization feature positions.

[0078] Please refer to Figure 4 The acquisition steps of the fused boundary tracking group are as follows:

[0079] S301: Call all position indexes in the set of channel synchronization feature positions, judge the continuous connection relationship according to the horizontal and vertical adjacency rules, combine the closed regions into the closed block structure in the image, exclude the edge incomplete regions and isolated pixel points, and generate the closed block number set;

[0080] All position indexes in the channel synchronization feature position set are called, for example, the set contains the point set {(950, 520), (951, 520), (950, 521), (966, 530)}, the system uses the four-neighbor rule to judge the continuous connection relationship, that is, whether there is another index in the set in the upper, lower, left and right four directions of each position index is checked, for (950, 520), there is (951, 520) to the right, and (950, 521) below, so the three points are regarded as connected to each other, and (966, 530) is not in the set, if it is an isolated point, it is excluded, by traversing all points in the set, all mutually connected points are divided into the same connected component to form a block, assuming that two blocks are formed, block 1 is composed of 15 continuous pixel points, and block 2 is composed of only point (966, 530), at this time, the system excludes block 2 as an isolated pixel point, and checks whether block 1 contacts the four boundaries of the image, assuming that block 1 is located in the center region of the image and does not contact the boundary, then it is a complete closed block structure, and the system assigns a number "Block_01" to it. The same combination and screening are performed on all connected components composed of channel synchronization feature position sets to generate a closed block number set.

[0081] S302: Based on the closed block number set, the edge pixel gray value of each block in the visible light image is extracted, and a gray direction sequence is generated by arranging the edge direction, the continuity of adjacent gray direction in the sequence is judged, the block number with a direction change amplitude less than the direction stability threshold is screened out, and a gray direction stable block set is obtained;

[0082] Based on the closed block number set, the block with the number "Block_01" is extracted, and its edge pixels are located in the visible light image. The determination method of the edge pixel is that if at least one of the four-neighbor pixels of any pixel in the block does not belong to the block, the pixel is an edge pixel. After extracting all the edge pixels of "Block_01", they are arranged in a clockwise direction to obtain an ordered edge pixel sequence, for example, the sequence is The gray gradient direction of each edge pixel position is calculated, a 3x3 Sobel operator is applied to the neighborhood gray value of the pixel to calculate its horizontal gradient and vertical gradient , and the gradient direction angle .

[0083] Table 1: Edge pixel gradient direction calculation table

[0084]

[0085] As shown in Table 1, the gradient calculation examples of the continuous three pixel points in the edge sequence are listed, and the direction stability threshold The setting of the reference refers to the edge gray direction change statistics of 100 different flame video samples, and the 95% quantile of the change rate is set as The continuity of the adjacent gray direction change in the sequence is judged, that is, the And Because And , it is indicated that the edge direction change of the segment is stable, and if the direction change amplitude of the entire block edge sequence is less than the threshold value, the number "Block_01" is retained, and a gray direction stable block set is obtained.

[0086] S303: According to the block number in the gray direction stable block set, the brightness value distribution of the corresponding region in the near-infrared image is extracted, a brightness trend curve is generated in the horizontal direction, it is judged whether the continuous pixel segment in the curve satisfies the increasing trend condition, and the intersection filtering with the result of the previous sub-step is performed, and a fusible boundary tracking graph group is obtained;

[0087] According to the block number "Block_01" in the gray direction stable block set, the corresponding pixel region in the near-infrared image is extracted, the average brightness value of each row of pixels is calculated by scanning the region row by row in the horizontal direction, and a brightness trend curve is generated. Assuming that the "Block_01" region spans the range from the vertical position To , the average brightness value sequence (i.e. the brightness trend curve) obtained by calculation is , it is judged whether the curve satisfies the continuous increasing trend condition, and the specific judgment process of this condition is as follows: the difference between adjacent elements in the sequence is calculated to obtain the difference sequence The increasing trend condition requires that the proportion of positive elements in the difference sequence exceeds a predetermined proportion, for example, 80%, and there is no case of two consecutive negative or zero differences. In this example, all the differences are positive, satisfying the condition of 100% positive, so it is determined that the brightness trend of the block is continuously increasing. The intersection filtering is performed on the block number and the result of the previous sub-step, and since "Block_01" simultaneously satisfies the gray direction stability and the brightness trend increasing, it is filtered out, thereby obtaining a fusible boundary tracking graph group.

[0088] Please refer to Figure 5 , and the acquisition steps of the heat change aggregation pattern block are as follows:

[0089] S401: In the specified region range in the fusible boundary tracking graph group, a plurality of frame sequence data corresponding to the thermal infrared image is extracted, the temperature values of each frame are recorded in sequence according to the pixel position, the temperature increment trend between consecutive frames is calculated, it is judged whether the increment is continuously positive, the continuously heated pixel position is retained, and the continuously heated pixel distribution value is obtained.

[0090] a process of judging whether the increment is continuously positive, specifically, requiring that the pixel temperature value of each frame of image is higher than the pixel temperature value at the same position of the previous frame of image within a preset number of continuous frames;

[0091] The "Block_01" region is specified in the mergable boundary tracking graph set, and the continuous 5-frame sequence data of the hot infrared image at the corresponding position of the region is extracted, with time stamps being T0, T0+100ms, T0+200ms, T0+300ms, and T0+400ms. The temperature values of each frame are recorded in order according to the pixel position, and the temperature increment between the continuous frames is calculated to judge whether the increment is continuously positive. The judgment process is specifically that the pixel temperature value of each frame of image is required to be higher than the pixel temperature value at the same position of the previous frame of image within a preset number of continuous frames, for example, 4 frames. As shown in Table 2, three pixel points P1, P2, and P3 in the "Block_01" region are selected for display:

[0092] Table 2: Multi-frame pixel temperature monitoring table

[0093]

[0094] As shown in Table 2, for the pixel P1, the continuous 4 increments are +5.5, +6.7, +6.4, and +6.5, all of which are positive values, so P1 is determined to be a continuously rising pixel. For the pixel P2, the temperature 77.9 in the third frame is lower than 78.2 in the second frame, and the increment is -0.3, which does not meet the condition of being continuously positive, so P2 is excluded. For the pixel P3, the increment is also positive, so P3 is determined to be a continuously rising pixel. This judgment is performed on all pixels in the "Block_01" region, and all pixel positions determined to be continuously rising are retained to obtain the distribution value of the continuously rising pixels.

[0095] S402: Based on the distribution value of the continuously rising pixels, the horizontal and vertical adjacent distribution relationship of the continuously rising pixels in the image space is detected, the aggregation density between the adjacent rising pixels is counted, the number of pixels and the distribution range in the continuously connected region are judged, and the region number meeting the aggregation density threshold is screened to obtain an index set of the rising aggregation region;

[0096] The process of judging the number of pixels and the distribution range in the continuously connected region further includes calculating the contour perimeter and the area of the continuously connected region, and screening the continuously connected region according to the ratio of the contour perimeter to the area.

[0097] Based on the distribution value of the continuously increasing temperature pixels, the transverse and longitudinal adjacent distribution relationship in the image space is detected, all the continuously increasing temperature pixels are taken as a set, the connected component analysis of the four adjacent connections is applied again, they are combined into multiple continuous connection regions, and the pixel number and distribution range of each region are judged, and this process further includes calculating the contour perimeter of each continuous connection region With the area , the area The total number of pixels in the region, the contour perimeter The number of pixels on the edge of the region, and the ratio of And The screening region is screened according to the shape factor The value is 1 for a circular region and less than 1 for other shapes, and the aggregation density threshold is set on the shape factor, and the shape factor of 50 groups of real flame samples and 50 groups of interference heat sources (such as light and reflection) samples is statistically analyzed.

[0098] Table 3: Shape factor statistical analysis table

[0099]

[0100] Table 3 lists the experimental statistical data for determining the threshold value, and it is found that the shape factor of the flame region is generally greater than 0.4, while the shape factor of the interference source is mostly below 0.3, in order to achieve better segmentation effect between the two, the aggregation density threshold is set to , assuming that the continuously increasing temperature pixels form a continuous region containing 28 pixels, the contour perimeter is calculated as 20 pixels, and the shape factor is Because The region meets the aggregation density requirement, and its number is reserved to obtain the index set of the aggregation region of the increasing temperature.

[0101] S403: Call the tile number in the index set of the aggregation region of the increasing temperature, extract the closed boundary information of each region in the thermal infrared image, mark the position of the corresponding space region in the image, integrate the boundary framework and position coordinates of all regions, establish the spatial identification layer of the continuous increasing temperature region, and generate the thermal change aggregation graphic block;

[0102] The process of generating the thermal change aggregation graphic block is to combine all the continuously connected regions in the image pixel set, and take a minimum convex hull that can completely cover all the merged pixels as the thermal change aggregation graphic block;

[0103] Call all tile numbers in the temperature rise aggregation region index set, extract the closed boundary information of each region in the thermal infrared image, and combine the pixel sets of all filtered continuous connection regions in the image. Assuming that the temperature rise aggregation region index set contains two regions "Region_A" and "Region_B" that meet the conditions, the system combines all pixel coordinates of the two regions into a large point set, and then calculates a minimum convex hull that can completely cover the merged point set. This process finds a minimum polygon by executing a convex hull algorithm, such as the Graham scan method or the Monotone Chain method, which takes points from the merged point set as vertices, and all points are inside or on the polygon. The region bounded by this convex polygon is the generated heat change aggregation graphic block.

[0104] Please refer to Figure 6 The acquisition step of the flame detection positioning result is:

[0105] S501: Extract all tile corresponding visible light image regions in the heat change aggregation graphic block, call the red and green channel values to calculate the pixel difference value trend, judge whether there is a continuous directional red-green channel difference enhancement phenomenon in the internal pixels of the tile, filter out the region number with sustained offset trend, and obtain the red-green offset trend value;

[0106] Extract all tile corresponding visible light image regions in the heat change aggregation graphic block, call the red and green channel values to calculate the pixel difference value , get a difference value distribution map, judge whether there is a continuous directional red-green channel difference enhancement phenomenon in the internal pixels of the tile, the specific execution process is to calculate the gradient of each pixel on this difference value distribution map. The gradient is a vector containing amplitude and direction, indicating the direction and rate of the fastest difference value change. The system further analyzes the direction distribution of all pixel gradient vectors in the region. If more than 70% of the pixels in a region have gradient directions pointing to a concentrated small range, such as direction angle distribution in interval, it is considered that there is a continuous directional enhancement phenomenon. Assuming that for the current block, the red-green channel difference value of most of the pixels gradually increases from 50 to 80 upwards, forming a gradient field with a relatively consistent direction, the region is considered to have a sustained offset trend, and the corresponding region number is retained. The red-green offset trend value is obtained.

[0107] S502: Based on the region number in the red-green offset trend value, extract the gray value of the corresponding region in the near-infrared image, calculate the region gray mean value, and compare it with the gray high response reference threshold to confirm whether it meets the high response condition. The numbers with gray mean values lower than the threshold are excluded to obtain a high response tile index set.

[0108] Based on the region number in the red-green offset trend value, the gray value of the corresponding region in the near-infrared image is extracted, and the gray mean of the region is calculated. Gray-scale high response reference threshold The setting method is to take the 98th percentile of the grayscale values ​​of all pixels in the current near-infrared image. This method ensures that the threshold can adapt to changes in the overall brightness of the image under different lighting conditions. Assuming that the 98th percentile of the grayscale value distribution of the current near-infrared image is 195.0, therefore... Now, calculate the mean gray value of the regions corresponding to the region numbers selected in the previous steps, and obtain... The mean is compared with the threshold because Once the region is confirmed to meet the high response condition, its number is retained. If the grayscale average of another region is 180.4, which is lower than the threshold, its number will be filtered out. All candidate regions are processed in this way to obtain the high response patch index set.

[0109] S503: Extract the distribution range of heated pixels in the thermal infrared image based on the number in the high-response patch index set, determine whether the heated pixels are concentrated in a single area, filter the patch numbers that simultaneously meet the requirements of red-green channel offset, near-infrared high response and thermal infrared concentrated heating, establish corresponding coordinate markers, and generate flame detection and positioning results.

[0110] Based on the index number of the high-response patch set, the distribution range of heated pixels in the thermal infrared image is extracted, and it is determined whether these heated pixels are concentrated in a single region. This determination is achieved by calculating the standard deviation of the heated pixel coordinates, and the horizontal coordinates of all continuously heated pixels (derived from S401) within that region are calculated respectively. and vertical coordinates Standard deviation and And calculate its spatial discreteness. Simultaneously calculate the dimensions of the region itself, such as the diagonal length of its bounding rectangle. ,like If the concentration coefficient is less than a preset value, such as 0.3, the heating locations are considered to be concentrated. This coefficient of 0.3 is obtained by statistically analyzing 100 flame samples and taking the 90th percentile of the ratio of dispersion to size. For the currently indexed region, if its red-green channel offset trend meets the requirements, the near-infrared grayscale shows a high response state, and its heating pixel distribution is... If the calculation result is 0.21, which is less than 0.3, then the area is identified as a flame. The system establishes the coordinate markers of its bounding rectangle, for example (x:945,y:515,width:30,height:25), and generates the flame detection and localization result under the fused image.

[0111] The above merely provides the specific implementation of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of the changes or replacements within the technical range disclosed by the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A flame detection method based on multimodal fusion images, characterized in that, Includes the following steps: S1: Acquire visible light, near-infrared and thermal infrared images of the same time point in the fire video monitoring area, extract pixel position indexes, establish a unified coordinate group based on the reference image, and generate an image position matching map group by matching the pixel position relationship between different images through coordinate mapping and interpolation. S2: Based on the position index in the image position matching group, extract the red-green difference of the visible light image, the grayscale shift of the near-infrared image, and the temperature distribution of the thermal infrared image, determine whether the corresponding pixel features of the three are shifted at the same time, extract the synchronous shift position, and generate a channel synchronous feature position set. The steps for obtaining the channel synchronization feature location set are as follows: S201: Based on the position index in the image position matching map group, extract the red and green channel values ​​of the corresponding positions in the visible light image, calculate the numerical difference between the two channels, compare the difference with the channel difference threshold, retain the position index that exceeds the threshold, and obtain the red and green channel difference position set. S202: Call the position index in the red-green channel difference position set, extract the gray value of the near-infrared image, calculate the difference magnitude of the difference relative to the overall gray value of the image, determine whether it exceeds the gray value offset threshold, filter out the position index that meets the condition, and obtain the gray value offset position set. S203: Based on the index in the grayscale offset position set, extract the temperature value of the thermal infrared image, compare the temperature with the high-order interval formed by the average temperature and the upper limit of the image temperature, filter out the position index that simultaneously meets the three-channel offset conditions, and obtain the channel synchronization feature position set. S3: Call the channel synchronization feature location set, combine closed regions according to the adjacency rule, extract edge direction changes in the visible light image, extract brightness trends in the near-infrared image, filter out patches with stable structural orientation and continuous brightness trends, and generate a fusionable boundary tracking map group. The steps for obtaining the fusionable boundary tracing map set are as follows: S301: Call all position indices in the channel synchronization feature position set, determine the continuous connection relationship according to the horizontal and vertical adjacency rules, combine the closed regions into a closed tile structure in the image, and generate a closed tile number set after excluding incomplete edge regions and isolated pixels. S302: Based on the closed patch number set, extract the gray values ​​of edge pixels of each patch in the visible light image, arrange them according to the edge direction to generate a gray direction sequence, make a coherence judgment on the adjacent gray change directions in the sequence, and filter out the patch numbers whose direction change amplitude is less than the direction stability threshold to obtain a gray direction stable patch set. S303: Based on the patch number in the gray-scale direction stable patch set, extract the brightness value distribution of the corresponding region in the near-infrared image, generate a brightness trend curve in the horizontal direction, determine whether the continuous pixel segments in the curve meet the increasing trend condition, and perform intersection filtering with the result of the previous sub-step to obtain a fusionable boundary tracking patch group. S4: Extract multi-frame data of thermal infrared image from the fusionable boundary tracking map group, analyze the temperature change trend of each pixel, identify continuous heating points, determine whether they are spatially clustered to form a region, extract the concentrated heating structure, and generate thermal change clustering graphic blocks. The steps for obtaining the thermal change clustered graphic block are as follows: S401: Within a specified area in the fusionable boundary tracking map group, extract multi-frame sequence data corresponding to the thermal infrared image, record the temperature value of each frame in sequence according to the pixel position, calculate the temperature increment trend between consecutive frames, determine whether the increment is continuous and positive, retain the pixel position of continuous temperature rise, and obtain the pixel distribution value of continuous temperature rise. S402: Based on the continuously heated pixel distribution value, detect its horizontal and vertical adjacent distribution relationship in the image space, count the cluster density between adjacent heated pixels, judge the number and distribution range of pixels in the continuous connected area, filter out the area number that meets the cluster density threshold, and obtain the heated cluster area index set. S403: Call the tile number in the index set of the heating accumulation area, extract the closed boundary information of each area in the thermal infrared image, mark the position of the corresponding spatial area in the image, integrate the boundary frame and position coordinates of all areas, establish a spatial identification layer of the continuous heating area, and generate thermal change accumulation graphic blocks. S5: Extract the red and green color shift trends in the visible light images of all blocks in the thermal change clustering graphic block, analyze the color shift of the visible light image, the grayscale response of the near-infrared image, and the distribution of the heating position in the thermal infrared image, determine whether flame features are present at the same time, mark the areas that meet the conditions, and generate flame detection and localization results under the fused image. The steps for obtaining the flame detection and positioning results are as follows: S501: Extract the visible light image regions corresponding to all blocks in the thermal change clustering graphic block, call the red and green channel values ​​to calculate the pixel difference change trend, determine whether there is a continuous red and green channel difference enhancement phenomenon in the pixels inside the block, filter out the region numbers with a continuous offset trend, and obtain the red and green offset trend value. S502: Based on the region number in the red-green offset trend value, extract the gray value of the corresponding region in the near-infrared image, calculate the region gray mean, and compare it with the gray high response reference threshold to confirm whether the high response condition is met. Remove the numbers with gray mean values ​​lower than the threshold to obtain the high response patch index set. S503: Extract the distribution range of heated pixels in the thermal infrared image based on the number in the high-response patch index set, determine whether the heated pixels are concentrated in a single area, filter the patch numbers that simultaneously meet the requirements of red-green channel offset, near-infrared high response and thermal infrared concentrated heating, establish corresponding coordinate markers, and generate flame detection and positioning results.

2. The flame detection method based on multimodal fusion image according to claim 1, characterized in that: The image location matching map set includes time tags, size parameters, and pixel location indexes for visible light images, near-infrared images, and thermal infrared images; the channel synchronization feature location set includes red-green channel differences, near-infrared image grayscale shift, thermal infrared image temperature values, and synchronization deviation location set; the fusionable boundary tracking map set includes patch edge grayscale direction sequence, brightness trend curve, direction stability, and brightness consistency region. The thermal change clustering graphic block includes continuous inter-frame temperature change trends, clustering regions, and pixels that are continuously heating up. The flame detection and localization results include the location of regions that match flame characteristics and the flame localization in the fused image.

3. The flame detection method based on multimodal fusion image according to claim 1, characterized in that, The steps for obtaining the image location matching map group are as follows: S101: Acquire visible light, near-infrared and thermal infrared images of the same time in the fire monitoring area, read the time label and size parameters of each type of image, extract the horizontal and vertical indexes of the image pixels, construct a set of coordinate indexes for the visible light images based on the time label matching, and establish a reference coordinate system based on the width and height in its size parameters to generate a visible light coordinate reference group. S102: Based on the visible light coordinate reference group, call the width and height parameters of the near-infrared and thermal infrared images, calculate the aspect ratio with the visible light image, perform coordinate scaling and offset processing, adjust the pixel position index of the corresponding image, establish the mapping relationship between coordinates, and obtain the coordinate mapping relationship group; S103: Call the coordinate mapping relationship group and the pixel intensity value of each image, select the surrounding pixels of the corresponding position according to each position under the reference coordinates, perform grayscale value interpolation and fill, integrate the pixel information of all channel images under the same coordinates, and obtain the image position matching map group.

4. The flame detection method based on multimodal fusion image according to claim 1, characterized in that, The channel difference threshold is a value dynamically determined based on the statistical distribution of the overall red and green channels of the visible light image; The grayscale offset threshold is a value dynamically set based on the statistical dispersion of grayscale values ​​in the near-infrared image. The process of comparing the temperature with the high-order interval formed by the average temperature of the image and the upper limit of the temperature specifically involves determining whether the temperature value of the thermal infrared image is greater than the average temperature of the image and less than the upper limit of the temperature.

5. The flame detection method based on multimodal fusion image according to claim 1, characterized in that, The process of determining whether the increment is continuously positive specifically requires that within a consecutive preset number of frames, the pixel temperature value of each frame image is higher than the pixel temperature value at the same position in the previous frame image. The process of judging the number and distribution range of pixels within a continuous connected region further includes calculating the perimeter and area of ​​the continuous connected region, and filtering the continuous connected region based on the ratio of the perimeter to the area. The process of generating the thermal change cluster graphic block specifically involves merging all the selected continuous connected regions in the image into a single pixel set, and using the smallest convex hull that can completely cover all the merged pixels as the thermal change cluster graphic block.

Citation Information

Patent Citations

  • Forest fire detection method based on near-infrared and thermal imaging image fusion

    CN116403160A

  • Multi-mode indoor fire identification method, device and equipment and storage medium

    CN119380156A