Multi-mode-based fire rescue detection method and system
Through the fire rescue detection method of multimodal information fusion, combined with infrared, visual and audio information, the shortcomings of single-modal detection are solved, more efficient positioning and temperature prediction of trapped people are achieved, and the accuracy and safety of rescue are improved.
Patent Information
- Application Number
- CN202510504592.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-08-05
AI Technical Summary
Existing fire detection tools are based only on single-modal imaging and cannot accurately and efficiently locate trapped people in harsh smoke environments. The model recognition recall rate is low, making it difficult to meet the needs of fast and efficient rescue.
Multimodal fire rescue detection method is adopted, combining infrared images, visual images and multi-channel audio information to extract and fusion multimodal feature, use optical flow prediction to provide prior information, output target position search prompts and temperature prediction, and assist rescue decisions.
It improves the identification accuracy and recall rate of the location information of trapped people, avoids the impact of building shading and smoke, and enhances the safety and efficiency of rescue.
Smart Images

Figure CN120431314A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology and provides a multi-modal fire rescue detection method and system. Background Art
[0002] With the acceleration of urbanization, fire scene rescue operations are facing higher demands. Traditional rescue methods face significant challenges during disaster relief and are unable to meet the current demand for rapid and efficient rescue. The rescue process requires immediate identification of trapped individuals and a swift decision-making process. Any delays in response and misjudgments during the rescue process will inevitably result in significant economic losses and casualties. At fire scenes, dense smoke can make it difficult to immediately identify trapped individuals, and obstruction by building walls can complicate the search. Of course, rescuers must not only rescue trapped individuals; their own protection is also paramount. Therefore, sensing changes in ambient temperature during the rescue process and providing them with information for reference is crucial.
[0003] In the existing technology, there are many auxiliary fire detection tools, such as fire temperature monitoring based on infrared thermal imaging, fire prevention detection based on surveillance cameras, etc.; Chinese Patent Publication No. CN118470342A - "Fire Detection Method, Device and Computer Equipment", this patent obtains a first detection target by passing a real-time image into a pre-trained fire detection model; calculates the relative position information of the second detection target in one or more historical images and the first detection target respectively; based on the relative position information, obtains a judgment result on whether the first detection target is a false detection; based on the judgment result, determines the fire detection result. This method requires fire detection through multiple fusion methods, and the model reasoning has high redundancy, which is not suitable for the demand for real-time fire prediction results. At the same time, fire monitoring based solely on image target detection is difficult to identify the position of people in a large smoke background environment, and the model recognition recall rate is low. Chinese patent CN202410547759.5, "Fire Detection and Alert Method, Device, and System for Bank Service Lobbies," uses visible and infrared cameras to capture visible light images of flames, smoke, and faces, providing a machine vision-based flame and smoke detection system method that can accurately and efficiently detect trapped individuals. This solution uses only infrared and camera vision modalities for personnel monitoring, making it difficult to monitor obstructions such as baffles or walls, significantly inconvenient for search and rescue personnel.
[0004] In summary, existing fire detection tools rely solely on infrared thermal imaging for environmental imaging, failing to accurately and efficiently assist firefighters in quickly locating trapped individuals through detection algorithms. Furthermore, many algorithms rely solely on detection for locating individuals, failing to account for harsh smoke and low-definition fire conditions, resulting in poor detection recall rates. Improvements are necessary. Summary of the Invention
[0005] The purpose of the present invention is to provide a multimodal fire rescue detection method and system to solve the problems raised by the above-mentioned background technology; the present invention not only performs target detection based on a single-frame image, but also performs optical flow prediction based on a continuous sequence, providing prior information for detection, capturing dynamic objects in the video, and improving the overall accuracy and recall rate of detection.
[0006] The first aspect of the present invention is achieved by providing a multimodal fire rescue detection method, the method comprising:
[0007] Acquiring relevant information collected in the rescue area, including infrared image information, visual image information, and multi-channel audio information;
[0008] Perform multimodal feature extraction: perform infrared dynamic and static feature extraction on the infrared image information to obtain infrared feature F cbb 、Infrared characteristics F cbi And the fire temperature distribution characteristics F uc ; Extract visual dynamic and static features of the visual image information to obtain optical flow features F raft and visual features F ffm ; Extract audio modal features from the multi-channel audio information to obtain speech features F KGS ;
[0009] Select multimodal feature fusion: Infrared feature F cbb With visual features F ffm After convolution and activation operations, the first fusion feature F of weighted fusion is obtained. cbff , and the first fusion feature F cbff With infrared characteristics F cbi Perform channel splicing and fusion, and then deconvolve to obtain the second fusion feature F cbffd ; The optical flow feature F raft and speech feature F kgs Fusion, and then with the second fusion feature F cbffd Add together to output prior information;
[0010] The prior information is deconvolved and up-sampled to output the target location search prompt; the fire temperature distribution feature F ucMultiply it by the calibrated temperature range to get the actual temperature at the next moment to assist in rescue.
[0011] Optionally, the infrared image information is subjected to infrared dynamic and static feature extraction to obtain infrared feature F cbb 、Infrared characteristics F cbi And the fire temperature distribution characteristics F uc , specifically including:
[0012] Infrared static feature extraction: The infrared image information is input into the FastNet Backbone network, RepVGG network and BiseNet network to obtain the infrared feature F fastnet 、Infrared characteristics F RepVGG and infrared signature F bi ; The infrared feature F fastnet After the deformable attention mechanism is operated and combined with the infrared feature F RepVGG Perform channel splicing and fusion to obtain infrared features F fr ; The infrared feature F fr Input to the convolution module and combine the output infrared features with the infrared features F bi Multiply to get the infrared feature F cbb ;
[0013] The infrared signature F bi Input to the convolution module to obtain infrared features F cbi ;
[0014] Infrared dynamic feature extraction: For the infrared image information, the infrared temperature of the next frame is predicted by extracting the infrared temperature of the consecutive frames. Each frame module is feature extracted through the vit block architecture, and the features of each frame are input into the Bi-LSTM model. The obtained features are further input into the deconvolution module to obtain the infrared feature F idv Then, the next frame fire temperature distribution feature F is finally obtained by connecting the convolution module and the Sigmoid module in series. uc .
[0015] Optionally, the visual image information is subjected to visual dynamic and static feature extraction to obtain optical flow features F raft and visual features F ffm , specifically including:
[0016] Visual dynamic feature extraction: Input the visual image information into the optical flow feature extraction module to obtain the optical flow feature F raft ;
[0017] Visual static feature extraction: Use the yolov11 backbone model and the mobileformer model to extract static image features from visual image information and obtain visual features F yb and visual features F mb ;
[0018] The visual feature F yb Input to the FAM module to obtain visual features F fam ;
[0019] The visual feature F mb Input GAM module to obtain visual features F CBRA , then input into the FAM module and output, and combined with the visual feature F fam Input them into the FFM module together to obtain the visual feature F ffm .
[0020] Optionally, the visual feature F mb Input GAM module to obtain visual features F CBRA , specifically including:
[0021] The visual feature F mb Input into the global maximum pooling and global average pooling operations respectively, and add the results of the two operations and perform convolution operations to obtain the visual feature F CBR ; At the same time, the visual feature F CBR Input to the 1x1 convolution module, LeakeyReLU activation function module, 1x1 convolution module and sigmoid activation function module to obtain the visual feature F CBRS , the visual feature F CBRS With visual features F CBR Multiply to get the visual feature F CBRSM ; The visual feature F CBRSM With visual features F CBR Add up to get the visual feature F CBRA .
[0022] Optionally, the multi-channel audio information is subjected to audio modal feature extraction to obtain speech feature F KGS , specifically including:
[0023] Process multi-channel audio information in segments;
[0024] Input different segments of continuous audio into the Fourier Transformer model to obtain speech features F FT , and the speech feature F FTThe speech features are input into the MFCCs module respectively, and the speech features obtained are input into the ResNet-Conformer model and the Mamba Modules model respectively to obtain the speech features F rc And the speech feature F voice_mamba ;
[0025] The speech feature F voice_mamba and speech feature F rc Perform channel splicing and fusion to obtain speech feature F rcv , and the speech feature F rc Input to the 1x1 convolution module and the sigmoid activation function module to obtain the attention feature, the attention feature and the speech feature F rcv Multiply to extract the key feature information of the audio and obtain the speech feature F through convolution operation rcvc ; For the speech feature F rcvc Perform time pooling, and then obtain SED information and DOA information through the KNN model of the two branches respectively. The SED information is the sound category, and the DOA information is the sound source location and is displayed in the form of a Gaussian graph; finally, the DOA information is multiplied by the SED information to obtain the speech feature F KGS ;
[0026] Voice feature F KGS Represents the sound position of the real sound source and the corresponding sound confidence output.
[0027] A second aspect of the present invention provides a multimodal fire rescue detection system for use in the method described above, the system comprising:
[0028] A data acquisition module is used to obtain relevant information collected in the rescue area, including infrared image information, visual image information and multi-channel audio information;
[0029] The multimodal feature extraction module is used to extract multimodal features: extract the infrared dynamic and static features of the infrared image information to obtain the infrared feature F cbb 、Infrared characteristics F cbi And the fire temperature distribution characteristics F uc ; Extract visual dynamic and static features of the visual image information to obtain optical flow features F raft and visual features F ffm ; Extract audio modal features from the multi-channel audio information to obtain speech features F KGS ;
[0030] Multimodal feature fusion module, used to select multimodal feature fusion: infrared feature F cbb With visual features F ffmAfter convolution and activation operations, the first fusion feature F of weighted fusion is obtained. cbff , and the first fusion feature F cbff With infrared characteristics F cbi Perform channel splicing and fusion, and then deconvolve to obtain the second fusion feature F cbffd ; The optical flow feature F raft and speech feature F kgs Fusion, and then with the second fusion feature F cbffd Add together to output prior information;
[0031] The detection result output module is used to output the target location search prompt through deconvolution and upsampling operations on the prior information; the fire temperature distribution feature F uc Multiply it by the calibrated temperature range to get the actual temperature at the next moment to assist in rescue.
[0032] The present invention provides a multimodal fire rescue detection method. The acquired multimodal information combines voice features, infrared features, and visual features, and performs sufficient feature fusion, thereby improving the accuracy and recall rate of identifying the location information of trapped persons. At the same time, the addition of the audio modality allows the detection method to avoid the limitations of the visual modality that cannot be detected due to building obstructions and smoke, making the detection effect more robust. The present invention not only combines single-frame information, but also adds continuous information of sequence frames. In the visual modality, based on dynamic features, it can assist in discovering moving objects. At the same time, based on the discovered static objects, the combination of the two provides better support for identifying trapped persons. In the infrared modality, based on continuous frame information, the temperature change and azimuth change trend of the fire spread can be predicted, which facilitates rescue personnel to make better rescue decisions and improves the safety of the rescue. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 A general flow chart of the principle of a multi-modal fire rescue detection method provided by an embodiment of the present invention;
[0034] Figure 2 Flowchart of a multi-modal fire rescue detection method according to one embodiment;
[0035] Figure 3 is a flow chart of an infrared feature extraction module in one embodiment;
[0036] Figure 4 is a flow chart of a visual feature extraction module in one embodiment;
[0037] Figure 5 is a flow chart of a speech feature extraction module in one embodiment;
[0038] Figure 6A flowchart of a multi-modal fire rescue detection method provided by an embodiment of the present invention;
[0039] Figure 7 A structural block diagram of a multimodal fire rescue detection system provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0040] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0041] like Figures 1-6 As shown, in one embodiment, a multimodal fire rescue detection method comprises the following steps S101 to S104;
[0042] S101: Acquire relevant information collected in the rescue area, including infrared image information, visual image information, and multi-channel audio information;
[0043] In this step, infrared image information can be collected by an infrared imager or infrared camera, and the data can be transmitted to a computer device or a portable laptop; visual image information can be collected by a camera, a camera or a mobile phone, and multi-channel audio information can be collected by one or more pickups or microphones, and the data can be transmitted to a computer device or a portable laptop to achieve data acquisition.
[0044] Exemplarily, relevant information can be acquired in three modalities and two modes, namely: infrared mode, video mode and voice mode, single frame mode and sequence frame mode; feature acquisition and processing of three modalities and two modes, fusion detection of multi-modal information, improve search and rescue efficiency, quickly locate the position of trapped people in distress, and help to quickly extinguish fire accidents.
[0045] S102: Perform multimodal feature extraction: perform infrared dynamic and static feature extraction on the infrared image information to obtain infrared feature F cbb 、Infrared characteristics F cbi And the fire temperature distribution characteristics F uc ; Extract visual dynamic and static features of the visual image information to obtain optical flow features F raft and visual features F ffm ; Extract audio modal features from the multi-channel audio information to obtain speech features F KGS ;
[0046] S103: Select multimodal feature fusion: Infrared feature F cbb With visual features F ffmAfter convolution and activation operations, the first fusion feature F of weighted fusion is obtained. cbff , and the first fusion feature F cbff With infrared characteristics F cbi Perform channel splicing and fusion, and then deconvolve to obtain the second fusion feature F cbffd ; The optical flow feature F raft and speech feature F kgs Fusion, and then with the second fusion feature F cbffd Add together to output prior information;
[0047] S104: Deconvolution and upsampling are performed on the prior information to output the target location search prompt; the fire temperature distribution feature F uc Multiply it by the calibrated temperature range to get the actual temperature at the next moment to assist in rescue.
[0048] In this embodiment, the detection process is not only based on single-frame image detection, but also based on continuous sequence optical flow prediction to provide prior information for detection; it helps to capture dynamic objects in the video (represented by visual images) and improves the overall accuracy and recall rate of detection. The feature fusion method adopted is the fusion of visual modality and sound modality, which makes the extracted features richer and improves the accuracy of prediction. Moreover, combined with the use of sound modality, it is beneficial for rescuers to still obtain the distress signals of trapped persons after being blocked by walls, etc., accurately locate rescuers, and provide great support for rescue work. At the same time, the temperature changes of the environmental area can be predicted during detection. It is convenient for rescuers, not only to achieve better rescue of people, but also to better protect themselves, thereby improving the safety of rescuers.
[0049] In step S102, the infrared image information is subjected to infrared dynamic and static feature extraction to obtain infrared feature F cbb 、Infrared characteristics F cbi And the fire temperature distribution characteristics F uc , specifically including:
[0050] Infrared static feature extraction: The infrared image information (specifically the infrared thermal map contained in the infrared image information) is input into the FastNet Backbone network, RepVGG network and BiseNet network to obtain the infrared feature F fastnet 、Infrared characteristics F RepVGG and infrared signature F bi ; The infrared feature F fastnet After the deformable attention mechanism operation (Deformable Attention) and the infrared feature F RepVGG Perform channel splicing and fusion (concate merging) to obtain infrared features F fr, realize feature alignment and fusion; the infrared feature F fr Input to the convolution module and combine the output infrared features with the infrared features F bi Multiply to get the infrared feature F cbb Among them, the FastNet Backbone network and the RepVGG network are efficient transformer models and CNN convolution deformation operators; the features obtained by the two feature extraction methods are concatenated and fused (or channel splicing and fusion) to obtain the thermal image feature F r At the same time, a semantic segmentation network is extracted based on BiseNet (i.e., bilateral segmentation network) to obtain the features after semantic segmentation. This feature mainly represents the semantic segmentation mask of the target object (people and pets). An auxiliary loss function, the MSE loss function, can be added to assist in network model training, enabling the model to predict more accurate target masks.
[0051] The infrared signature F bi Input to the convolution module to obtain infrared features F cbi ; Among them, the convolution module consists of a convolution operation layer, an activation function operation layer, and a BatchNorm operation layer.
[0052] Infrared dynamic feature extraction: For the infrared image information, the infrared temperature of the next frame is predicted by extracting the infrared temperature of the consecutive frames. Each frame module is feature extracted through the vit block architecture, and the features of each frame are input into the Bi-LSTM model. The obtained features are further input into the deconvolution module (DeConv) to obtain the infrared feature F idv , and then through the series convolution module and the Sigmoid module (also called Sigmoid activation function module, or Sigmoid) to finally get the next frame of fire temperature distribution feature F uc Fire temperature distribution characteristics F uc The characteristic range is infrared temperature of 0-1. It should be emphasized that the interval between frames is 1 second, and the loss function used is DICE loss function. The final predicted temperature is obtained by the fire temperature distribution feature F uc Multiply it by the temperature range to get the actual temperature at the next second, which makes it easier for rescuers to choose rescue measures based on the dynamics of the fire.
[0053] In step S102, the visual image information is subjected to visual dynamic and static feature extraction to obtain the optical flow feature F raft and visual features F ffm , specifically including: visual dynamic feature extraction, visual static feature extraction;
[0054] For example, visual dynamic feature extraction can be achieved through a dynamic feature extraction module, which is mainly responsible for capturing the position of moving objects in the visual scene; visual static feature extraction can be achieved through a static feature extraction module, which samples small lightweight feature extraction modules, namely the yolov11 backbone model and the mobileformer model. The static feature extraction module is mainly responsible for extracting the actual position of the stationary target object. The feature extraction module of the dynamic target object here uses the RAFT optical flow extraction model, and inputs the multi-frame video RGB image segment information (i.e., visual image information) into the RAFT optical flow extraction model to obtain the feature F raft The model uses a pre-trained model and can obtain the optical flow feature F required by this embodiment without training and direct inference. raft The two methods of yolov11 backbone model and mobileformer model extract richer feature information.
[0055] For example, visual dynamic feature extraction: the visual image information is input into the optical flow feature extraction module to obtain the optical flow feature F raft ; Visual static feature extraction: Use the yolov11 backbone model and the mobileformer model to extract static image features from visual image information and obtain visual features F yb and visual features F mb ;
[0056] The visual feature F yb Input to the FAM module to obtain visual features F fam ; The FAM module is composed of a global maximum pooling layer, a 1x1 convolution layer, and a BN and sigmoid layer in series (or a global maximum pooling module, a 1x1 convolution module, and a BN and sigmoid module in series).
[0057] The visual feature F mb Input GAM module to obtain visual features F CBRA , then input into the FAM module and output, and combined with the visual feature F fam Input them into the FFM module together to obtain the visual feature F ffm ; This feature has a strong target location attribute feature.
[0058] In fact, the visual feature F mb Input GAM module to obtain visual features F CBRA , specifically including:
[0059] The visual feature F mbInput into the global maximum pooling and global average pooling operations respectively, and add the results of the two operations and perform convolution operations to obtain the visual feature F CBR ; At the same time, the visual feature F CBR Input to the 1x1 convolution module, LeakeyReLU activation function module, 1x1 convolution module and sigmoid activation function module to obtain the visual feature F CBRS , the visual feature F CBRS With visual features F CBR Multiply to get the visual feature F CBRSM ; The visual feature F CBRSM With visual features F CBR Add up to get the visual feature F CBRA The introduced residual mechanism can prevent the model memory from decaying.
[0060] In steps S102 and S103, the audio modal features of the multi-channel audio information are extracted to obtain the speech features F KGS , specifically including:
[0061] The multi-channel audio information is segmented and processed; the length of each audio segment corresponds to the segment length of the visual and infrared modes.
[0062] Input different segments of continuous audio into the Fourier Transformer model to obtain speech features F FT , and the speech feature F FT The speech features are input into the MFCCs module respectively, and the speech features obtained are input into the ResNet-Conformer model and the Mamba Modules model respectively to obtain the speech features F rc And the speech feature F voice_mamba ;
[0063] The speech feature F voice_mamba and speech feature F rc Perform channel splicing and fusion to obtain speech feature F rcv , and the speech feature F rc Input to the 1x1 convolution module and the sigmoid activation function module to obtain the attention feature, the attention feature and the speech feature F rcv Multiply to extract the key feature information of the audio and obtain the speech feature F through convolution operation rcvc ; For the speech feature F rcvc Perform time pooling, and then obtain SED information and DOA information through the KNN model of the two branches respectively. The SED information is the sound category, and the DOA information is the sound source location and is displayed in the form of a Gaussian graph; finally, the DOA information is multiplied by the SED information to obtain the speech feature F KGS; Among them, by multiplying the DOA information with the SED information, the true sound source's sound position and the corresponding sound confidence output feature can be obtained, that is, the speech feature F KGS .
[0064] Voice feature F KGS , which is consistent with the spatial position of the visual feature map, and the value is also 0-1, representing the confidence of the sound source. The higher the score, the greater the probability of the target distress call; it represents the sound position of the real sound source and the corresponding sound confidence output.
[0065] For example, for DOA information and SED information, the classification SED loss function used is the cross entropy loss function, and the DOA loss function is the smooth L1 loss function.
[0066] In this embodiment, feature extraction is combined with the audio modality, which can better solve the problem of the visual modality being unable to quickly identify the location information of trapped persons due to building occlusion.
[0067] In step S103, when the optical flow and sound feature information are fused, the optical flow feature F raft and speech feature F kgs The features are fully fused by multiplication, because both represent the features of a continuous sequence. The continuous sequence features are then combined with the second fusion feature F cbffd Adding is equivalent to fusing static information with dynamic information to give the features more prior information.
[0068] like Figure 7 As shown, in another embodiment, a multimodal fire rescue detection system is used in the method described above, and the system 100 includes:
[0069] The data acquisition module 110 is used to acquire relevant information collected in the rescue area, including infrared image information, visual image information and multi-channel audio information;
[0070] The multimodal feature extraction module 120 is used to extract multimodal features: extract the infrared dynamic and static features of the infrared image information to obtain the infrared feature F cbb 、Infrared characteristics F cbi And the fire temperature distribution characteristics F uc ; Extract visual dynamic and static features of the visual image information to obtain optical flow features F raft and visual features F ffm ; Extract audio modal features from the multi-channel audio information to obtain speech features F KGS ;
[0071] The multimodal feature fusion module 130 is used to select multimodal feature fusion: the infrared feature F cbb With visual features F ffm After convolution and activation operations, the first fusion feature F of weighted fusion is obtained. cbff , and the first fusion feature F cbff With infrared characteristics F cbi Perform channel splicing and fusion, and then deconvolve to obtain the second fusion feature F cbffd ; The optical flow feature F raft and speech feature F kgs Fusion, and then with the second fusion feature F cbffd Add together to output prior information;
[0072] The detection result output module 140 is used to output the target location search prompt by deconvolution and upsampling the prior information; uc Multiply it by the calibrated temperature range to get the actual temperature at the next moment to assist in rescue.
[0073] This embodiment provides a multimodal fire rescue detection method, and based on this method, provides a multimodal fire rescue detection system. During the detection process, the method obtains multimodal information that combines voice features, infrared features, and visual features, and performs sufficient feature fusion, thereby improving the accuracy and recall rate of identifying the location information of trapped persons. At the same time, the addition of audio modality allows the detection method to avoid the limitations of visual modality that cannot be detected due to building obstructions and smoke, making the detection effect more robust. This method not only combines single-frame information, but also adds continuous information of sequence frames. In the visual modality, dynamic features can assist in discovering moving objects, and based on the discovered static objects, the combination of the two provides better support for identifying trapped persons. In the infrared modality, based on continuous frame information, the temperature change and azimuth change trend of the fire spread can be predicted, which facilitates rescue personnel to make better rescue decisions and improves the safety of the rescue.
[0074] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
[0075] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A multimodal fire rescue detection method, characterized in that: The method comprises: Acquiring relevant information collected in the rescue area, including infrared image information, visual image information, and multi-channel audio information; Perform multimodal feature extraction: extract the infrared dynamic and static features of the infrared image information to obtain the infrared feature F cbb 、Infrared characteristics F cbi And the fire temperature distribution characteristics F uc ; Extract visual dynamic and static features of the visual image information to obtain optical flow features F raft and visual features F ffm ; Extract audio modal features from the multi-channel audio information to obtain speech features F KGS ; Select multimodal feature fusion: Infrared feature F cbb With visual features F ffm After convolution and activation operations, the first fusion feature F of weighted fusion is obtained. cbff , and the first fusion feature F cbff With infrared characteristics F cbi Perform channel splicing and fusion, and then deconvolve to obtain the second fusion feature F cbffd ; The optical flow feature F raft and speech feature F kgs Fusion, and then with the second fusion feature F cbffd Add together to output prior information; The prior information is deconvolved and up-sampled to output the target location search prompt; the fire temperature distribution feature F uc Multiply it by the calibrated temperature range to get the actual temperature at the next moment to assist in rescue.
2. The method according to claim 1, characterized in that The infrared image information is subjected to infrared dynamic and static feature extraction to obtain infrared feature F cbb 、Infrared characteristics F cbi And the fire temperature distribution characteristics F uc , specifically including: Infrared static feature extraction: The infrared image information is input into the FastNet Backbone network, RepVGG network and BiseNet network to obtain the infrared feature F fastnet 、Infrared characteristics F RepVGG and infrared signature F bi ; The infrared feature F fastnet After the deformable attention mechanism is operated and combined with the infrared feature F RepVGG Perform channel splicing and fusion to obtain infrared features F fr ; The infrared feature F fr Input to the convolution module and combine the output infrared features with the infrared features F bi Multiply to get the infrared feature F cbb ; The infrared signature F bi Input to the convolution module to obtain infrared features F cbi ; Infrared dynamic feature extraction: For the infrared image information, the infrared temperature of the next frame is predicted by extracting the infrared temperature of the consecutive frames. Each frame module is feature extracted through the vit block architecture, and the features of each frame are input into the Bi-LSTM model. The obtained features are further input into the deconvolution module to obtain the infrared feature F idv Then, the next frame fire temperature distribution feature F is finally obtained by connecting the convolution module and the Sigmoid module in series. uc .
3. The method according to claim 1, characterized in that The visual image information is subjected to visual dynamic and static feature extraction to obtain the optical flow feature F raft and visual features F ffm , specifically including: Visual dynamic feature extraction: Input the visual image information into the optical flow feature extraction module to obtain the optical flow feature F raft ; Visual static feature extraction: Use the yolov11 backbone model and the mobileformer model to extract static image features from visual image information and obtain visual features F yb and visual features F mb ; The visual feature F yb Input to the FAM module to obtain the visual feature F fam ; The visual feature F mb Input GAM module to obtain visual features F CBRA , then input into the FAM module and output, and combined with the visual feature F fam Input them into the FFM module together to obtain the visual feature F ffm .
4. The method according to claim 3, characterized in that The visual feature F mb Input GAM module to obtain visual features F CBRA , specifically including: The visual feature F mb Input into the global maximum pooling and global average pooling operations respectively, and add the results of the two operations and perform convolution operations to obtain the visual feature F CBR ; At the same time, the visual feature F CBR Input to the 1x1 convolution module, LeakeyReLU activation function module, 1x1 convolution module and sigmoid activation function module to obtain the visual feature F CBRS , the visual feature F CBRS With visual features F CBR Multiply to get the visual feature F CBRSM ; The visual feature F CBRSM With visual features F CBR Add up to get the visual feature F CBRA .
5. The method according to claim 1, characterized in that The audio modal feature extraction is performed on the multi-channel audio information to obtain the speech feature F KGS , specifically including: Process multi-channel audio information in segments; Input different segments of continuous audio into the Fourier Transformer model to obtain speech features F FT , and the speech feature F FT The speech features are input into the MFCCs module respectively, and the speech features obtained are input into the ResNet-Conformer model and the Mamba Modules model respectively to obtain the speech features F rc And the speech feature F voice_mamba ; The speech feature F voice_mamba and speech feature F rc Perform channel splicing and fusion to obtain speech feature F rcv , and the speech feature F rc Input to the 1x1 convolution module and the sigmoid activation function module to obtain the attention feature, the attention feature and the speech feature F rcv Multiply to extract the key feature information of the audio and obtain the speech feature F through convolution operation rcvc ; For the speech feature F rcvc Perform time pooling, and then obtain SED information and DOA information through the KNN model of the two branches respectively. The SED information is the sound category, and the DOA information is the sound source location and is displayed in the form of a Gaussian graph; finally, the DOA information is multiplied by the SED information to obtain the speech feature F KGS ; Voice feature F KGS Represents the sound position of the real sound source and the corresponding sound confidence output.
6. A multimodal fire rescue detection system, used in the method according to any one of claims 1 to 5, characterized in that: The system comprises: A data acquisition module is used to obtain relevant information collected in the rescue area, including infrared image information, visual image information and multi-channel audio information; The multimodal feature extraction module is used to extract multimodal features: extract the infrared dynamic and static features of the infrared image information to obtain the infrared feature F cbb 、Infrared characteristics F cbi And the fire temperature distribution characteristics F uc ; Extract visual dynamic and static features of the visual image information to obtain optical flow features F raft and visual features F ffm ; Extract audio modal features from the multi-channel audio information to obtain speech features F KGS ; Multimodal feature fusion module, used to select multimodal feature fusion: infrared feature F cbb With visual features F ffm After convolution and activation operations, the first fusion feature F of weighted fusion is obtained. cbff , and the first fusion feature F cbff With infrared characteristics F cbi Perform channel splicing and fusion, and then deconvolve to obtain the second fusion feature F cbffd ; The optical flow feature F raft and speech feature F kgs Fusion, and then with the second fusion feature F cbffd Add together to output prior information; The detection result output module is used to output the target location search prompt through deconvolution and upsampling operations on the prior information; the fire temperature distribution feature F uc Multiply it by the calibrated temperature range to get the actual temperature at the next moment to assist in rescue.
Citation Information
Patent Citations
Flame and smoke detection system and method based on machine vision
CN118397785A
Fire behavior detection method and device and computer equipment
CN118470342A
Cited By
Emergency rescue real-time human body detection method and equipment based on time sequence motion feature enhancement
CN121305671A