A Fire Detection and Early Warning Method Based on Attention Mechanism and Multi-Scale Features
The Fire-YOLOv5 network with adjustable depth and width, combined with attention mechanisms and multi-scale features, enhances fire detection accuracy and flexibility across hardware platforms for early fire warnings.
Patent Information
- Application Number
- CN202310003454.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-03
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-01-03
AI Technical Summary
The existing fire detection technology based on vision sensors is insufficient in the detection of large and medium-sized flame and smoke targets, and the deep neural network cannot be flexibly adjusted to adapt to different hardware devices, resulting in low detection accuracy and high error detection rate.
The dynamic neural network Fire-YOLOv5 with variable depth and width is adopted, combining the isometric attention mechanism and multi-scale feature fusion, enhances the weight representation of the target position, and realizes early fire warning through the video frame voting mechanism.
It improves the average accuracy of flame and smoke target detection, solves the problem of imbalance between large, medium and small targets, realizes flexible deployment on different hardware devices, and realizes real-time early warning of early fires.
Smart Images

Figure CN116343077B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of image processing and deep learning, and particularly relates to a fire detection and early warning method based on an attention mechanism and multi-scale features. Background Art
[0002] As one of the disasters recognized worldwide, fire seriously endangers human life and property safety. For the security construction of smart cities, early and effective fire detection and early warning are crucial. Physical signal-based sensors, such as smoke sensors, heat release infrared flame sensors, ultraviolet flame sensors, etc., are widely used in fire alarm systems. Since these traditional physical sensors are limited to the near-fire source location and cannot work effectively in large semi-enclosed space buildings and open underground spaces, and cannot provide detailed disaster information such as the fire location, fire size, and combustion degree, the fire detection technology based on visual sensors can meet these needs.
[0003] There is a method (a method and device for monitoring fire situation based on video, application number 2021112915514) that can obtain the streaming media data of a video camera, preprocess the streaming media data to obtain a target image; use the YOLO-V4 algorithm to detect the target image to determine the bounding boxes of the target image, where the bounding boxes include: a fire situation bounding box and a smoke bounding box; perform superpixel segmentation on the image within the bounding box to obtain superpixel slices, and classify the superpixel slices to obtain an initial fire situation monitoring result; construct an external rectangle box based on the initial fire situation monitoring result and superimpose the external rectangle box on the streaming media data to obtain a target fire situation monitoring result. However, it has the following disadvantages: it is applicable to large and medium-sized fire target data samples and cannot detect small target flames and smoke; there are problems of high missed detection rate and false detection rate in fire detection in multiple scenarios, resulting in low average detection accuracy; the depth and width of the deep neural network model cannot be flexibly adjusted and cannot be better deployed to different hardware devices. Summary of the Invention
[0004] In order to overcome the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a fire detection and early warning method based on an attention mechanism and multi-scale features, which solves the problem of unbalanced large, medium, and small flame and smoke targets, uses a variable-depth and width dynamic neural network to adjust the size of the network model for deployment to different hardware devices, proposes an improved deep learning network model Fire-YOLOv5, introduces a co-location attention mechanism in the backbone network to enhance the weight representation of the target position, realizes better fusion of features at each scale; and realizes real-time early warning of early fires through a video frame voting mechanism.
[0005] In order to achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0006] A fire detection and early warning method based on attention mechanism and multi-scale features, comprising the following steps:
[0007] Step S1, establish a multi-scenario fire data set, preprocess the data, and obtain a training sample set {train1,…,train d ,…,train m} and a test sample set {test1,…,test e ,…,test n};
[0008] Step S2, build an improved deep learning network model Fire-YOLOv5;
[0009] Step S201, set the depth and width coefficients of the neural network to adjust the size of the network model to adapt to different hardware platforms, set the parameter vector of data augmentation, and perform affine transformation, perspective transformation and combined transformation on the image samples to enrich the data set;
[0010] Step S202, use the co-located attention module CAB optimized by the Mish activation function to replace the CSP2_X module in the YOLOv5 backbone network to enhance the weight parameter representation of the region of interest;
[0011] Step S203, use Concat to connect the bidirectional cross-scale link to fuse feature maps of different scales and achieve multi-layer fusion of semantics;
[0012] Step S204, add a set of small target anchor boxes and detection heads to achieve the detection of pixel-level targets with 32 times downsampling of the original image;
[0013] Step S3, continuously iterate and train to minimize the loss function, obtain the trained Fire-YOLOv5 model, and deploy it to the edge server for tunnel monitoring;
[0014] Step S4, the tunnel monitoring acquisition module obtains streaming media data, preprocesses the collected video by normalization, and obtains the image frame sequence of the video;
[0015] Step S5, use the trained Fire-YOLOv5 model to perform fire and smoke detection frame by frame on the image frame sequence of the video;
[0016] Step S6, judge the detection result through the video frame voting mechanism and give an early warning of the occurrence of a fire.
[0017] The beneficial effects of the present invention are:
[0018] Since a multi-scenario flame and smoke image data set is constructed and a variety of data augmentation methods are used, the problem of imbalance of large, medium and small flame and smoke targets is solved;
[0019] Due to the use of variable depth and width dynamic neural networks to adjust the size of the network model to deploy to different hardware devices;
[0020] In order to improve the average accuracy of detection, a deep learning network model Fire-YOLOv5 is proposed. The same-position attention mechanism is introduced in the backbone network to enhance the weight representation of the target position. Based on the principle of bidirectional feature pyramid network, some path aggregation networks are converted into bidirectional cross-scale connections. Through simple splicing operations, better fusion of features at various scales can be achieved. At the same time, a small target detection layer is designed to focus on detecting small targets in visual tasks, and real-time early warning of fires is achieved through a video frame voting mechanism. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 The present invention is a flowchart for implementing the method of the present invention.
[0022] Figure 2 This is a diagram of the Fire-YOLOv5 network structure of an embodiment of the present invention.
[0023] Figure 3 This is a CAB network structure diagram of the attention mechanism module of an embodiment of the present invention.
[0024] Figure 4 This is a network performance diagram of an embodiment of the present invention.
[0025] Figure 5 This is a performance comparison chart of the embodiment of the present invention and other methods. DETAILED DESCRIPTION
[0026] The present invention is described in detail below with reference to the embodiments.
[0027] The network training uses the open source Pytorch deep learning framework, Ubuntu 20.04 system environment, Cuda10.0 and Python3.7 programming environment, the hardware platform GPU model is NVIDIA GeForce RTX 2070Max-Q, the video memory size is 8G, the CPU model is Intel(R) Core(TM) i-10750HCPU@2.60GHz, the memory size is 12G, and training and testing are completed. Due to the limitation of hardware equipment, the training batch size is set to 2, the SGD learning optimizer is used, and the global initial learning rate is set to 0.001.
[0028] Example 1, reference Figure 1 , a fire detection and early warning method based on attention mechanism and multi-scale features, comprising the following steps:
[0029] Step S1: Establish a multi-scenario fire dataset, preprocess the data, and obtain a training sample set {train1, …, train d , …, train m} and a test sample set {test1, …, test e , …, test n};
[0030] Step S2: Build an improved deep learning network model Fire-YOLOv5, as Figure 2 shown;
[0031] Step S201: Set the depth and width coefficients of the neural network to adjust the size of the network model to adapt to different hardware platforms, set the parameter vector of data augmentation, and perform affine transformation, perspective transformation, and combined transformation on the image samples to enrich the dataset;
[0032] Step S202: Use the co-located attention module CAB optimized by the Mish activation function to replace the CSP2_X module in the YOLOv5 backbone network to enhance the weight parameter representation of the region of interest;
[0033] Step S203: Use Concat to connect the bidirectional cross-scale link to fuse feature maps of different scales and achieve multi-layer fusion of semantics;
[0034] Step S204: Add a set of small target anchor boxes and detection heads to achieve the detection of target at the pixel level of 32 times downsampling of the original image;
[0035] Step S3: Continuously iterate and train to minimize the loss function to obtain a trained Fire-YOLOv5 model, and deploy it to the edge server for tunnel monitoring;
[0036] Step S4: The tunnel monitoring acquisition module obtains streaming media data, preprocesses the collected video by normalization, and obtains the image frame sequence of the video;
[0037] Step S5: Use the trained Fire-YOLOv5 model to perform fire and smoke detection frame by frame on the image frame sequence of the video;
[0038] Step S6: The detection results are judged through the video frame voting mechanism and a fire warning is issued.
[0039] Specifically, Step S1 is as follows:
[0040] S101. Obtain multi-scenario fire images Image = {Image1, …, Image i , …, Image N}, produce sample labels in a unified format Label = {Label1, … Label i …, Label N}, each label Label i represents the center point position coordinates (X i of the j-th target in the corresponding sample Image ij , Y ij ), the width and height of the target (W ij , H ij ), and the category {0, 1},, Image i represents the i-th sample in the dataset, i ∈ [0, N], N represents the total number of images, and the categories {0, 1} represent {flame, smoke} respectively;
[0041] S102. Normalize each sample in the dataset to 640 pixels * 640 pixels, and fill the background with gray;
[0042] S103. Divide the normalized dataset into a training set part Train and a test set part Test. For each type of image, 80% is selected as the training set, and the remaining 20% is used as the test set.
[0043] In step S102, the scaling ratio of images with different aspect ratios The image is scaled to where max and min are the maximum and minimum values between the two, w represents the image width, h represents the image height, is rounded up, and the gray filling value is (114, 114, 114).
[0044] Step S2 is specifically as follows,
[0045] In step S201, the network depth of the deep neural network, that is, the number of network layers, and the network width, that is, the network output channels, are controlled by the depth factor DM and the width factor WM respectively. The number of network layers is max(round(number * DM), 1), where number is the number of network layers of different modules, and round is rounding to the nearest integer. The network output channels are where channel is the number of channels of different modules, is rounded up.
[0046] In step S202, referring to Figure 3 , Fire-YOLOv5 introduces an efficient channel attention mechanism module CAB. After the feature pyramid pooling layer, two spatial range pooling kernels are used to perform one-dimensional feature encoding on each channel along the horizontal and vertical coordinates respectively; the two one-dimensional feature encoding outputs of the c-th channel are
[0047] Where W and H are the width and height of the cth channel, and a 1×1 convolution kernel is used to convert the number of channels and the Mish activation function to obtain global spatial information in the horizontal and vertical directions. The output of the intermediate feature map f = δ(F1([z h ,z w ])), [z h ,z w ] represents the two-way tensor concatenation operation along the horizontal and vertical directions, which divides the intermediate feature map into two independent tensors along the spatial dimension and uses two 1×1 convolutions to convert the channels to be consistent with the input channels; the conversion process
[0048] Among them, F h and F w represents two 1×1 convolution transformations, σ represents the Mish activation function; the two tensors g h and g w As the weight parameter of attention. Output of the same-position attention module
[0049] The Mish activation function used is y=x*tanh(ln(1+e x )), this function is a smooth curve, which is not completely truncated in the negative part, allowing relatively small negative gradients to flow into the neural network and more favorable information to get higher accuracy and generalization. With the increase of layer depth, the ReLU activation function will cause the training accuracy to drop rapidly, while the Mish activation function has comprehensive improvements in training stability, average accuracy, peak accuracy, etc.
[0050] In step S203, the Fire-YOLOv5 model combines the principle of the bidirectional feature pyramid network to connect the input nodes and output nodes of the same layer across layers, shortening the path for transmitting low-level semantics to high-level layers, and merging adjacent layers by splicing rather than adding to organically combine the rich semantic features of the high-level with the features at the low-level, significantly improving the accuracy of the prediction; a bidirectional cross-scale connection method that eliminates weights is used to perform feature fusion, aiming to improve detection accuracy without affecting the network's reasoning speed.
[0051] In step S204, in the Fire-YOLOv5 model, the downsampling multiple is too large many times, resulting in the loss of small target information. Considering the limited available resolution and context information of the model, a set of anchor boxes and small target detection layers are added to solve the problem that small fire targets cannot be detected. The feature map output by the 18th layer CBS structure is upsampled to obtain a feature map with a size of 160X160, which is concatenated with the feature map output by the second layer in the backbone network. Then, a CSP_2X layer and a convolutional layer are connected; the input image size is uniformly adjusted to 640X640 pixels. The 160X160 feature map is used to detect targets larger than 4X4 pixels, the 80X80 feature map is used to detect targets larger than 8X8 pixels, the 40X40 feature map is used to detect targets larger than 16X16 pixels, and the 20X20 feature map is used to detect targets larger than 32X32 pixels. After adding the small target detection layer in this way, the four detection structures can cover different receptive fields, realizing the fast detection and accurate positioning of ultra-small pixel targets.
[0052] Step S3 is specifically as follows:
[0053] S301. Set the maximum number of iterations Itera, the learning rate η, the training batch size B. Each time, input B pictures from the training data set {train1,…,train d ,…,train m}. The input times Num is where m is the total number of samples in the training data set; the loss function is the sum of the classification loss, the localization loss, and the positive and negative sample confidence loss, L = L class +L CIoU +L obj +L noobj ;
[0054] S302. Use the gradient descent method to minimize the loss function and perform iterative optimization on the network. Use the SGD learning optimizer, and the global initial learning rate is η, where ω t+1 is used as the network parameter and for prediction, ω t is the current network weight parameter, is the gradient value for the next iteration;
[0055] S303. When the number of iterations has not reached the set minimum number of iterations Itera, if the loss function L no longer decreases, stop training; when the number of iterations reaches the set minimum number of iterations Itera, stop training to obtain the trained network model; otherwise, continue with iterative optimization.
[0056] In step S301, the loss function is specifically as follows:
[0057] T is the number of output feature maps t, S2 is the number of cells in the grid for partitioning the feature map, N is the number of anchor boxes on each grid n, w is the width of the predicted box, h is the height of the predicted box, 1 r<4 The condition for being judged as a positive sample is set such that the ratio of the width and height of the calibration box to the width and height of the predicted box is less than 4;
[0058] Calculate the error between the category inferred by the classification loss and the corresponding calibrated classification:
[0059] where x i is one of the N calibrated categories, taking values {0, 1, …, N - 1}, y i is the normalized category probability, is the probability of the target category inferred by the network;
[0060] Calculate the error between the predicted box and the calibration box for the localization loss:
[0061]
[0062] where w gt is the width of the calibration box, h gt is the height of the calibration box, IoU is the ratio of the intersection over union of the calibration box and the predicted box, ρ 2 (b, b gt ) is the distance between the center points of the calibration box and the predicted box;
[0063] Calculate the confidence of the network for the positive and negative sample confidence losses:
[0064]
[0065] where C is the calibrated confidence, taking values {0, 1}, 0 represents not being the target, 1 represents being the target, gr is the set probability factor, is the inferred confidence, and the confidence of the negative sample is zero;
[0066] In step S4, the tunnel monitoring acquisition module obtains the streaming media data and saves the input video stream output as a picture sequence at interval frames.
[0067] In step S5, use the trained Fire - YOLOv5 model to perform fire and smoke detection frame by frame, draw the target area in the picture sequence and label its category and probability, and finally form a video by grouping the frames.
[0068] In step S6, use the deep neural network to detect N consecutive frames in the video respectively, compare the obtained fire category probabilities with the threshold to infer N prediction voting values, and use these N voting values for decision - making to achieve early - stage fire warning.
[0069] The experimental results are referred toFigure 4 and Figure 5 In the flame and smoke detection tasks, Fire-YOLOv5x achieves a good balance between performance and efficiency and is more robust. The parameters of this network are 70.7M, which is 18.0% less than those of the YOLOv5x network. The detection accuracy reaches 93.5%, which is 2.0% higher than that of YOLOv5x. When the IoU threshold is set to 0.5, the average detection accuracy reaches 71.8%, which is 0.2% higher than before. The inference speed is comparable to that of YOLOv5x. From the F1 value, precision, and recall rate curves of Fire-YOLOv5x, it can be seen that the average class precision and recall rate of detection reach 93.5% and 96% respectively. It can be seen that the new method in this paper has higher detection accuracy and lower missed detection rate. Using a public dataset for testing, the detection accuracy of Fire-YOLOv5x is 1.6% and 2% higher than that of EfficientDet-D4 and YOLOv5 respectively. The detection recall rate is 1.7% higher than that of EfficientDet-D4, and the average detection accuracy when the IoU threshold is 0.5 is 14.5% higher than that of EfficientDet-D4. The detection speed is comparable to that of EfficientDet-D4. Especially when dealing with ultra-small pixels and dense fire targets, the performance is better than existing deep learning-based flame and smoke detection methods. The detection results of tunnel fire videos show that it can achieve fast detection and timely warning of fires. The depth and width of the deep neural network model can be flexibly adjusted, and networks of different scales can be trained and deployed to hardware devices with different computing powers.
Claims
1. A fire detection and early warning method based on attention mechanism and multi-scale features, characterized in that, It includes the following steps: Step S1, establish a multi-scenario fire dataset, preprocess the data, and obtain a training sample set {train1,…,train d ,…,train m} and a test sample set {test1,…,test e ,…,test n}; Step S2, build an improved deep learning network model Fire-YOLOv5; Step S201, set the depth and width coefficients of the neural network to adjust the size of the network model to adapt to different hardware platforms, set the parameter vector of data augmentation, and perform affine transformation, perspective transformation and combined transformation on the image samples to enrich the data set; Step S202, use the co-location attention mechanism module CAB optimized by the Mish activation function to replace the CSP2_X module in the YOLOv5 backbone network to enhance the weight parameter representation of the region of interest; Step S203, use Concat to connect the bidirectional cross-scale link to fuse feature maps of different scales and achieve multi-layer fusion of semantics; The Fire-YOLOv5 model combines the principle of the bidirectional feature pyramid network, cross-connects the input nodes and output nodes of the same layer, shortens the path for low-level semantics to be transmitted to the high level, and organically combines the rich semantic features at the high level with the features at the low level by splicing instead of adding adjacent layers, improving the prediction accuracy; adopts a bidirectional cross-scale connection method with eliminated weights for feature fusion, improving the detection accuracy without affecting the inference operation speed of the network; Step S204, add a group of small target anchor boxes and detection heads to realize the detection of object at the pixel level with 32 times downsampling of the original image; Step S3, continuously iterate and train to minimize the loss function to obtain a trained Fire-YOLOv5 model and deploy it to the edge server for tunnel monitoring; Step S4, the tunnel monitoring acquisition module obtains streaming media data, preprocesses the collected video by normalization to obtain the image frame sequence of the video; Step S5, use the trained Fire-YOLOv5 model to perform fire and smoke detection frame by frame on the image frame sequence of the video; Step S6, judge the detection result through the video frame voting mechanism and give an early warning of the occurrence of fire.
2. The method according to claim 1, wherein Step S1 is specifically: S101. Obtain multi-scenario fire images containing two types of targets, namely flames and smoke, from the open-source dataset Image = {Image1, …, Image i …, Image N}, and create sample labels Label = {Label1, …, Label i …, Label N} in a unified format. Each label Label i represents the center point position coordinates (X i , Y ij ), width and height (W ij , H ij ) of the j-th target in the corresponding sample Image ij ), and the category {0, 1}. Image i represents the i-th sample in the dataset, i ∈ [0, N], where N represents the total number of images. The category {0, 1} represents {flame, smoke} respectively; S102. Normalize each sample in the data set to 640 pixels * 640 pixels and fill the background with gray; S103. Divide the normalized data set into a training set part Train and a test set part Test. For each type of image, select 80% as the training set and the remaining 20% as the test set.
3. The method according to claim 2, characterized in that, In step S102, the scaling ratios of images with different aspect ratios The image is scaled to where max and min are the maximum and minimum values between the two, w represents the image width, and h represents the image height, is rounded up, and the gray filling value is (114, 114, 114).
4. The method according to claim 1, characterized in that, In step 201, the network depth (i.e., the number of network layers) and the network width (i.e., the network output channels) of the deep neural network are controlled by the depth factor DM and the width factor WM respectively. The number of network layers is max(round(number*DM), 1), where number is the number of network layers of different modules, and round is rounding to the nearest integer. The network output channels are where channel is the number of channels of different modules, which is rounding up.
5. The method according to claim 1, wherein In step S202, Fire-YOLOv5 introduces an efficient co-location attention mechanism module CAB. After the feature pyramid pooling layer, two pooling kernels with different spatial ranges are used to perform one-dimensional feature encoding on each channel along the horizontal and vertical coordinates respectively; The two one-dimensional feature encoding outputs of the c-th channel are where W and H are the width and height of the c-th channel, and a 1×1 convolutional kernel is used to convert the number of channels and the Mish activation function is used to obtain the global spatial information in the horizontal and vertical directions. The output of the intermediate feature map is f = δ(F1([z h ,z w )), [z h ,z w represents the concatenation operation of two directional tensors in the horizontal and vertical directions. The intermediate feature map is divided into two independent tensors along the spatial dimension, and two 1×1 convolutions are used to convert the channels to be consistent with the input channels; the conversion process where F h and F w represent two 1×1 convolution transforms, and σ represents the Mish activation function; the two resulting tensors g h and g w serve as the weight parameters of the attention; the output of the co-location attention mechanism module The Mish activation function used is y = x * tanh(ln(1 + e x )), which is a smooth curve and is not completely truncated in the negative value part, allowing relatively small negative gradients to flow in and more favorable information to penetrate deep into the neural network, thus achieving higher accuracy and generalization ability.
6. The method according to claim 1, wherein In step S204, the excessive downsampling multiples in the Fire-YOLOv5 model lead to the loss of small target information. A set of anchor boxes and small target detection layers are added to solve the problem that small fire targets cannot be detected. The feature map output by the 18th layer CBS structure is upsampled to obtain a feature map with a size of 160X160, which is concatenated with the feature map output by the 2nd layer in the backbone network. Then, a CSP_2X layer and a convolutional layer are connected. The input image size is uniformly adjusted to 640X640 pixels. The 160X160 feature map is used to detect targets larger than 4X4 pixels, the 80X80 feature map is used to detect targets larger than 8X8 pixels, the 40X40 feature map is used to detect targets larger than 16X16 pixels, and the 20X20 feature map is used to detect targets larger than 32X32 pixels.
7. The method according to claim 1, characterized in that, Step S3 specifically is, S301. Set the maximum number of iterations Itera, the learning rate η, the training batch size B. Each time, input B images from the training dataset {train1, …, train d , …, train m}, and the number of input times Num is , where m is the total number of samples in the training dataset; the loss function is the sum of the classification loss, the localization loss, and the positive and negative sample confidence losses, L = L class + L CIoU + L obj + L noobj ; S302. Use the gradient descent method Iteratively optimize the network by minimizing the loss function, and adopt the SGD learning optimizer. The global initial learning rate is η, where ω t+1 is used as the network parameter for prediction, and ω t is the current network weight parameter, and is the gradient value for the next iteration; S303. When the number of iterations has not reached the set minimum number of iterations Itera, if the loss function L no longer decreases, then stop training; when the number of iterations reaches the set minimum number of iterations Itera, then stop training to obtain the trained network model; otherwise, continue iterative optimization.
8. The method according to claim 7, wherein In step S301, the loss function is specifically as follows: T is the number of output feature maps t, S 2 is the number of grid cells for feature map division, N is the number of anchor boxes on each grid n, w is the width of the predicted box, h is the height of the predicted box, 1 r<4 The condition for being judged as a positive sample is set such that the ratio r of the width and height of the calibration box to the width and height of the predicted box is less than 4; The classification loss calculates the error between the inferred category of the classification loss and the corresponding calibrated classification. where x i is one of the N calibrated categories, taking values {0, 1, …, N - 1}, and y i is the normalized category probability, and is the probability that the network infers the target category; The localization loss calculates the error between the predicted bounding box and the calibrated bounding box. where w gt is the width of the calibration box, h gt is the height of the calibration box, IoU is the ratio of the intersection over union of the calibration box and the prediction box, ρ 2 (b, b gt ) is the distance between the center points of the calibration box and the prediction box; The positive and negative sample confidence loss calculates the confidence of the network. Where C is the calibrated confidence, taking values in {0, 1}, 0 means not the target, 1 means it is the target, and gr is the set probability factor. The confidence of the inference, and the confidence of the negative sample is zero.
9. The method according to claim 1, wherein In step S6, the deep neural network is used to detect N consecutive frames in the video respectively. The obtained fire category probabilities are compared with the threshold to infer N predicted voting values, and these N voting values are used for decision-making to realize early-stage fire warning.
Citation Information
Patent Citations
Lightweight fusion global and local feature network-based finger vein recognition method
CN114299559A
Driver concentration detection method based on YOLO neural network
CN114387586A