Fire detection method, device and system and nonvolatile storage medium
A fire detection method that utilizes cloud and edge devices in collaboration, employing pre-trained models and data augmentation techniques, solves the problem of high fire detection costs and achieves highly accurate fire detection.
Patent Information
- Application Number
- CN202511285555.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-12-16
AI Technical Summary
In existing technologies, fire detection requires multiple detectors to work together for confirmation, which leads to high costs and the problem that edge devices are not equipped with detectors.
By employing a collaborative approach between cloud and edge devices, video stream data is processed using a first model deployed on the edge device and a second model deployed on the cloud device. The target visual model is trained using pre-trained models and data augmentation techniques to achieve accurate identification of flame and environmental information.
It achieves high-accuracy fire detection without increasing edge device hardware, thus reducing the cost of fire detection.
Smart Images

Figure CN121148084A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and more specifically, to a fire detection method, apparatus, system, and non-volatile storage medium. Background Technology
[0002] In related technologies, fire detection typically requires multiple different types of detectors to work together to confirm the fire. The problem with this approach is that many edge devices with video capture capabilities are not yet equipped with these detectors, and installing detectors on these edge devices would be costly.
[0003] There is currently no effective solution to the above problems. Summary of the Invention
[0004] This application provides a fire detection method, apparatus, system, and non-volatile storage medium to at least solve the technical problem of high fire detection costs caused by the need for multiple detectors to detect fires in related technologies.
[0005] According to one aspect of the embodiments of this application, a fire detection method is provided, comprising: a cloud device acquiring a first processing result obtained after processing video stream data by a first model, wherein the first processing result includes candidate image regions determined by the first model in the video stream data, the candidate image regions being regions where the first model determines that flames exist, and the first model being deployed in an edge device; analyzing the candidate image regions through a second processing model to obtain a second processing result, wherein the second processing model is deployed in a cloud device, and the second processing result includes the presence of flames in the candidate image regions and the absence of flames in the candidate image regions; and determining the candidate image regions where flames exist as fire image regions.
[0006] Optionally, processing the first processing result through a second processing model to obtain a second processing result includes: obtaining a preset prompt word template, wherein the prompt word template includes a first type of prompt word and a second type of prompt word, the first type of prompt word being used to guide the second processing model to process flame information in the candidate image area, and the second type of prompt word being used to guide the second processing model to process environmental information in the candidate image area; using the prompt word template to drive the second processing model to analyze the first processing result to obtain a second processing result, wherein the second processing model is used to extract flame information and environmental information from the image area according to the prompt word template, and to determine the second processing result based on the flame information and environmental information.
[0007] Optionally, the flame information includes at least one of the following: flame shape, flame color, flame brightness, smoke concentration, and smoke color; the environmental information includes at least one of the following: flammable material information, wind direction information, and wind speed information.
[0008] Optionally, before the cloud device obtains the first processing result after processing the video stream data by the first model, the fire detection method further includes: obtaining a first training dataset, wherein the first training dataset includes historical fire image data and flame and smoke data; performing image enhancement processing on the first training dataset to obtain a second training dataset; using the second training dataset to train the pre-trained initial visual model to obtain a target visual model; and training the target visual model to obtain a first processing model.
[0009] Optionally, training the target visual model to obtain the first processing model includes: processing the second training dataset with the target visual model to obtain data labels for the training data in the second training dataset; training the initial image processing model with the second training dataset containing data labels to obtain the target image processing model; and training the target image processing model with a third training dataset to obtain the first processing model, wherein the third training dataset is a manually labeled dataset.
[0010] Optionally, performing image enhancement processing on the first training dataset to obtain the second training dataset includes: performing histogram equalization processing on the image data in the first training dataset; and selecting target data points in the image data according to a preset probability, and performing at least one of the following processing on the target data points to obtain the second training dataset: random rotation processing, horizontal flip processing, or vertical flip processing.
[0011] Optionally, the fire detection method further includes: determining the shooting time and the identification of the shooting device corresponding to the fire image area; determining the location information of the shooting device based on the identification of the shooting device; and generating a fire alarm based on the shooting time and the location information of the shooting device.
[0012] According to another aspect of the embodiments of this application, a fire detection device is also provided, suitable for cloud devices, comprising: a first processing module, configured to acquire a first processing result obtained after a first model processes video stream data, wherein the first processing result includes candidate image regions determined by the first model in the video stream data, the candidate image regions being regions where the first model determines the presence of flames, and the first model being deployed in an edge device; a second processing module, configured to analyze the candidate image regions through a second processing model to obtain a second processing result, wherein the second processing model is deployed in a cloud device, and the second processing result includes the presence of flames in the candidate image regions and the absence of flames in the candidate image regions; and a third processing module, configured to determine the candidate image regions where flames are present as fire image regions.
[0013] According to another aspect of the embodiments of this application, a fire detection system is also provided, including a cloud device and an edge device. The edge device is used to collect video stream data of a target area and process the video stream data using a first model to obtain a first processing result. The first processing result includes candidate image regions determined by the first model in the video stream data. The candidate image regions are areas where the first model determines the presence of flames. The first model is deployed in the edge device. The cloud device is used to analyze the candidate image regions using a second processing model. From the first processing result, a second processing result is obtained. The second processing model is deployed in the cloud device. The second processing result includes a judgment result for the candidate image regions, which includes whether flames exist in the candidate image regions or not. Candidate image regions with a judgment result indicating the presence of flames are determined as fire image regions.
[0014] According to another aspect of the embodiments of this application, a non-volatile storage medium is also provided, in which a program is stored, wherein a fire detection method is executed when the program is run.
[0015] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the program executes a fire detection method when it runs.
[0016] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that implements a fire detection method when executed by a processor.
[0017] In this embodiment, a cloud device is used to acquire a first processing result obtained after processing video stream data by a first model. The first processing result includes candidate image regions determined by the first model in the video stream data. The candidate image regions are areas where the first model determines that flames exist. The first model is deployed in an edge device. A second processing model is used to analyze the candidate image regions to obtain a second processing result. The second processing model is deployed in a cloud device. The second processing result includes the presence of flames in the candidate image regions and the absence of flames in the candidate image regions. By determining the candidate image regions with flames as fire image regions, and by using a cloud-edge collaborative approach to process the video stream data, the fire detection area is determined. This achieves the goal of high-accuracy fire detection even with only video stream data. This achieves the technical effect of accurately determining whether there is a fire without adding multiple detectors to the edge device, thereby solving the technical problem of high fire detection costs caused by the need for multiple detectors in related technologies. Attached Figure Description
[0018] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0019] Figure 1 This is a schematic diagram of a fire detection system provided according to an embodiment of this application;
[0020] Figure 2 This is a schematic flowchart of a fire detection method provided according to an embodiment of this application;
[0021] Figure 3 This is a schematic flowchart of a fire detection process provided according to an embodiment of this application;
[0022] Figure 4 This is a schematic diagram of a fire detection device provided according to an embodiment of this application;
[0023] Figure 5 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of this application. Detailed Implementation
[0024] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0026] To better understand the embodiments of this application, the technical terms involved in the embodiments of this application are explained below:
[0027] GroundingDINO: GroundingDINO is an advanced visual-language pre-trained model. Its main function is visual localization, which involves associating natural language descriptions with specific regions in an image. GroundingDINO is pre-trained on a large amount of image and text data, learning rich visual and linguistic knowledge, providing a powerful technical means for the fusion of vision and language.
[0028] Fine-tuning refers to adapting a pre-trained model to a specific task by adjusting parameters, changing its structure, etc. It utilizes the general features of the pre-trained model, trains it on a new dataset, optimizes model performance, and improves accuracy for a specific task.
[0029] Model refinement: a technique for transferring knowledge from large models to lightweight models. It uses soft labels to convey feature representation capabilities and solves the accuracy loss problem caused by model lightweighting.
[0030] Fire detection technologies in related fields encompass various detectors, such as heat detectors, smoke detectors, and flame detectors. Heat detectors detect fires by monitoring changes in ambient temperature, but their response is slow in the early stages of a fire due to the slow rate of temperature increase. Smoke detectors identify fires by analyzing changes in smoke parameters, but their response is slower in smokeless fires or when smoke has not fully formed. Flame detectors directly detect flames, but their detection effectiveness is affected when flames are obscured or hidden. Furthermore, these detectors require hardware upgrades to existing camera equipment, posing challenges to business expansion and promotion.
[0031] Currently, advancements in computer vision and artificial intelligence technologies have driven the development of image processing-based fire detection algorithms, which have become a research hotspot in the field. These algorithms detect and issue alarms by deeply analyzing flame features in video or image sequences, such as color, shape, texture, and dynamic changes. However, in complex real-world environments, such as those with changing lighting and weather conditions, and multi-target interference, these algorithms are prone to misjudgment, leading to high false alarm rates. Furthermore, due to excessive algorithm complexity or computational resource limitations, detection speeds are slow, and adaptability and robustness to environmental changes are insufficient. The overall system still has room for improvement in terms of real-time performance and accuracy, making it difficult to efficiently handle complex fire scenarios.
[0032] In conclusion, in the current complex and ever-changing practical application environment, how to effectively overcome adverse factors such as changes in light, weather effects, and multi-target interference, significantly improve the accuracy and real-time performance of fire detection, and at the same time greatly reduce the probability of false alarms and missed alarms, has become a core problem that urgently needs to be solved in the field of fire detection technology.
[0033] To address the aforementioned issues, this application provides relevant solutions, which are detailed below.
[0034] According to an embodiment of this application, a method for fire detection is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0035] The method embodiments provided in this application can be implemented as follows: Figure 1 This is executed within the fire detection system shown. From... Figure 1 As can be seen, the system includes a cloud device 10 and an edge device 12. The edge device 12 is used to collect video stream data of the target area and process the video stream data through a first model to obtain a first processing result. The first processing result includes candidate image areas determined by the first model in the video stream data. The candidate image areas are areas where the first model determines that flames exist. The first model is deployed in the edge device 12. The cloud device 10 is used to analyze the candidate image areas through a second processing model. The first processing result yields a second processing result. The second processing model is deployed in the cloud device 10. The second processing result includes the judgment result of the candidate image areas. The judgment result includes whether there are flames in the candidate image areas or not. The candidate image areas with the judgment result of flames are determined to be fire image areas.
[0036] The number of the aforementioned edge devices 12 can be one or more.
[0037] Under the above operating environment, this application provides a fire detection method, such as... Figure 2 As shown, the method includes the following steps:
[0038] Step S202: The cloud device obtains the first processing result after the first model processes the video stream data. The first processing result includes the candidate image region determined by the first model in the video stream data. The candidate image region is the region where the first model determines that there is a flame. The first model is deployed in the edge device.
[0039] In the technical solution provided in step S202, before the cloud device obtains the first processing result after processing the video stream data of the first model, the fire detection method further includes: obtaining a first training dataset, wherein the first training dataset includes historical fire image data and flame and smoke data; performing image enhancement processing on the first training dataset to obtain a second training dataset; using the second training dataset to train the pre-trained initial visual model to obtain a target visual model; and training the target visual model to obtain a first processing model.
[0040] As an optional implementation, the step of training the target visual model to obtain the first processing model includes: processing the second training dataset with the target visual model to obtain the data labels of the training data in the second training dataset; training the initial image processing model with the second training dataset with data labels to obtain the target image processing model; and training the target image processing model with a third training dataset to obtain the first processing model, wherein the third training dataset is a manually labeled dataset.
[0041] In some embodiments of this application, such as Figure 3 As shown, the first processing model (i.e.) is obtained. Figure 3 The process (based on the actual deployment model in the model) includes the following steps:
[0042] The first step is to build the dataset.
[0043] In some embodiments of this application, video images of historical fire incidents and open-source flame and smoke datasets obtained through internet crawling can be collected to obtain a first training dataset. Subsequently, image enhancement processing is performed on the collected video images to obtain higher-quality enhanced images. These enhanced images are then combined with the originally collected images into a constructed dataset to obtain a second training dataset. Finally, this dataset is divided into training, validation, and test sets according to a certain ratio, providing a solid data foundation for subsequent model training. Constructing a high-quality and rich dataset can effectively improve the accuracy of the model's detection results.
[0044] In some embodiments of this application, the step of performing image enhancement processing on the first training dataset to obtain the second training dataset includes: performing histogram equalization processing on the image data in the first training dataset; and selecting target data points in the image data according to a preset probability, and performing at least one of the following processing on the target data points to obtain the second training dataset: random rotation processing, horizontal flip processing, or vertical flip processing.
[0045] The second step is model fine-tuning.
[0046] Based on the training set prepared in the first step, Zero Shot inference is performed on models such as the Grounding DINO model. The inferred dataset, after manual verification, is used as a fine-tuning dataset to fine-tune a general visual model, resulting in the target visual model. This fine-tuning process allows the Grounding DINO model to better adapt to the feature requirements of fire detection tasks, thus obtaining a finely tuned and optimized fire visual model. This enables the entire detection method to more accurately identify flame-related features, effectively improving the ability to extract and judge fire features in complex scenes.
[0047] The third step is model refinement and dataset inference.
[0048] The target visual model is used to infer the dataset constructed in the first step, thereby generating a labeled dataset. This approach further optimizes the model's performance, and the labeled dataset generated through inference provides more targeted data for the training of subsequent lightweight models. This allows the model to learn on more accurately labeled data, improving its accuracy and stability in identifying fires.
[0049] The fourth step is to train the initial image processing model.
[0050] The labeled dataset generated in step three is used to train image processing models such as the YOLOv8 model, thereby obtaining the target image processing model. During this process, the YOLOv8 model and other image processing models, leveraging their inherent object detection capabilities and combined with the processed high-quality dataset, can quickly learn the characteristic patterns of fires, improving the accuracy and speed of the entire detection system and enhancing its ability to discriminate fires in complex scenarios.
[0051] The fifth step is to train the target image processing model to obtain the first processing model.
[0052] By manually annotating the training dataset and using the model obtained in step four as a pre-trained model, and then retraining the pre-trained model with the annotated dataset, a first processing model that can be practically deployed on edge devices can be obtained. Furthermore, by using finely annotated data to further optimize and adjust the model, the final target image processing model can better adapt to various complex situations in real-world application scenarios, improving the model's reliability and practicality in real-world environments, and ensuring that the final first processing model can accurately and stably detect fires.
[0053] In some embodiments of this application, the step of performing image enhancement processing on the first training dataset to obtain the second training dataset includes: performing histogram equalization processing on the image data in the first training dataset; and selecting target data points in the image data according to a preset probability, and performing at least one of the following processing on the target data points to obtain the second training dataset: random rotation processing, horizontal flip processing, or vertical flip processing.
[0054] In some embodiments of this application, a model training process for training a large visual model to obtain a target visual model is also provided. Taking the Grounding Dino model as an example, the process includes the following steps:
[0055] The first step is to build the dataset.
[0056] After a fire occurs or during regular fire-related data collection, data collection personnel or programs can begin collecting video images and datasets. The collected historical fire event video images and open-source flame and smoke datasets obtained through web crawling can be further processed and used as training data.
[0057] As an optional implementation, image enhancement processing can be performed on the collected video images to improve data quality. Specifically, a histogram equalization-based image enhancement method can be used, which significantly improves image contrast by adjusting the histogram distribution of the image. Furthermore, for color images, histogram equalization is performed on the red, green, and blue channels separately to enhance color saturation and detail. Simultaneously, random rotation and flipping operations can also be used to enhance the image. For example, for each pixel in the image, a certain probability can be used to determine whether to rotate or flip that pixel, thus increasing image diversity. For instance, the random rotation angle range can be set between 0° and 360°, and the flipping operations include horizontal and vertical flipping.
[0058] Furthermore, Gaussian blurring can be used to smooth images, effectively reducing image noise. The formula for the Gaussian distribution function is as follows:
[0059]
[0060] Where σ is the standard deviation, and (x,y) represents the coordinates of each pixel in the image coordinate system. The image coordinate system is a Cartesian coordinate system established in the image.
[0061] In some embodiments of this application, the aforementioned Gaussian distribution function can be used to perform a weighted average of the neighborhood of each pixel, resulting in smoother image edges. The enhanced image and the original image are then included in the constructed dataset, and divided into training, validation, and test sets at a ratio of 80%, 10%, and 10%, respectively. This high-quality training, validation, and test sets provide a solid data foundation for subsequent model training, helping to improve the accuracy and reliability of model training.
[0062] The second step is Zeroshot reasoning.
[0063] After preparing the training dataset, you can begin Zero Shot inference and fine-tuning of the Grounding Dino model, a process performed by the cloud server.
[0064] The Grounding DINO model is built upon an image Transformer detector and a language-based localization pre-trained model, possessing powerful feature extraction and understanding capabilities. Through the Transformer architecture, the model can effectively encode and represent features of images and text, achieving alignment and understanding between images and text, thus providing a foundation for object detection. This model also exhibits multi-task adaptability and, by combining with language models (such as BERT), can locate and detect objects in images based on given text descriptions. Furthermore, by adjusting the textual cues for specific object detection, the model can achieve multi-task scene detection, such as closed-set object detection, open-vocabulary object detection, phase grounding, and understanding referential expressions.
[0065] In some embodiments of this application, the model can match the feature vectors of an image with predefined concept embeddings. These predefined concepts include flame-related feature information, such as the color, shape, and texture of the flame. The model generates a confidence score by analyzing the image features, representing the probability of a flame existing in the image. For example, for an image of a fire scene, the model might output a probability value, such as 0.85, indicating a high probability of the flame existing. During inference, the model also outputs the flame's location information, representing its position in the image as a bounding box. The model outputs a rectangle enclosing the flame's location, with coordinates (x1, y1, x2, y2), where (x1, y1) is the coordinate of the top-left corner of the rectangle, and (x2, y2) is the coordinate of the bottom-right corner.
[0066] The third step is manual correction.
[0067] After reasoning is completed, the results can be reviewed and corrected manually. First, the accuracy of the flame location information can be checked, such as whether the flame is correctly framed and whether its location matches reality. The flame category classification will be adjusted based on the actual situation. For example, the flame might be misclassified as another type of object, or its category might be refined into different combustion states. Furthermore, a comprehensive analysis can be conducted by combining the actual conditions of the fire scene, background information from the images, and other relevant factors.
[0068] The fourth step is to standardize the data format.
[0069] The Grounding DINO model can be fine-tuned using inference data that has been manually corrected. First, the data format needs to be standardized using pre-defined formats, such as the Object Detection (OD) format (containing image filename, height, width, and detection instance information, including bounding box coordinates, label, and category) and the phrase grounding (VG) format (containing image filename, height, width, descriptive text, and related region information, including bounding box coordinates, phrases, and the character position of the phrases in the descriptive text). This ensures data standardization and consistency, facilitating model training and processing. Second, data augmentation is necessary. To improve the model's generalization ability, data augmentation operations can be performed before training, such as random rotation, flipping, and scaling of images, increasing data diversity.
[0070] Step 5: Load the model.
[0071] When selecting and loading pre-trained models, the general-purpose visual large model Grounding DINO model, which is pre-trained on a large-scale dataset, can be used as the initialization. The weights of the grounding_dino_swin_b pre-trained model can be selected for loading. By leveraging the advantages of pre-trained models in general feature learning, the convergence speed on specific tasks can be accelerated.
[0072] Step 6: Parameter adjustment and optimization.
[0073] During fine-tuning, the model's weight parameters can be adjusted to better adapt to the characteristic requirements of fire detection tasks. Gradient descent can be used, with AdamW as the optimizer, as shown in the following expression:
[0074] θ t =θ t-1 -γλθ t-1
[0075] In the above expression, t and t-1 represent training rounds, θ represents the set of model weight parameters, γ is the learning rate, and λ is the weight decay coefficient.
[0076] Since the regularization coefficients added to the gradient are ultimately updated on the weights, no additional gradients are needed in between. AdamW directly decays the weights, eliminating the need for gradient additions, and thus outperforms Adam in convergence speed. It calculates the gradient using backpropagation based on the loss function, updates the model parameters to minimize the loss function, and improves model accuracy. Simultaneously, regularization techniques (L1 or L2 regularization) are introduced to prevent overfitting, ensuring that the model's good performance on training data generalizes to real-world applications. Appropriate training strategies need to be developed based on the specific task and dataset characteristics.
[0077] During training, some specific training parameters can be set by the user. For example, you can set the number of training epochs to 20, the learning rate adjustment strategy to MultiStepLR, and the batch size to 1.
[0078] During fine-tuning, the model adjusts its internal feature representations based on the corrected inference data. The model also adjusts its attention mechanism to focus more on the characteristic information of the flame.
[0079] Step 7: Fine-tune monitoring.
[0080] During fine-tuning, model performance monitoring and adjustments are necessary, along with regular model evaluation to track performance changes. Based on the evaluation results, training parameters such as the learning rate and regularization coefficient should be adjusted promptly. Appropriate optimization of the model structure (e.g., adjusting the feature extraction module) can also improve performance. Furthermore, if the model's accuracy on the validation set stops improving or even begins to decline during training, it indicates overfitting. In this case, the learning rate can be reduced or the regularization strength increased.
[0081] Step 8: Testing and Verification.
[0082] During the testing and validation of the fine-tuned model, various evaluation metrics can be used to measure the model's performance during training, including Average Precision (AP) and Average Recall (AR). These metrics can comprehensively reflect the model's detection performance under different IoU thresholds, different target area ranges, and different maximum detection counts, providing an objective basis for judging the model's quality.
[0083] This finely tuned and optimized model, known as the large-scale visual model of the fire situation (or the large-scale visual model of the target), can be applied to real-world scenarios such as fire detection, outputting a finely tuned large-scale visual model of the fire scene. By processing images or video streams of the fire scene, it accurately detects the location, shape, size, and other features of the flames.
[0084] In some embodiments of this application, after obtaining a large visual model of the fire situation, the large visual model can be used to refine image processing models such as the Yolov8 model. Taking the Yolov8 model as an example, the model refinement process includes the following steps:
[0085] The first step is model refinement and inference on real datasets;
[0086] After the model is fine-tuned, the fire vision big model deployed on cloud devices such as servers can be called to segment and infer the video stream data collected by multiple edge devices, and use the inference results as prior knowledge to generate a Yolov8 pre-trained model for this fire scene, which is to achieve model refinement.
[0087] During the refinement process, the server acts as the execution entity, performing refinement operations on the fine-tuned fire situation visual model. Based on its unique architecture, the general visual model, after fine-tuning, possesses a deep understanding of images and text. This embodiment of the application leverages this feature, using the model to perform inference on a real-world dataset. During inference, the fire situation model analyzes each image in the dataset based on its learned features and knowledge. It combines text prompts to accurately identify target objects in the images and generate a labeled dataset. These labels not only contain category information of the target objects but also reflect semantic information related to the text prompts.
[0088] During inference, the fire detection model possesses the ability to handle detection with arbitrary input category names. In this detection mode, the input category names are "fire" and "smoke," and it can detect flames and smoke based on specific inference logic. This feature breaks through the limitations of traditional closed-set detection and expands the detection range. It can also automatically detect bounding boxes corresponding to noun phrases involved in user-input language descriptions. In this process, users can operate in two ways: one is to automatically extract noun phrases using the NLTK library; the other is to manually specify noun phrases in the sentence, such as "probably fire," to enhance the model's learning ability. This inference method provides a flexible and effective means for achieving object detection under complex language descriptions.
[0089] The final result is a training dataset of real-world scenes with pseudo-labels. The pseudo-label generation rule is to retain only predicted bounding boxes with a confidence score > 0.7 (reducing noisy labels), using soft label encoding.
[0090]
[0091] Through manual verification, each target object in the image was labeled using the professional labeling tool Labelme. The labeling format was YOLO, which included the target object's category and location coordinates, serving as the dataset for training the final YOLO model.
[0092] The second step is YOLOv8 pre-model training;
[0093] After completing step S3 (acquiring the training dataset), the pre-training process of the Yolov8 model is triggered, with the server acting as the main execution unit. The Yolov8 model is trained on the dataset using a single Nvidia A10 graphics card on the server. In this stage, an open-source pre-trained model based on Cocos datasets is used, fully utilizing its existing parameters and feature representations to accelerate the convergence speed of the Yolov8 model and improve its performance. Following deep learning training principles, the backpropagation algorithm is used to update the parameters of the Yolov8 model, and the learning rate, batch size, and training epochs can be flexibly adjusted according to specific needs. Simultaneously, the SGD optimization algorithm is employed to effectively ensure the efficiency and stability of the training.
[0094] After completing this step, you can obtain a Yolov8 pre-model that is applicable to multiple scenarios, easy to deploy, and capable of real-time inference.
[0095] The third step is dataset processing;
[0096] To further enhance the model's performance without affecting its inference speed and size, the pre-trained model obtained in step S2 needs to undergo further training.
[0097] In some embodiments of this application, 10,000 image data points from multiple devices can be collected. Approximately 30% are urban street scene images, 20% are indoor environment images, 40% are industrial plant scene images, and 10% are other scene images. It should be noted that the specific figures above are for illustrative purposes only and do not represent a limitation on the solution provided in this application. The following are specific preprocessing methods: Image normalization: The collected images are normalized, adjusting the pixel values to a uniform range of 0-1. This helps improve the stability and efficiency of model training. For an 8-bit image, the pixel value range is originally 0-255; through normalization, it is transformed into a range of 0-1. Image cropping is performed according to actual needs. For an image containing multiple targets, to focus on a specific target, the surrounding redundant parts are cropped. The cropping size is set to 512×512 pixels to ensure a moderate image size for subsequent processing. In addition, the image is scaled to a suitable size. For some high-resolution images, to reduce computation and storage space, they need to be scaled down to a certain ratio. The image is reduced to one-quarter of its original size, that is, scaled from 1920×1280 pixels to 480×320 pixels. Meanwhile, some smaller images may need to be enlarged to meet the requirements of model training.
[0098] The collected image data was divided into training, validation, and test sets according to a certain ratio. 70% of the images (7000 images) were selected as the training set. These images will be used to train the model, allowing iterative training to enable the model to learn the features and patterns of the data. 20% of the remaining images (2000 images) were selected as the validation set. The validation set is used to adjust the model's hyperparameters. Analysis and evaluation of the validation set ensure that the model achieves optimal performance during training. During training, the learning rate, regularization parameters, etc., are adjusted, and the changes in the loss function on the validation set are observed to determine the optimal hyperparameter settings. The remaining 10% of the images (1000 images) are used as the test set. The test set is used to evaluate the model's performance and verify its effectiveness in practical applications. Testing on the test set yields metrics such as accuracy, recall, and average precision, allowing for a comprehensive evaluation of the model's performance.
[0099] The fourth step is model training;
[0100] The model is trained using a training set containing a large amount of image data and corresponding annotations. During training, the backpropagation algorithm is used to optimize the model. The specific steps are as follows:
[0101] Forward propagation: The input image data is passed through the various layers of the model to obtain the prediction results. After the input image passes through convolutional layers, pooling layers, and fully connected layers, the predicted class probabilities are obtained.
[0102] Loss calculation: A loss function is used to measure the difference between the predicted result and the true label. In this embodiment, the cross-entropy loss function is used:
[0103]
[0104] Where N is the number of samples, C is the number of categories, and y ij It is the label of the true category j of sample i. This represents the probability of class j predicted by the model. Backpropagation involves propagating the loss value backward to each layer of the model and calculating the gradient. The model's weights and biases are then updated using gradient descent. Specific training parameter selection includes batching, i.e., dividing the training set into multiple batches, each containing a certain number of samples. For example, a batch size of 60 can be set, training one batch at a time, thus utilizing more sample information in each training iteration and improving training efficiency. The gradient descent method can employ the Stochastic Gradient Descent (SGD) algorithm, randomly selecting a batch for training in each iteration. SGD's advantages include low computational cost and fast convergence. Simultaneously, the learning rate can be adjusted to avoid local optima during gradient descent. Furthermore, cosine annealing can be used to optimize the learning rate, dynamically adjusting it according to the training process to prevent the model from getting trapped in local optima prematurely, thus improving the model's generalization ability and overall performance. The cosine annealing learning rate strategy is based on the properties of the cosine function; as the number of training rounds increases, the learning rate gradually decreases in the form of a cosine function. Specifically, the formula for how the learning rate changes with the number of training epochs is:
[0105]
[0106] Where is lr max The initial maximum learning rate, lr min Here, T is the minimum learning rate, and T is the total number of training epochs. This strategy allows the model to maintain a high learning rate in the early stages of training, enabling rapid convergence. As the number of training epochs increases, the learning rate is gradually reduced to avoid overfitting during training.
[0107] During the model training phase, multi-dimensional data augmentation strategies can be employed to improve the robustness of the model in complex fire scenarios. For example, one or more of the following strategies can be used:
[0108] 1. Mosaic Enhancement: Four training images are randomly stitched together with a probability of 0.8 to simulate a scenario with multiple fire sources coexisting, thereby enhancing the model's ability to detect dense flame targets.
[0109] 2. HSV color perturbation: Apply a ±15° offset to the H (hue) channel, a ±0.5 gain to the S (saturation) channel, and a ±0.3 fluctuation to the V (brightness) channel to enhance the model's adaptability to changes in lighting.
[0110] 3. CutOut Occlusion Enhancement: 10 rectangular mask regions are randomly generated on each image, with the maximum area of a single mask being 15%. This improves the model's anti-interference ability by simulating a scene where flames partially obscure the image.
[0111] 4. MixUp Enhancement: Image mixing coefficients are generated using a Beta distribution (parameter α = 0.5), and the two images are weighted and fused at the pixel level, which effectively enhances the model's sensitivity to small-scale flame targets.
[0112] Step S204: Analyze the candidate image region using the second processing model to obtain the second processing result. The second processing model is deployed in a cloud device. The second processing result includes whether there is a flame in the candidate image region or not.
[0113] In the technical solution provided in step S204, the step of processing the first processing result through the second processing model to obtain the second processing result includes: obtaining a preset prompt word template, wherein the prompt word template includes a first type of prompt word and a second type of prompt word, the first type of prompt word is used to guide the second processing model to process flame information in the candidate image area, and the second type of prompt word is used to guide the second processing model to process environmental information in the candidate image area; using the prompt word template to drive the second processing model to analyze the first processing result to obtain the second processing result, wherein the second processing model is used to extract flame information and environmental information from the image area according to the prompt word template, and determine the second processing result based on the flame information and environmental information.
[0114] In some embodiments of this application, flame information includes at least one of the following: flame shape, flame color, flame brightness, smoke concentration, and smoke color; environmental information includes at least one of the following: flammable material information, wind direction information, and wind speed information.
[0115] Step S206: Determine the candidate image area containing flames as the fire image area.
[0116] In the technical solution provided in step S206, after determining the fire image area, the fire detection method further includes: determining the shooting time and shooting device identifier corresponding to the fire image area; determining the shooting device location information based on the shooting device identifier; and generating a fire alarm based on the shooting time and shooting device location information.
[0117] In some embodiments of this application, the first processing model may be an image processing model trained on the basis of the YOLOv8 model, and the second processing model may be a multimodal large model deployed on cloud devices such as cloud servers. Figure 3 As shown, the process of determining the fire situation using a cloud-edge collaborative approach may include the following steps:
[0118] The first step is model reasoning.
[0119] First, the YOLOv8 model deployed on edge devices performs inference on real-time acquired image data. This image data may originate from monitoring equipment in various scenarios, such as factory workshops, warehouses, forests, and residential areas, covering different lighting conditions, viewing angles, and various environments where fires may occur. Based on its trained target detection capabilities, the YOLOv8 model analyzes the images, attempting to identify whether there are fire-related features and targets, such as flames, smoke, and burning objects.
[0120] When the YOLOv8 model's inference results indicate a potential fire, to avoid unnecessary interference and resource waste caused by misjudgments, the corresponding image data is transmitted to a multimodal large model deployed on a cloud server. This transmission process is not a simple data transfer but requires consideration of many practical factors. Since edge devices are typically located in environments with relatively unstable network conditions, efficient data transmission protocols and encryption technologies are employed to ensure the integrity and reliability of data transmission. This application's embodiment uses reliable transmission based on the TCP / IP protocol, combined with SSL / TLS encryption, to guarantee the security and integrity of image data during transmission. Simultaneously, to address network bandwidth limitations, the image data is appropriately compressed. While ensuring that critical image information is not lost, the image is converted into a more compact base64 data format to reduce the amount of data transmitted and improve transmission efficiency.
[0121] The second step is multimodal analysis.
[0122] The multimodal large model deployed on a cloud server is a powerful comprehensive analysis system that integrates data processing capabilities across multiple modalities, including images and text, to perform comprehensive and in-depth analysis of input data. Upon receiving images transmitted from edge devices, the multimodal large model can be invoked to make further judgments based on prompt statements. These prompt statements are carefully selected and optimized, guiding the multimodal large model not only to focus on direct fire features in the image, such as the shape, color, and brightness of flames, and the concentration and color changes of smoke, but also to comprehensively consider surrounding environmental information, such as the presence of flammable materials, wind direction, and wind speed, and their impact on the fire situation. An example of a prompt statement is as follows: "Based on the provided image, carefully analyze whether there are any burning flames, and note whether the color range of the flames conforms to common combustion phenomena."
[0123] The input data structure is as follows:
[0124] json
[0125] {"image_data":"base64 encoded JPEG image",
[0126] "detection":{
[0127] "bbox":[x1,y1,x2,y2], / / normalized coordinates
[0128] "confidence": 0.82 / / Confidence of the edge model
[0129] },
[0130] "prompt": "Determines whether this region meets the requirements of the prompt word."
[0131] }
[0132] The third step is to screen and output the results.
[0133] Through pre-defined prompt statements, the multimodal large model can make judgments on images from a more comprehensive perspective. Its internal working mechanism integrates various advanced deep learning technologies, including image feature extraction based on convolutional neural networks, semantic understanding based on natural language processing, and environmental information association analysis based on deep learning. It maps information from images to its existing knowledge graph, which is trained using a large amount of real-world fire scene data and related environmental data. This knowledge graph contains information such as the various characteristics of different types of fires, environmental influencing factors, and possible development trends.
[0134] If the multimodal large model confirms, based on the above comprehensive analysis, that a fire does indeed exist in the image, it will output an alarm message. However, if the multimodal large model determines that the situation in the image is not a real fire, but rather a misjudgment by the YOLOv8 model due to factors such as light interference or misidentification of similar objects, the system will cancel the alarm operation to avoid causing unnecessary panic and resource mobilization.
[0135] Throughout the collaborative process, system performance monitoring and optimization are also involved. Optionally, real-time monitoring of communication latency between edge devices and cloud servers can be performed to ensure timely data transmission and processing; the analysis time of the multimodal large model can be statistically analyzed to continuously improve processing speed in subsequent optimizations; and the parameters of the YOLOv8 model and the multimodal large model can be dynamically adjusted based on misjudgments and correct judgments to further improve the overall system performance.
[0136] By using this collaborative processing method of YOLOv8 models in edge applications and multimodal large models in the cloud, not only is the rapid inference advantage of YOLOv8 models utilized, but also the powerful analytical capabilities of multimodal large models are leveraged to provide a more accurate, reliable, and practically valuable solution for fire detection, providing strong technical support for fire prevention and emergency response in various environments.
[0137] In some embodiments of this application, to further improve the accuracy of identifying fire image regions, a prediction model can be deployed in the cloud. This prediction model can be used to predict possible subsequent fire images based on the fire images input into the model. After obtaining the second processing result, the determined fire image regions can be arranged in chronological order according to timestamps or shooting times. The top-ranked fire image regions are then input into the prediction model to obtain the predicted fire image regions at a preset time. The specific number of image regions input into the model can be set independently. Each image region can be considered an independent image. The predicted fire image regions at the preset time can then be compared with the actual fire image regions. If the similarity between the predicted and actual fire image regions is lower than a preset similarity threshold, the actual fire image regions at the preset time and the fire image regions input into the prediction model are considered abnormal. Manual review is then used to reconfirm whether the actual fire image regions and the input images actually contain flames.
[0138] In some embodiments of this application, real fire videos and flame data can be used to train the prediction model. Furthermore, during training, the training process can be adjusted to ensure that the prediction model's output conforms to the physical laws of the actual combustion process. This ensures that, assuming no anomalies in the input data, the prediction model's output conforms to the physical laws of the actual combustion process. Therefore, if the prediction model's output is found to be inconsistent with the actual collected data, it can be assumed that the input image may contain image regions that do not actually contain flames or other characteristics expected in the combustion process but are mistakenly identified as combustion features. Alternatively, it could be that the actually collected image region is not a fire image region but is mistakenly identified as one. This approach can avoid misjudgments caused by large model illusions, further improving the accuracy of fire assessment results.
[0139] The first processing result is obtained by acquiring video stream data through a cloud device using a first model. This first processing result includes candidate image regions identified by the first model within the video stream data. These candidate image regions are areas where the first model determines flames exist. The first model is deployed on an edge device. A second processing model analyzes these candidate image regions to obtain a second processing result. This second processing model is deployed on a cloud device. The second processing result includes candidate image regions containing flames and candidate image regions where flames are not present. This method of determining candidate image regions containing flames as fire detection areas, through cloud-edge collaboration in processing video stream data, achieves high-accuracy fire detection even with only video stream data. This eliminates the need for additional detectors on edge devices to accurately determine the presence of a fire, thus solving the technical problem of high fire detection costs caused by the need for multiple detectors in related technologies.
[0140] This application provides a fire detection device suitable for cloud-based devices. Figure 4 This is a schematic diagram of the device. From... Figure 4As can be seen from the diagram, the device includes: a first processing module 40, used to acquire a first processing result obtained after the first model processes video stream data, wherein the first processing result includes candidate image regions determined by the first model in the video stream data, the candidate image regions being regions where the first model determines that flames exist, and the first model is deployed in an edge device; a second processing module 42, used to analyze the candidate image regions through a second processing model to obtain a second processing result, wherein the second processing model is deployed in a cloud device, and the second processing result includes the presence of flames in the candidate image regions and the absence of flames in the candidate image regions; and a third processing module 44, used to determine that the candidate image regions containing flames are fire image regions.
[0141] In some embodiments of this application, before the cloud device obtains the first processing result after processing the video stream data of the first model, the fire detection device is further configured to: obtain a first training dataset, wherein the first training dataset includes historical fire image data and flame and smoke data; perform image enhancement processing on the first training dataset to obtain a second training dataset; use the second training dataset to train the pre-trained initial visual model to obtain a target visual model; and train the target visual model to obtain a first processing model.
[0142] In some embodiments of this application, the step of training a target visual model to obtain a first processing model by the fire detection device includes: processing a second training dataset using the target visual model to obtain data labels for the training data in the second training dataset; training an initial image processing model using the second training dataset with data labels to obtain a target image processing model; and training the target image processing model using a third training dataset to obtain the first processing model, wherein the third training dataset is a manually labeled dataset.
[0143] In some embodiments of this application, the fire detection device performs image enhancement processing on the first training dataset to obtain the second training dataset, including: performing histogram equalization processing on the image data in the first training dataset; and selecting target data points in the image data according to a preset probability, and performing at least one of the following processing on the target data points to obtain the second training dataset: random rotation processing, horizontal flipping processing, or vertical flipping processing.
[0144] In some embodiments of this application, the second processing module 42 processes the first processing result through a second processing model to obtain a second processing result, including: obtaining a preset prompt word template, wherein the prompt word template includes a first type of prompt word and a second type of prompt word, the first type of prompt word being used to guide the second processing model to process flame information in the candidate image area, and the second type of prompt word being used to guide the second processing model to process environmental information in the candidate image area; using the prompt word template to drive the second processing model to analyze the first processing result to obtain a second processing result, wherein the second processing model is used to extract flame information and environmental information from the image area according to the prompt word template, and to determine the second processing result based on the flame information and environmental information.
[0145] In some embodiments of this application, flame information includes at least one of the following: flame shape, flame color, flame brightness, smoke concentration, and smoke color; environmental information includes at least one of the following: flammable material information, wind direction information, and wind speed information.
[0146] In some embodiments of this application, after determining the fire image area, the third processing module 44 is further configured to: determine the shooting time and shooting device identifier corresponding to the fire image area; determine the shooting device location information based on the shooting device identifier; and generate a fire alarm based on the shooting time and shooting device location information.
[0147] It should be noted that each module in the above-mentioned fire detection device can be a program module (for example, a set of program instructions to implement a certain function) or a hardware module. For the latter, it can be manifested in the following forms, but is not limited to them: each of the above modules is manifested as a processor, or the functions of each of the above modules are implemented by a processor.
[0148] Figure 5 A hardware block diagram of an electronic device for implementing a fire detection method is shown. Figure 5 As shown, the electronic device 50 may include one or more processors 502 (shown as 502a, 502b, ..., 502n in the figure) 502 (processor 502 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 504 for storing data, and a transmission device 506 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 5 The structure shown is for illustrative purposes only and does not limit the structure of the electronic device described above. For example, electronic device 50 may also include... Figure 5The more or fewer components shown, or having the same Figure 5 The different configurations shown.
[0149] It should be noted that the aforementioned one or more processors 502 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 50 (or mobile device). As described in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0150] The memory 504 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the fire detection method in this embodiment. The processor 502 executes various functional applications and data processing by running the software programs and modules stored in the memory 504, thereby realizing the aforementioned fire detection method. The memory 504 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 504 may further include memory remotely located relative to the processor 502, and these remote memories can be connected to the electronic device 50 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0151] The transmission device 506 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the electronic device 50. In one example, the transmission device 506 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 506 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0152] The display may be, for example, a touchscreen liquid crystal display (LCD), which allows a user to interact with the user interface of the electronic device 50.
[0153] In some embodiments of this application, a non-volatile storage medium is also provided, which stores a program that executes the following fire detection method during program execution: a cloud device acquires a first processing result obtained after a first model processes video stream data, wherein the first processing result includes candidate image regions determined by the first model in the video stream data, the candidate image regions being areas where the first model determines the presence of flames, and the first model is deployed in an edge device; the candidate image regions are analyzed by a second processing model to obtain a second processing result, wherein the second processing model is deployed in a cloud device, and the second processing result includes the presence of flames in the candidate image regions and the absence of flames in the candidate image regions; the candidate image regions where flames are determined are identified as fire image regions.
[0154] In some embodiments of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the following fire detection method: a cloud device acquires a first processing result obtained after processing video stream data by a first model, wherein the first processing result includes candidate image regions determined by the first model in the video stream data, the candidate image regions being areas where the first model determines the presence of flames, and the first model is deployed in an edge device; the candidate image regions are analyzed by a second processing model to obtain a second processing result, wherein the second processing model is deployed in a cloud device, and the second processing result includes the presence of flames in the candidate image regions and the absence of flames in the candidate image regions; the candidate image regions where flames are determined are identified as fire image regions.
[0155] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0156] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.
[0157] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0158] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0159] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to related technologies, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0160] The above are merely preferred embodiments of this application. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A fire detection method, characterized in that, include: The cloud device obtains a first processing result after the first model processes the video stream data. The first processing result includes candidate image regions determined by the first model in the video stream data. The candidate image regions are areas where the first model determines that flames exist. The first model is deployed in an edge device. The candidate image region is analyzed by a second processing model to obtain a second processing result. The second processing model is deployed in the cloud device. The second processing result includes whether there is a flame in the candidate image region or whether there is no flame in the candidate image region. The candidate image area where flames are present is determined to be a fire image area.
2. The fire detection method according to claim 1, characterized in that, The second processing result is obtained by processing the first processing result through the second processing model, including: Obtain a preset prompt word template, wherein the prompt word template includes a first type of prompt word and a second type of prompt word, the first type of prompt word is used to guide the second processing model to process the flame information in the candidate image area, and the second type of prompt word is used to guide the second processing model to process the environmental information in the candidate image area; The second processing model is driven by the prompt word template to analyze the first processing result and obtain the second processing result. The second processing model is used to extract the flame information and the environment information from the image area according to the prompt word template, and to determine the second processing result based on the flame information and the environment information.
3. The fire detection method according to claim 2, characterized in that, The flame information includes at least one of the following: flame shape, flame color, flame brightness, smoke concentration, and smoke color; The environmental information includes at least one of the following: flammable material information, wind direction information, and wind speed information.
4. The fire detection method according to claim 1, characterized in that, Before the cloud device obtains the first processing result after processing the video stream data using the first model, the fire detection method further includes: Obtain a first training dataset, wherein the first training dataset includes historical fire image data and flame and smoke data; The first training dataset is subjected to image enhancement processing to obtain the second training dataset; The pre-trained initial visual model is trained using the second training dataset to obtain the target visual model; The target visual model is trained to obtain the first processing model.
5. The fire detection method according to claim 4, characterized in that, Training the target visual model to obtain the first processing model includes: The target visual model is used to process the second training dataset to obtain the data labels of the training data in the second training dataset. The initial image processing model is trained using the second training dataset with the data labels to obtain the target image processing model; The target image processing model is trained using a third training dataset to obtain the first processing model, wherein the third training dataset is a manually labeled dataset.
6. The fire detection method according to claim 4, characterized in that, The first training dataset is subjected to image augmentation processing to obtain the second training dataset, which includes: Histogram equalization is performed on the image data in the first training dataset; and, Target data points are selected from the image data according to a preset probability, and the target data points are processed by at least one of the following methods to obtain the second training dataset: random rotation processing, horizontal flip processing, or vertical flip processing.
7. The fire detection method according to claim 1, characterized in that, The fire detection method also includes: Determine the shooting time and shooting device identifier corresponding to the fire image area; The location information of the shooting device is determined based on the identification of the shooting device; A fire alarm is generated based on the shooting time and the location information of the shooting device.
8. A fire detection device, suitable for cloud-based equipment, characterized in that, include: The first processing module is used to obtain a first processing result obtained after the first model processes video stream data. The first processing result includes candidate image regions determined by the first model in the video stream data. The candidate image regions are areas where the first model determines that flames exist. The first model is deployed in an edge device. The second processing module is used to analyze the candidate image region through a second processing model to obtain a second processing result, wherein the second processing model is deployed in the cloud device, and the second processing result includes whether there is a flame in the candidate image region or whether there is no flame in the candidate image region; The third processing module is used to determine the candidate image area where flames exist as a fire image area.
9. A fire detection system, characterized in that, Including cloud devices and edge devices, among which, The edge device is used to collect video stream data of a target area and process the video stream data through a first model to obtain a first processing result. The first processing result includes candidate image areas determined by the first model in the video stream data. The candidate image areas are areas where the first model determines that there is a flame. The first model is deployed in the edge device. The cloud device is used to analyze the candidate image regions through a second processing model. The first processing result yields a second processing result. The second processing model is deployed in the cloud device. The second processing result includes a judgment result of the candidate image regions, which includes whether there is a flame in the candidate image regions or whether there is no flame in the candidate image regions. The candidate image regions with the judgment result indicating the presence of flames are determined to be fire image regions.
10. A non-volatile storage medium, characterized in that, The non-volatile storage medium stores a program, wherein when the program is executed, it controls the device containing the non-volatile storage medium to perform the fire detection method according to any one of claims 1 to 7.
11. An electronic device, characterized in that, include: A memory and a processor, the processor being configured to run a program stored in the memory, wherein the program, when running, executes the fire detection method according to any one of claims 1 to 7.
12. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the fire detection method according to any one of claims 1 to 7.