Fire identification method, system, electronic device and computer readable storage medium
A fire detection system that works collaboratively between cloud servers and edge devices utilizes large language models and lightweight models for multi-round recognition, solving the accuracy and real-time issues of traditional fire detection and achieving high-precision and low-cost fire detection.
Patent Information
- Application Number
- CN202511514812.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-22
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2045-10-22
AI Technical Summary
Traditional fire detection methods struggle to detect fires promptly and accurately in complex environments, and the equipment is expensive and has a limited monitoring range.
The fire identification system, which employs a collaborative approach between cloud servers and edge devices, utilizes a large-scale fire identification language model and a lightweight model. It identifies fires through multiple rounds of identification and summary instructions, and combines model training samples with image degradation enhancement technology to improve identification accuracy and interpretability.
It achieves high-precision fire identification in complex environments, reduces equipment costs, and improves the real-time performance and accuracy of identification through cloud-edge collaboration.
Smart Images

Figure CN120997774B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of disaster monitoring, and more particularly to a fire identification method, system, electronic device, and computer-readable storage medium. Background Technology
[0002] Traditional fire detection is usually based on monitoring systems equipped with sensors; however, this method often fails to detect fires in a timely and accurate manner in complex environments, and has limitations such as high equipment costs and limited monitoring range. Summary of the Invention
[0003] The main objective of this invention is to propose a fire identification method, system, electronic device, and computer-readable storage medium, aiming to solve the problems of poor accuracy and timeliness in fire detection in the prior art.
[0004] To achieve the above objectives, the present invention provides a fire detection method applied to a cloud server, the method comprising the following steps:
[0005] Acquire scene images uploaded from the edge terminal and input the scene images into the fire recognition language model, wherein the scene images are uploaded when a fire is detected at the edge terminal;
[0006] Input multi-round recognition instructions to the fire recognition big language model, so that the fire recognition big language model outputs recognition sub-results corresponding to the multi-round recognition instructions based on the scene image, wherein the multi-round recognition instructions are for different objects in the scene image;
[0007] Input a summary command into the fire identification language model so that the fire identification language model integrates multiple identification sub-results to obtain a fire identification result.
[0008] Optionally, the input of multi-round recognition instructions to the fire recognition large language model includes:
[0009] Input scene recognition instructions into the fire recognition big language model, so that the fire recognition big language model outputs scene recognition results based on the scene image;
[0010] An environmental element recognition instruction is input into the fire recognition big language model, so that the fire recognition big language model outputs environmental element recognition results based on the scene image;
[0011] Input fire feature recognition instructions into the fire recognition big language model so that the fire recognition big language model outputs fire feature recognition results based on the scene image.
[0012] Optionally, the method further includes:
[0013] Obtain model training samples, wherein the model training samples include scene image samples, recognition instruction samples, and expected result samples;
[0014] Obtain an initial large language model for fire identification;
[0015] The scene image samples and the recognition instruction samples are input into the initial fire recognition large language model to obtain the training recognition result output by the initial fire recognition large language model;
[0016] The training recognition results are compared with the expected result samples to obtain the result differences;
[0017] The initial fire identification language model is updated based on the differences in the results to obtain the fire identification language model.
[0018] Optionally, updating the initial fire identification large language model based on the result differences includes:
[0019] Determine whether a preset keyword exists in the training recognition result or the expected result sample;
[0020] If the training recognition results or the expected result samples contain preset keywords, the initial fire recognition big language model is updated based on the result differences.
[0021] To achieve the above objectives, the present invention also provides a fire detection method applied at the edge, the fire detection method comprising:
[0022] Acquire scene images and input the scene images into a lightweight fire identification model;
[0023] Obtain the edge results corresponding to the scene image output by the lightweight fire recognition model;
[0024] Determine whether the edge result indicates a fire has occurred;
[0025] If the edge result indicates a fire has occurred, the scene image will be uploaded to the cloud server.
[0026] Optionally, the step of inputting the scene image into the lightweight fire recognition model includes:
[0027] Obtain the scenario type for the application of the lightweight fire identification model;
[0028] Obtain model training samples corresponding to the scene type;
[0029] Obtain an initial lightweight fire identification model;
[0030] The initial lightweight fire identification model is trained using the model training samples to obtain the lightweight fire identification model.
[0031] Optionally, training the initial fire identification lightweight model using the model training samples includes:
[0032] The training samples of the model are subjected to image degradation enhancement to obtain degraded training samples;
[0033] The initial lightweight fire identification model is trained using the degraded training samples.
[0034] To achieve the above objectives, the present invention also provides a fire detection system, which includes a cloud server and an edge terminal; the cloud server is used for:
[0035] Acquire scene images uploaded from the edge terminal and input the scene images into the fire recognition language model, wherein the scene images are uploaded when a fire is detected at the edge terminal;
[0036] Input multi-round recognition instructions to the fire recognition big language model, so that the fire recognition big language model outputs recognition sub-results corresponding to the multi-round recognition instructions based on the scene image, wherein the multi-round recognition instructions are for different objects in the scene image;
[0037] Input a summary command into the fire identification language model so that the fire identification language model integrates multiple identification sub-results to obtain a fire identification result;
[0038] The edge end is used for:
[0039] Acquire scene images and input the scene images into a lightweight fire identification model;
[0040] Obtain the edge results corresponding to the scene image output by the lightweight fire recognition model;
[0041] Determine whether the edge result indicates a fire has occurred;
[0042] If the edge result indicates a fire has occurred, the scene image will be uploaded to the cloud server.
[0043] To achieve the above objectives, the present invention also provides an electronic device, the electronic device including a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the fire identification method as described above.
[0044] To achieve the above objectives, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the fire identification method as described above.
[0045] This invention proposes a fire identification method, system, electronic device, and computer-readable storage medium. The method involves acquiring scene images uploaded from an edge device and inputting these images into a fire identification language model. The scene images are uploaded when a fire is detected at the edge device. Multiple rounds of identification instructions are input to the fire identification language model, causing it to output identification sub-results corresponding to these instructions based on the scene images. These instructions target different objects within the scene images. A summary instruction is then input to the fire identification language model, enabling it to synthesize the multiple identification sub-results to obtain a fire identification result. Using a large language model for fire identification leverages its multimodal capabilities to improve accuracy. The multi-round identification instructions guide the fire identification language model to fully perceive the content of the scene images, enhancing its reasoning depth and discrimination accuracy. Furthermore, by performing fire identification on both the cloud server and the edge device, collaborative detection between the cloud and edge devices is achieved, resulting in a dual improvement in fire identification accuracy and interpretability. Attached Figure Description
[0046] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0047] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0048] Figure 1 This is a flowchart illustrating the first embodiment of the fire identification method of the present invention;
[0049] Figure 2 This is a detailed flowchart of the fire identification method of the present invention;
[0050] Figure 3 This is a schematic diagram of the module structure of the electronic device of the present invention. Detailed Implementation
[0051] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.
[0052] This invention provides a fire detection method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the fire detection method of the present invention, applied to a cloud server. The method includes the following steps:
[0053] Step S11: Obtain the scene image uploaded by the edge terminal and input the scene image into the fire recognition big language model, wherein the scene image is uploaded when the fire is detected at the edge terminal;
[0054] The scene image is an image captured by an image acquisition device located at the edge.
[0055] Edge systems are deployed close to the fire monitoring site. Edge systems can be equipped with image acquisition devices, computing devices, etc. Image acquisition devices include cameras. Computing devices can perform information processing tasks at the edge system. The number of edge systems can be set according to actual needs, such as setting up multiple edge systems to monitor multiple different fire monitoring sites.
[0056] A cloud server is a high-performance computing platform located in the cloud; a communication connection is established between the cloud server and the edge.
[0057] The fire recognition big language model is a model set up on a cloud server to generate the final fire recognition result from scene images; the specific type of the fire recognition big language model can be set according to actual needs, such as LLaVA, BLIP-2, and Qwen-VL.
[0058] In this embodiment, the edge device acquires scene images in real time through a camera and performs fire identification based on the scene images. When a fire is detected, the scene image is uploaded to the cloud server, and the final fire detection is achieved through a fire identification big data language model.
[0059] Understandably, since the edge only performs preliminary fire detection, lightweight models can be deployed at the edge to detect fires, thereby improving detection efficiency and reducing deployment costs.
[0060] Step S12: Input multi-round recognition instructions to the fire recognition big language model, so that the fire recognition big language model outputs recognition sub-results corresponding to the multi-round recognition instructions according to the scene image, wherein the multi-round recognition instructions are for different objects in the scene image;
[0061] The recognition command is a natural language guidance command set for fire recognition, such as prompt words. After the recognition command is input into the fire recognition language model, the fire recognition language model generates the corresponding recognition sub-result based on the recognition command. If the recognition command is "Please judge whether there is a fire in the picture and explain the reason", the recognition sub-result output by the fire recognition language model will include the detection answer of whether there is a fire and the reason for the answer.
[0062] In this embodiment, a multi-round recognition instruction is set, that is, by inputting recognition instructions multiple times, the fire recognition big data language model outputs multiple rounds of corresponding recognition sub-results. Since the multi-round recognition instructions correspond to different objects in the scene image, the fire recognition big data language model can perform semantic analysis on different objects when outputting recognition sub-results for different recognition instructions. Thus, through multi-round analysis, the fire recognition big data language model can more comprehensively identify the scene image and clearly understand the specific state of different types of objects in the scene image, thereby providing a reliable basis for the final output of subsequent fire recognition results.
[0063] Step S13: Input a summary command into the fire identification big language model so that the fire identification big language model integrates multiple identification sub-results to obtain a fire identification result.
[0064] The summary instruction is used to instruct the fire identification big language model to output the final fire identification result. The fire identification big language model has the ability to connect contexts. Therefore, after receiving the summary instruction, it will output the fire identification result by combining the identification sub-results obtained from multiple rounds of identification instructions.
[0065] The fire identification result is the final indication of the fire occurrence obtained from the cloud server; the content of the fire identification result can be set according to actual needs, such as whether a fire has occurred, the location of the fire, the fire intensity, and cause analysis.
[0066] In practice, the scene image can be converted into a visual feature vector; and multiple recognition sub-results can be generated sequentially based on the corresponding recognition instructions; then, the fire recognition result can be obtained by judging according to the preset rules.
[0067] Based on the fire identification results, specific response strategies can be implemented. When the fire identification results indicate that a fire has occurred, a platform-level fire warning is triggered, including pushing alarms to the command center, linking cameras to focus, and activating fire-fighting robots. When the fire identification results indicate that no fire has occurred, it is recorded as a false alarm sample and automatically archived. When the fire identification results are uncertain, the system guides the backend to assist in the review or marks it as a sample to be trained and adds it to the subsequent model incremental learning library.
[0068] The cloud server can periodically return the reviewed samples and their semantic paths to the edge, which can be used to update the model parameters of the lightweight fire identification model or optimize the feature extraction logic, thereby achieving continuous iterative optimization of cloud-edge joint learning.
[0069] Applied to the edge, the fire identification method includes:
[0070] Step S21: Acquire a scene image and input the scene image into the lightweight fire recognition model;
[0071] Step S22: Obtain the edge results corresponding to the scene image output by the lightweight fire recognition model;
[0072] Step S23: Determine whether the edge result indicates a fire has occurred;
[0073] Step S24: If the edge result indicates a fire has occurred, then the scene image is uploaded to the cloud server.
[0074] If the edge result indicates that no fire has occurred, the scene image is not uploaded to the cloud server. In other embodiments, the scene image can be uploaded continuously. When the edge result indicates that a fire has occurred, the fire indicator is associated with the scene image and uploaded to the cloud server; when the edge result indicates that no fire has occurred, the no-fire indicator is associated with the scene image and uploaded to the cloud server.
[0075] Understandably, since edge detection only performs preliminary fire detection, lightweight models can be deployed at the edge to improve detection efficiency and reduce deployment costs. The specific type of lightweight fire detection model can be chosen based on actual needs, such as MobileNet, EfficientNet-lite, or YOLO-Lite. The lightweight fire detection model outputs a fire probability based on the scene image. If the fire probability is greater than a fire threshold, the edge result is that a fire has occurred; if the fire probability is less than or equal to the fire threshold, the edge result is that no fire has occurred.
[0076] The edge result is the recognition result output by the lightweight fire recognition model based on the scene image.
[0077] If the edge result indicates a fire, it means that a fire may exist at this time. Therefore, the scene image is uploaded to the cloud server so that the cloud server can determine the final fire identification result based on the scene image.
[0078] When the edge device indicates a fire, it continuously uploads real-time scene images to the cloud server, enabling the cloud server to maintain fire identification based on continuous scene images. When the cloud server determines that there is no fire, it sends a false alarm command to the cloud server, which then stops uploading scene images. However, it continues to collect scene images and performs fire identification using the lightweight fire identification model scene images.
[0079] This embodiment proposes a cloud-edge collaborative fire detection mechanism based on edge short-thinking and cloud long-thinking. Short-thinking refers to the rapid initial screening and efficient response of the lightweight fire detection model at the edge to large-scale monitored scene images, primarily achieving low-latency detection of suspected fires. Long-thinking refers to the deep semantic analysis and multi-round logical reasoning of the large-scale fire detection language model set up on the cloud server, providing a high-precision and interpretable final judgment. The two complement each other through a collaborative verification mechanism, ensuring both real-time performance and accuracy and robustness in complex environments.
[0080] This embodiment uses a large language model for fire identification, leveraging its multimodal capabilities to improve accuracy. By progressively obtaining identification results through multiple rounds of identification instructions, the fire identification large language model can be guided to fully perceive the content of the scene image, thereby enhancing its reasoning depth and discrimination accuracy. Simultaneously, by performing fire identification on scene images at both the cloud server and the edge, collaborative detection between the cloud and edge can be constructed, achieving a dual improvement in fire identification accuracy and interpretability.
[0081] Furthermore, see also Figure 2 In the second embodiment of the fire identification method of the present invention based on the first embodiment, step S11 includes the following steps:
[0082] Step S111: Input scene recognition instructions into the fire recognition big language model so that the fire recognition big language model outputs scene recognition results based on the scene image;
[0083] Step S112: Input environmental element recognition instructions to the fire recognition big language model, so that the fire recognition big language model outputs environmental element recognition results based on the scene image;
[0084] Step S113: Input fire feature recognition instructions to the fire recognition big language model so that the fire recognition big language model outputs fire feature recognition results based on the scene image.
[0085] Scene recognition instructions are used to prompt the fire recognition big data language model to output a linguistic description of the overall environment and objects in the scene image, in order to extract scene information and significant visual elements from the scene image; this enables the fire recognition big data language model to establish preliminary image semantic cognition, forming the basis for subsequent reasoning; scene recognition instructions can be set according to actual needs, such as "Please describe the main objects and scenes in the picture"; the scene recognition results output by the fire recognition big data language model are such as "There is a forest in the picture, some tree areas show orange-red light, accompanied by gray-black smoke".
[0086] The environmental element recognition command prompts the fire identification big data model to analyze environmental information in scene images. This environmental information includes image metadata such as time, geographical location, and equipment number, as well as environmental status data such as weather, wind speed, and rainfall. By introducing environmental elements, the fire identification big data model can help determine the likelihood of a fire. For example, the risk of a fire is lower in rainy conditions, while the risk increases significantly in dry, windy conditions. The environmental element recognition command can be customized based on actual needs, such as "Please retrieve the current weather conditions for the area where this scene is located and analyze it in conjunction with the image." The scene recognition result output by the fire identification big data model might be something like, "Based on geographical location and time information, the local weather is sunny with high wind speed and no rainfall; combined with the image, the flames and smoke are more likely to be a real fire than a natural weather phenomenon."
[0087] Fire feature recognition instructions are used to prompt the fire recognition big data model to extract typical fire features from scene images and compare and distinguish them with common interference factors. Typical fire features include flames and smoke; common interference factors include morning fog, cooking smoke, and light reflection. This ensures the reliability of fire feature recognition and reduces false alarms caused by environmental interference. Fire feature recognition instructions can be set according to actual needs, such as "Please determine whether there are typical fire features (flames, smoke, etc.) in the image, and explain the differences between them and common environmental interference factors (such as morning fog, cooking smoke, and light reflection)." The scene recognition results output by the fire recognition big data model are as follows: "Orange-red flames and rising black smoke are visible in the image; compared with morning fog, its shape is uneven and its color is darker; compared with cooking smoke, its scale is larger and its diffusion speed is faster; compared with light reflection, its edges are irregular and accompanied by smoke diffusion."
[0088] In practice, you can start from a broad level and then refine it to a smaller level. For example, first input the scene recognition command, then the environmental element recognition command, and then the fire feature recognition command. After that, you can get the final fire recognition result by inputting the summary command. The summary command can be set according to actual needs, such as "Based on the above analysis, please make a comprehensive judgment on whether there is a fire in this image and give the reason." The fire recognition big language model outputs the fire recognition result as "Yes, there is a fire in this image, based on the following: a large area of orange-red flames and black smoke, combined with clear weather and wind speed conditions, which are consistent with the typical characteristics of a forest fire."
[0089] In this embodiment, the design of multi-turn dialogue enables the fire identification big language model to gradually construct the fire judgment logic from different angles and dimensions, thereby improving the accuracy of fire identification results.
[0090] Furthermore, in the third embodiment of the fire identification method of the present invention based on the first embodiment, the method further includes the step of:
[0091] Step S14: Obtain model training samples, wherein the model training samples include scene image samples, recognition instruction samples, and expected result samples;
[0092] Step S15: Obtain the initial fire identification large language model;
[0093] Step S16: Input the scene image sample and the recognition instruction sample into the initial fire recognition big language model to obtain the training recognition result output by the initial fire recognition big language model;
[0094] Step S17: Compare the training recognition result with the expected result sample to obtain the result difference;
[0095] Step S18: Update the initial fire identification language model based on the result difference to obtain the fire identification language model.
[0096] The model training samples are used to train the initial fire identification large language model.
[0097] The model training samples consist of multiple sets of training data, and the training data consists of triples (x, I, A). exp It consists of: x represents a scene image sample; I represents a recognition instruction sample designed for that scene image sample; A expThis represents the expected result sample corresponding to the identification instruction sample. Specifically, it can be labeled by experts. When setting up training data, positive and negative samples can be set. Positive samples indicate the conclusion of a fire and supporting reasons, such as "There is a fire; the image shows obvious orange-red flames and thick smoke, which matches typical fire characteristics." Negative samples indicate the conclusion of no fire and supporting reasons, such as "There is no fire; the red clouds in the image are a sunset scene, lacking flame and smoke characteristics." The large-scale model fine-tuning uses the Supervised Fine-Tuning (SFT) method. To ensure the effectiveness of fine-tuning training, the model training samples need to cover fire characteristics in different scenarios, such as common interference factors like sunlight and industrial steam. The scale of the model training samples should be large enough to ensure that the model can learn rich fire discrimination knowledge and cope with interference factors in different environments. In terms of quality, ensure that the annotation of each image is correct and consistent, avoid labeling errors, and ensure a balance between positive and negative samples.
[0098] The specific training method can be set according to actual needs. For example, SFT (Supervised Fine-Tuning) can be used. Its basic principle is to combine visual input with natural language instructions to construct sample pairs of recognition instructions and expected results. Through training, the large model can generate expected answers given scene images and recognition instructions. The goal of training is to guide the model to learn the multimodal association between scene images and recognition instructions, so that it can not only identify fire features such as flames and smoke in its answers, but also provide reasonable explanations based on environmental interference factors, thereby achieving a comprehensive discrimination ability of "recognition + reasoning + explanation".
[0099] The result difference indicates the discrepancy between the output of the fire identification big language model and the expected output. Therefore, updating the fire identification big language model based on the result difference can bring it closer to the desired output and minimize the difference between the training recognition result predicted by the model and the expected result sample provided by the expert. Specifically, the fire identification big language model can be updated by setting a loss function, which can be set according to actual needs.
[0100] Further, step S18 includes the following steps:
[0101] Step S181: Determine whether there are preset keywords in the training recognition results or the expected result samples;
[0102] Step S182: If the training recognition result or the expected result sample contains preset keywords, then update the initial fire recognition big language model according to the result difference.
[0103] In this embodiment, preset keywords are set, which are keywords related to the fire domain, such as the preset keyword set K={fire, flame, smoke, non-fire, interference, ...}. If the preset keywords are present in the training recognition results or expected result samples, it is considered that key factors related to fire are involved; therefore, this part needs to be strictly constrained. The specific loss function can be expressed as:
[0104]
[0105] in, This represents the probability of the i-th word generated by the model in round t; 1 represents the nth word expected in the nth round; Ⅱ(∙) is an indicator function that takes the value 1 only when the preset keyword belongs to the preset keyword set K, and 0 otherwise.
[0106] For example, the expected result sample is "There is a fire, and flames and smoke appear in the image, which is consistent with the characteristics of a typical forest fire";
[0107] The training recognition result was "There are bright spots in the image, but no smoke was detected".
[0108] At this point, both "flame" and "smoke" belong to set K. Therefore, the loss function imposes strong constraints on errors in the positions of these two words, while not imposing constraints on non-keywords such as "in the image" and "bright light spots." In this way, training can focus on guiding the model to correctly learn the generation of key semantics in the domain, thereby ensuring the reliability of the final answer in terms of professionalism and completeness.
[0109] When the initial fire recognition language model reaches the training completion condition after training, the trained fire recognition language model is obtained. The specific training completion condition can be set according to actual needs, such as hyperparameters like learning rate and batch size, or training optimization objectives. The training optimization objective can be set as accurately identifying the core visual features of flames and smoke, such as color, shape, texture, and dynamics; understanding the manifestation of "fire" in specific scenarios, such as the different characteristics of forest fires, urban fires, and industrial fires; distinguishing real fires from common interferences in corresponding scenarios, such as red clouds, warm light, water vapor, dust, and reflections, to reduce false alarms; and generating judgment reasons that conform to the professional logic of fire recognition, such as "Fire confirmed: the flames in the image are orange-red, and the smoke is dense, which is consistent with the characteristics of a forest fire." When the initial fire recognition language model meets the set training optimization objectives, it is considered to have met the training completion condition.
[0110] It is understandable that after training, the fire identification language model may make recognition errors in applications due to complex scenarios or interference factors. Therefore, the fire identification language model can be corrected and updated based on actual detection results during application. Specifically, this can include:
[0111] Error judgment identification: If the fire identification result is different from the actual scene, the answer will be judged as incorrect by a human.
[0112] Error cause analysis: Identify possible sources of misjudgment in the fire identification big data model, such as lighting interference and morphological similarity, and use relevant data, such as scene images, recognition commands, and fire identification results, as samples for misjudgment.
[0113] Error correction process: By comparing the fire identification results with the real scene, the model is guided to update its inference logic; this can be performed analogously to the training process.
[0114] Incremental training: Misclassified samples are recorded and fed back into the model's training data. Through incremental training, these misclassified samples are used to retrain the model, enabling it to reduce similar errors in future judgments.
[0115] Multi-round correction mechanism: In multi-round dialogues, if the model makes a misjudgment in a certain round of response, the system will adjust in subsequent rounds. For example, if the model misjudges the existence of a fire in the second round of response, subsequent responses can guide the model to re-analyze previous responses based on expert logic, thereby correcting the error and enhancing its understanding of the scenario.
[0116] Automatic feedback loop: To ensure continuous system optimization, all misjudgments and corrections are recorded within the system, forming a feedback mechanism. Misjudged samples are periodically generated to update and fine-tune the system, ensuring the model improves its accuracy through continuous operation.
[0117] Furthermore, in the fourth embodiment of the fire identification method of the present invention based on the first embodiment of the present invention, step S21 includes the following steps:
[0118] Step S25: Obtain the scenario type of the fire identification lightweight model application;
[0119] Step S26: Obtain model training samples corresponding to the scene type;
[0120] Step S27: Obtain the initial lightweight fire identification model;
[0121] Step S28: Train the initial lightweight fire identification model using the model training samples to obtain the lightweight fire identification model.
[0122] It is understandable that the state of a fire will differ depending on the type of scenario. The specific scenario type can be set based on actual needs, such as forest, city, industry, and oil field.
[0123] Since the edge terminal is deployed in the front-end system close to the fire monitoring site, the fire identification of the edge terminal directly corresponds to the scene type. Therefore, in this embodiment, corresponding model training samples are set for different scene types, so that a lightweight fire identification model can be trained for fire identification in a specific scene.
[0124] The specific training samples for the model can be set based on the characteristics of the actual scenario; for example:
[0125] If the scene type is forest, images of forest fires will be collected, with a focus on the dynamic characteristics of flames and smoke. Data will be collected using drones and ground-based high-point monitoring cameras, covering images under different weather conditions.
[0126] The scenario type is urban, and the urban fire data mainly comes from building fires, traffic fires, etc. Samples are collected using urban surveillance cameras and traffic monitoring equipment to ensure that images are covered under different lighting and environmental interference conditions.
[0127] The scenario type is industrial. Fire characteristics in industrial environments differ from other scenarios, primarily focusing on high-temperature equipment and chemical leaks. Monitoring equipment in the industrial area is used to collect data on fires involving different equipment, with particular emphasis on oil fires, smoke, and other specific fire characteristics.
[0128] If the scenario is an oil field, then oil field fires typically occur in areas such as oil wells and storage tanks, characterized by intense flames and dense smoke. We collect fire images in these specific environments using specialized monitoring equipment and drones, with data covering the effects of special factors such as high temperatures and oil vapors.
[0129] Data collection in different scenarios can be carried out through different terminal devices, ensuring the diversity and breadth of data. For example, drones are suitable for large areas such as forests and oil fields, providing fire images at different heights and angles to capture fire information over a wide area. Ground-based fixed camera data is suitable for urban and industrial fire data collection. These devices provide stable, localized fire images, helping models to more accurately identify local fire characteristics. Mobile inspection equipment, such as fire inspection robots, robot dogs, and vehicle-mounted monitoring systems, can effectively supplement fixed monitoring blind spots and are suitable for fire data acquisition in scenarios such as industrial facilities and complex urban areas, enhancing the system's ability to perceive real-time fire changes.
[0130] To improve the reliability of the samples, image samples can be enhanced after collection, such as through noise reduction and illumination adjustment. The images can also be standardized to improve data quality and ensure that the model can effectively learn the fire characteristics under different environments during training.
[0131] After acquiring image samples, structured annotation is required to form labeled data pairs; for example, for each image sample i, a label pair L is set. i :
[0132]
[0133] Among them, y i ∈{0, 1}, indicating whether a fire exists, where 1 indicates existence and 0 indicates non-existence. S i ∈S represents the environmental scene label to which the image sample belongs, S={forest, city, industry, oil field}.
[0134] For each image sample, in addition to labeling whether a fire exists, images containing interfering content, such as sunlight and smoke, can also be labeled to avoid misjudgment by the model.
[0135] Further, step S28 includes the following steps:
[0136] Step S281: Perform image degradation enhancement on the model training samples to obtain degraded training samples;
[0137] Step S282: Train the initial fire identification lightweight model using the degraded training samples.
[0138] In real-world monitoring scenarios, images can be affected by various environmental factors, such as drone motion blur, direct sunlight overexposure, and dust / haze interference. These factors all affect the sensor's visual imaging quality, interfere with the model's recognition accuracy, leading to an increased false alarm rate and reduced recognition efficiency. Therefore, to improve the model's robustness and accuracy, image degradation enhancement processing can be applied to the model's training samples. The purpose of image degradation enhancement is to simulate interference conditions in real-world scenarios and enhance the adaptability of the lightweight fire recognition model to these interferences. For example, let the original image be I... raw The degraded and enhanced image is I aug The degenerate transformation function is τ(∙;θ aug Then, the degraded training samples after image degradation enhancement are:
[0139]
[0140] θ aug Parameters for degradation enhancement.
[0141] The specific model for image degradation enhancement can be set based on the type of interference in the scene; for example, a motion blur model can be set to address motion blur interference; or an image stretching and blurring caused by the rapid movement of a drone or camera can be simulated. Assume K... motion Given a linear motion kernel, the image degradation model is as follows:
[0142]
[0143] Where I' is the image after motion blur processing; K motion For linear motion kernels, length l:
[0144]
[0145] Where θ is the direction of motion, v is the speed of movement, t is the shutter speed, d is the target distance, w is the image width, and FOV is the field of view.
[0146] Set up a fog interference model to address fog interference;
[0147] Based on an atmospheric scattering model, the reduction in visibility caused by haze, dust, or smog is simulated. The impact of haze can be simulated using the following formula:
[0148]
[0149]
[0150] Where t(x) is the transmittance, β∈[0.5,1.2] is the scattering coefficient, d(x) is the pixel depth, and A∈[0.8,1.0] is the atmospheric light intensity.
[0151] Set up a solar overexposure model to address strong light interference;
[0152] To simulate strong light interference, a radial mask is used to generate highlight areas. Overexposure correction can be performed using the following formula:
[0153]
[0154] Where M(x) is a radial mask that gradually changes from the center outward, η is the overexposure intensity (e.g., 60~100), and 255 is the maximum pixel value, ensuring that some areas of the image produce saturated highlights.
[0155] Through these degradation enhancement models, we can effectively simulate the impact of various environmental factors on image quality, thereby training a more robust initial fire recognition lightweight model and reducing false alarms and false negatives in practical applications.
[0156] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0157] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0158] This application also provides a fire identification system for implementing the above-described fire identification method, the fire identification system comprising a cloud server and an edge terminal; the cloud server is used for:
[0159] Acquire scene images uploaded from the edge terminal and input the scene images into the fire recognition language model, wherein the scene images are uploaded when a fire is detected at the edge terminal;
[0160] Input multi-round recognition instructions to the fire recognition big language model, so that the fire recognition big language model outputs recognition sub-results corresponding to the multi-round recognition instructions based on the scene image, wherein the multi-round recognition instructions are for different objects in the scene image;
[0161] Input a summary command into the fire identification language model so that the fire identification language model integrates multiple identification sub-results to obtain a fire identification result;
[0162] The edge end is used for:
[0163] Acquire scene images and input the scene images into a lightweight fire identification model;
[0164] Obtain the edge results corresponding to the scene image output by the lightweight fire recognition model;
[0165] Determine whether the edge result indicates a fire has occurred;
[0166] If the edge result indicates a fire has occurred, the scene image will be uploaded to the cloud server.
[0167] This system uses a large language model for fire identification, leveraging its multimodal capabilities to improve accuracy. By progressively obtaining identification results through multiple rounds of instructions, the system guides the fire identification language model to fully perceive the content of the scene image, thereby enhancing its reasoning depth and discrimination accuracy. Simultaneously, by performing fire identification on scene images both on the cloud server and at the edge, collaborative detection between the cloud and edge is achieved, resulting in a dual improvement in fire identification accuracy and interpretability.
[0168] Furthermore, the input of multi-round recognition instructions to the fire recognition large language model includes:
[0169] Input scene recognition instructions into the fire recognition big language model, so that the fire recognition big language model outputs scene recognition results based on the scene image;
[0170] An environmental element recognition instruction is input into the fire recognition big language model, so that the fire recognition big language model outputs environmental element recognition results based on the scene image;
[0171] Input fire feature recognition instructions into the fire recognition big language model so that the fire recognition big language model outputs fire feature recognition results based on the scene image.
[0172] Furthermore, the method also includes:
[0173] Obtain model training samples, wherein the model training samples include scene image samples, recognition instruction samples, and expected result samples;
[0174] Obtain an initial large language model for fire identification;
[0175] The scene image samples and the recognition instruction samples are input into the initial fire recognition large language model to obtain the training recognition result output by the initial fire recognition large language model;
[0176] The training recognition results are compared with the expected result samples to obtain the result differences;
[0177] The initial fire identification language model is updated based on the differences in the results to obtain the fire identification language model.
[0178] Furthermore, updating the initial fire identification large language model based on the result differences includes:
[0179] Determine whether a preset keyword exists in the training recognition result or the expected result sample;
[0180] If the training recognition results or the expected result samples contain preset keywords, the initial fire recognition big language model is updated based on the result differences.
[0181] Furthermore, the step of inputting the scene image into the lightweight fire recognition model includes:
[0182] Obtain the scenario type for the application of the lightweight fire identification model;
[0183] Obtain model training samples corresponding to the scene type;
[0184] Obtain an initial lightweight fire identification model;
[0185] The initial lightweight fire identification model is trained using the model training samples to obtain the lightweight fire identification model.
[0186] Further, training the initial fire identification lightweight model using the model training samples includes:
[0187] The training samples of the model are subjected to image degradation enhancement to obtain degraded training samples;
[0188] The initial lightweight fire identification model is trained using the degraded training samples.
[0189] Reference Figure 3 In terms of hardware structure, the electronic device may include components such as a communication module 10, a memory 20, and a processor 30. In the electronic device, the processor 30 is connected to both the memory 20 and the communication module 10. The memory 20 stores a computer program, which is executed by the processor 30. When the computer program is executed, it implements the steps of the above-described method embodiments.
[0190] The communication module 10 can connect to external communication devices via a network. The communication module 10 can receive requests from the external communication devices and can also send requests, instructions, and information to the external communication devices. The external communication devices can be other electronic devices, servers, or IoT devices, such as televisions, etc.
[0191] The memory 20 can be used to store software programs and various data. The memory 20 may primarily include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as acquiring scene images uploaded from the edge and inputting the scene images into a fire identification large language model), etc.; the data storage area may include a database, and may store data or information created based on system usage. Furthermore, the memory 20 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0192] The processor 30 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 20, and by calling data stored in the memory 20, it performs various functions and processes data, thereby providing overall monitoring of the electronic device. The processor 30 may include one or more processing units; optionally, the processor 30 may integrate an application processor and a modem processor. The application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 30.
[0193] although Figure 3 Not shown, but the above-described electronic device may further include a circuit control module for connecting to a power supply to ensure the normal operation of other components. Those skilled in the art will understand that... Figure 3 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0194] The present invention also proposes a computer-readable storage medium having a computer program stored thereon. The computer-readable storage medium may be... Figure 3 The memory 20 in the electronic device may also be at least one of ROM (Read-Only Memory) / RAM (Random Access Memory), magnetic disk, optical disk, etc. The computer-readable storage medium includes a number of instructions to cause a terminal device with a processor (which may be a television, automobile, mobile phone, computer, server, terminal, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0195] In this invention, the terms "first," "second," "third," "fourth," and "fifth" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0196] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0197] Although embodiments of the present invention have been shown and described above, the scope of protection of the present invention is not limited thereto. It is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, and substitutions to the above embodiments within the scope of the present invention, and such changes, modifications, and substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A fire identification method characterized by, The fire identification method applied to the cloud server comprises: obtaining a scene image uploaded by an edge terminal and inputting the scene image into a fire identification large language model, wherein the scene image is uploaded when the edge terminal detects a fire; inputting a plurality of rounds of identification instructions into the fire identification large language model to enable the fire identification large language model to output identification sub-results corresponding to the plurality of rounds of identification instructions according to the scene image, wherein the plurality of rounds of identification instructions are directed to different objects in the scene image; inputting a summary instruction into the fire identification large language model to enable the fire identification large language model to obtain a fire identification result by synthesizing a plurality of identification sub-results; The fire identification method applied to the edge terminal comprises: obtaining a scene image and inputting the scene image into a fire identification light model; obtaining an edge result corresponding to the scene image output by the fire identification light model; determining whether the edge result is a fire; if the edge result is a fire, uploading the scene image to a cloud server.
2. The fire identification method of claim 1, wherein, The inputting of the plurality of rounds of identification instructions into the fire identification large language model comprises: inputting a scene identification instruction into the fire identification large language model to enable the fire identification large language model to output a scene identification result according to the scene image; inputting an environmental element identification instruction into the fire identification large language model to enable the fire identification large language model to output an environmental element identification result according to the scene image; inputting a fire feature identification instruction into the fire identification large language model to enable the fire identification large language model to output a fire feature identification result according to the scene image.
3. The fire identification method of claim 1, wherein, The method further comprises: obtaining a model training sample, wherein the model training sample comprises a scene image sample, an identification instruction sample, and an expected result sample; obtaining an initial fire identification large language model; inputting the scene image sample and the identification instruction sample into the initial fire identification large language model to obtain a training identification result output by the initial fire identification large language model; comparing the training identification result with the expected result sample to obtain a result difference; updating the initial fire identification large language model according to the result difference to obtain the fire identification large language model.
4. The fire identification method of claim 3, wherein, The updating of the initial fire identification large language model according to the result difference comprises: determining whether there is a preset keyword in the training identification result or the expected result sample; if there is a preset keyword in the training identification result or the expected result sample, updating the initial fire identification large language model according to the result difference.
5. The fire identification method of claim 1, wherein, The inputting of the scene image into the fire identification light model comprises: obtaining a scene type to which the fire identification light model is applied; obtaining a model training sample corresponding to the scene type; obtaining an initial fire identification light model; training the initial fire identification light model by using the model training sample to obtain the fire identification light model.
6. The fire identification method of claim 5, wherein, The training of the initial fire identification light model by using the model training sample comprises: The model training sample is subjected to image degradation enhancement to obtain a degraded training sample; The initial fire identification lightweight model is trained through the degraded training sample.
7. A fire detection system characterized by, The fire identification system comprises a cloud server and an edge terminal; the cloud server is configured to: obtain a scene image uploaded by the edge terminal and input the scene image into a fire identification large language model, wherein the scene image is uploaded when the edge terminal detects a fire; input a multi-round identification instruction into the fire identification large language model to enable the fire identification large language model to output an identification sub-result corresponding to the multi-round identification instruction according to the scene image, wherein the multi-round identification instruction is directed to different objects in the scene image; input a summary instruction into the fire identification large language model to enable the fire identification large language model to obtain a fire identification result by synthesizing multiple identification sub-results; the edge terminal is configured to: obtain a scene image and input the scene image into a fire identification lightweight model; obtain an edge result corresponding to the scene image output by the fire identification lightweight model; determine whether the edge result indicates a fire; if the edge result indicates a fire, upload the scene image to the cloud server.
8. An electronic device, comprising: The electronic device comprises a memory, a processor and a computer program stored on the memory and executable on the processor, and the computer program is executed by the processor to implement the steps of the fire identification method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps of the fire identification method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Hidden disaster hazard identification method, system and device, and storage medium
CN119807817A
Fire monitoring method, fire monitoring server and fire monitoring system
CN120412183A