Fire identification method and system, electronic equipment and computer readable storage medium

A fire detection system that works collaboratively between cloud servers and edge devices utilizes large language models and lightweight models for multi-round recognition instruction guidance, solving the accuracy and real-time problems of traditional fire detection and achieving efficient and low-cost fire detection.

CN120997774AActive Publication Date: 2025-11-21SHENZHEN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511514812.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2025-11-21
Estimated Expiration
2045-10-22

AI Technical Summary

Technical Problem

Traditional fire detection methods struggle to detect fires promptly and accurately in complex environments, and the equipment is expensive and has a limited monitoring range.

Method used

The fire identification system, which uses cloud servers and edge devices to work together, uses a large fire identification language model and a lightweight model to guide multi-round identification instructions, and combines scene images to identify fires. The system gradually obtains identification results through the multimodal capabilities of the large language model and multi-round identification instructions.

Benefits of technology

It improves the accuracy and interpretability of fire identification, achieves real-time performance and accuracy and robustness in complex environments, and reduces deployment costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997774A_ABST
    Figure CN120997774A_ABST
Patent Text Reader

Abstract

According to the fire identification method and system, the electronic equipment and the computer readable storage medium provided by the invention, fire identification is carried out through the large language model, so that the fire identification accuracy can be improved by utilizing the multi-modal capability of the large language model, the identification result is gradually obtained through multiple rounds of identification instructions, and the fire identification efficiency is improved. Therefore, the fire identification large language model can fully perceive the content in the scene image in a guiding manner, thereby enhancing the reasoning depth and discrimination accuracy of the fire identification large language model. And meanwhile, fire recognition is performed on the scene image at the cloud server and the edge end, so that cooperative detection of the cloud end and the edge end can be constructed, and double improvement of fire recognition accuracy and interpretability is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of disaster monitoring, and in particular to a fire identification method and system, an electronic device, and a computer readable storage medium. BACKGROUND

[0002] Traditional fire detection is usually achieved based on a monitoring system with sensors; however, this way is often difficult to discover fire in time and accurately in complex environments, and has limitations such as high equipment cost and limited monitoring range. SUMMARY

[0003] The main purpose of the present application is to provide a fire identification method, system, electronic device and computer readable storage medium, which aims to solve the problem of poor accuracy and timeliness of fire detection in the prior art.

[0004] To achieve the above-mentioned purpose, the present application provides a fire identification method applied to a cloud server, comprising the steps of: obtaining a scene image uploaded by an edge terminal and inputting the scene image into a fire identification large language model, wherein the scene image is uploaded when the edge terminal detects a fire; inputting a plurality of rounds of identification instructions into the fire identification large language model, so that the fire identification large language model outputs identification sub-results corresponding to the plurality of rounds of identification instructions according to the scene image, wherein the plurality of rounds of identification instructions are directed to different objects in the scene image; inputting a summary instruction into the fire identification large language model, so that the fire identification large language model synthesizes a plurality of identification sub-results to obtain a fire identification result.

[0005] Optionally, the inputting a plurality of rounds of identification instructions into the fire identification large language model comprises: inputting a scene identification instruction into the fire identification large language model, so that the fire identification large language model outputs a scene identification result according to the scene image; inputting an environmental element identification instruction into the fire identification large language model, so that the fire identification large language model outputs an environmental element identification result according to the scene image; inputting a fire feature identification instruction into the fire identification large language model, so that the fire identification large language model outputs a fire feature identification result according to the scene image.

[0006] Optionally, the method further comprises: obtaining a model training sample, wherein the model training sample comprises a scene image sample, an identification instruction sample, and an expected result sample; obtaining an initial fire identification large language model; inputting the scene image sample and the identification instruction sample into the initial fire identification large language model to obtain a training identification result output by the initial fire identification large language model; comparing the training identification result with the expected result sample to obtain a result difference; updating the initial fire identification large language model according to the result difference to obtain the fire identification large language model.

[0007] Optionally, the updating the initial fire identification large language model according to the result difference comprises: judging whether there is a preset keyword in the training identification result or the expected result sample; if there is a preset keyword in the training identification result or the expected result sample, updating the initial fire identification large language model according to the result difference.

[0008] To achieve the above-mentioned purpose, the present application further provides a fire identification method applied to an edge end, which comprises: acquiring a scene image and inputting the scene image into a fire identification light model; acquiring an edge end result corresponding to the scene image output by the fire identification light model; judging whether the edge end result is a fire occurrence; if the edge end result is a fire occurrence, uploading the scene image to a cloud server.

[0009] Optionally, the method further comprises, before the inputting the scene image into the fire identification light model: acquiring a scene type applied by the fire identification light model; acquiring a model training sample corresponding to the scene type; acquiring an initial fire identification light model; training the initial fire identification light model through the model training sample to obtain the fire identification light model.

[0010] Optionally, the training the initial fire identification light model through the model training sample comprises: performing image degradation enhancement on the model training sample to obtain a degradation training sample; training the initial fire identification light model through the degradation training sample.

[0011] To achieve the above-mentioned purpose, the present application further provides a fire identification system, which comprises a cloud server and an edge end; the cloud server is used for: Acquire a scene image uploaded on an edge terminal, and input the scene image into a fire identification large language model, wherein the scene image is uploaded when the edge terminal monitors a fire; Input a multi-round identification instruction into the fire identification large language model, so that the fire identification large language model outputs an identification sub-result corresponding to the multi-round identification instruction according to the scene image, wherein the multi-round identification instruction is for different objects in the scene image; Input a summary instruction into the fire identification large language model, so that the fire identification large language model synthesizes a plurality of identification sub-results to obtain a fire identification result; The edge terminal is used for: Acquire a scene image, and input the scene image into a fire identification light model; Acquire an edge result corresponding to the scene image output by the fire identification light model; Determine whether the edge result is a fire occurrence; If the edge result is a fire occurrence, upload the scene image to a cloud server.

[0012] To achieve the above object, the present application further provides an electronic device, which comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the computer program implements the steps of the fire identification method when executed by the processor.

[0013] To achieve the above object, the present application further provides a computer readable storage medium, which stores a computer program, and the computer program implements the steps of the fire identification method when executed by a processor.

[0014] The fire identification method, system, electronic device and computer readable storage medium provided by the present application are as follows: a scene image uploaded by an edge end is acquired, and the scene image is input into a fire identification large language model, wherein the scene image is uploaded when the edge end monitors a fire; a plurality of rounds of identification instructions are input into the fire identification large language model, so that the fire identification large language model outputs an identification sub-result corresponding to the plurality of rounds of identification instructions according to the scene image, wherein the plurality of rounds of identification instructions are directed to different objects in the scene image; a summary instruction is input into the fire identification large language model, so that the fire identification large language model comprehensively obtains a fire identification result from a plurality of identification sub-results. The fire identification is performed by using the large language model, so that the multi-modal capability of the large language model can be used to improve the identification accuracy of the fire. The identification result is gradually obtained by using the plurality of rounds of identification instructions, so that the fire identification large language model can be guided to fully perceive the content in the scene image, thereby enhancing the reasoning depth and discrimination accuracy of the fire identification large language model. At the same time, the scene image is identified by using the cloud server and the edge end, so that the collaborative detection of the cloud and the edge can be constructed, and the dual improvement of the fire identification accuracy and the interpretability can be realized. BRIEF DESCRIPTION OF DRAWINGS

[0015] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the application.

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, those skilled in the art can obtain other drawings according to these drawings without any creative effort.

[0017] Figure 1 The flowchart of the first embodiment of the fire identification method of the present application is shown in the figure. Figure 2 The detailed flowchart of the fire identification method of the present application is shown in the figure. Figure 3 The module structure diagram of the electronic device of the present application is shown in the figure. DETAILED DESCRIPTION

[0018] It should be understood that the specific embodiments described herein are merely illustrative of the present application and are not intended to limit the present application. In order to enable persons skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by persons skilled in the art without creative labor should be within the scope of protection of the present application.

[0019] The present application provides a fire identification method, referring to Figure 1 , Figure 1 The flowchart of the first embodiment of the fire identification method of the present application is applied to a cloud server, and the method comprises the following steps: In step S11, a scene image uploaded by an edge end is acquired, and the scene image is input into a fire identification large language model, wherein the scene image is uploaded when the edge end detects a fire; The scene image is an image collected by an image collection device arranged at the edge end.

[0020] The edge end is arranged at a front-end system close to a fire monitoring site; the edge end can be specifically provided with an image collection device, a computing device, etc.; the image collection device is, for example, a camera; the computing device can specifically perform information processing work of the edge end; the number of edge ends can be specifically set based on actual needs, for example, multiple edge ends can be set to monitor multiple different fire monitoring sites.

[0021] The cloud server is a high-performance computing platform arranged in the cloud; a communication connection is established between the cloud server and the edge end.

[0022] The fire identification large language model is a model arranged in the cloud server for generating a final fire identification result of a scene image; the specific type of the fire identification large language model can be set based on actual needs, for example, LLaVA, BLIP-2, Qwen-VL.

[0023] In the present embodiment, the edge end collects a scene image in real time through a camera, and performs fire identification according to the scene image; when a fire is detected, the scene image is uploaded to the cloud server, and further fire detection is realized through the fire identification large language model.

[0024] It can be understood that, since the edge end only performs preliminary fire detection, a light-weight model can be deployed at the edge end to realize fire detection, thereby improving detection efficiency and reducing deployment cost.

[0025] In step S12, a multi-round recognition instruction is input to the fire recognition large language model, so that the fire recognition large language model outputs a recognition sub-result corresponding to the multi-round recognition instruction according to the scene image, wherein the multi-round recognition instruction is for different objects in the scene image. The recognition instruction is a natural language guide instruction set for fire recognition, such as a prompt word. After the recognition instruction is input to the fire recognition large language model, the fire recognition large language model generates a corresponding recognition sub-result according to the recognition instruction. If the recognition instruction is "please judge whether there is fire in the picture and explain the reason", the fire recognition large language model will output a detection answer whether there is fire and the reason for the answer in the recognition sub-result.

[0026] In this embodiment, a multi-round recognition instruction is set, that is, the recognition instruction is input multiple times to make the fire recognition large language model output a multi-round corresponding recognition sub-result. Since the multi-round recognition instruction corresponds to different objects in the scene image, when the fire recognition large language model outputs the recognition sub-result for different recognition instructions, it can perform semantic analysis for different objects, so that through multi-round analysis, the fire recognition large language model can more comprehensively identify the scene image, and can clearly understand the specific state of different types of objects in the scene image, thereby providing a reliable basis for the final output of the subsequent fire recognition result.

[0027] In step S13, a summary instruction is input to the fire recognition large language model, so that the fire recognition large language model synthesizes multiple recognition sub-results to obtain a fire recognition result.

[0028] The summary instruction is used to instruct the fire recognition large language model to output the final fire recognition result. The fire recognition large language model has the ability to link the context, so after receiving the summary instruction, it will output the fire recognition result by synthesizing the recognition sub-results obtained by the multi-round recognition instruction.

[0029] The fire recognition result is the result obtained by the cloud server indicating the occurrence of fire. The content contained in the fire recognition result can be set based on actual needs, such as whether the fire occurs, the location of the fire, the situation of the fire, the cause analysis, etc.

[0030] In specific implementation, the scene image can be converted into a visual feature vector, and multiple recognition sub-results are generated based on corresponding recognition instructions in sequence, and then the fire recognition result is obtained by judging according to a preset rule.

[0031] According to the fire identification result, a response strategy can be specifically executed; when the fire identification result indicates that a fire occurs, a platform-level fire warning is triggered, including pushing an alarm to a command center, linking a camera to focus, starting a fire-fighting robot, and the like; when the fire identification result indicates that no fire occurs, a false positive sample is recorded and automatically archived; when the fire identification result is not uncertain, a background manual review is guided, or the sample is marked as a to-be-trained sample and added to a subsequent model incremental learning library.

[0032] The cloud server can periodically return the reviewed samples and their semantic paths to the edge side to update the model parameters of the fire identification lightweight model or optimize the feature extraction logic, thereby realizing continuous iterative optimization of cloud-edge joint learning.

[0033] The fire identification method applied to the edge side includes the following steps. In step S21, a scene image is acquired and input into a fire identification lightweight model. In step S22, an edge-side result corresponding to the scene image output by the fire identification lightweight model is acquired. In step S23, it is determined whether the edge-side result indicates that a fire occurs. In step S24, if the edge-side result indicates that a fire occurs, the scene image is uploaded to a cloud server.

[0034] If the edge-side result indicates that no fire occurs, the scene image is not uploaded to the cloud server. In other embodiments, the uploading of the scene image can be maintained at all times, and when the edge-side result indicates that a fire occurs, an identifier indicating that a fire occurs is associated with the scene image and uploaded to the cloud server; when the edge-side result indicates that no fire occurs, an identifier indicating that no fire occurs is associated with the scene image and uploaded to the cloud server.

[0035] As can be understood, since the edge side only performs preliminary fire detection, a lightweight model can be deployed on the edge side to realize fire detection, thereby improving detection efficiency and reducing deployment costs. The specific type of the fire identification lightweight model can be set based on actual needs, such as MobileNet, EfficientNet-lite, and YOLO-Lite. The fire identification lightweight model outputs a fire probability based on a scene image, and if the fire probability is greater than a fire threshold, the edge-side result indicates that a fire occurs; if the fire probability is less than or equal to the fire threshold, the edge-side result indicates that no fire occurs.

[0036] The edge-side result is an identification result output by the fire identification lightweight model based on the scene image.

[0037] If the edge result indicates that a fire occurs, it indicates that there may be a fire at this time, so the scene image is uploaded to the cloud server to enable the cloud server to determine the final fire identification result based on the scene image.

[0038] When the edge result indicates that a fire occurs, the edge continues to upload the real-time collected scene image to the cloud server to enable the cloud server to continuously identify the fire based on the continuous scene image; when the cloud server determines that there is no fire, a misjudgment instruction is sent to the cloud server, the cloud server stops uploading the scene image, and then still maintains the collection of the scene image and performs fire identification on the scene image through the fire identification lightweight model.

[0039] The embodiment specifically proposes a cloud-edge collaborative fire identification mechanism of edge short thinking-cloud long thinking. The short thinking refers to the rapid preliminary screening and efficient response of the fire identification lightweight model of the edge to the large-scale monitored scene image, mainly realizing low-latency suspected fire detection; the long thinking refers to the deep semantic analysis and multi-round logical reasoning of the fire identification large language model set by the cloud server, providing high-precision and interpretable final determination. The two complement each other through a collaborative verification mechanism, which guarantees real-time performance and takes into account accuracy and robustness in complex environments.

[0040] The embodiment uses a large language model to identify a fire, which can improve the identification accuracy of the fire by using the multi-modal capability of the large language model, and gradually obtains the identification result through multiple rounds of identification instructions, so as to guide the fire identification large language model to fully perceive the content in the scene image, thereby enhancing the reasoning depth and discrimination accuracy of the fire identification large language model; at the same time, the scene image is identified by the cloud server and the edge, so that collaborative detection of the cloud and the edge can be constructed to realize double improvement of fire identification accuracy and interpretability.

[0041] Further, in combination with Figure 2 In the second embodiment of the fire identification method of the present application based on the first embodiment of the present application, the step S11 comprises the steps of: Step S111, inputting a scene identification instruction to the fire identification large language model, so that the fire identification large language model outputs a scene identification result according to the scene image; Step S112, inputting an environment element identification instruction to the fire identification large language model, so that the fire identification large language model outputs an environment element identification result according to the scene image; Step S113, inputting a fire feature identification instruction to the fire identification large language model, so that the fire identification large language model outputs a fire feature identification result according to the scene image.

[0042] The scene recognition instruction is used to prompt the fire identification large language model to output a language description of the overall environment and objects of the scene image, so as to extract scene information and significant visual elements in the scene image; so that the fire identification large language model can establish preliminary image semantic cognition and form the basis for subsequent reasoning; the scene recognition instruction can be set based on actual needs, such as “please describe the main objects and scenes in the picture”; the scene recognition result output by the fire identification large language model is, for example, “there is a forest in the picture, and part of the tree area appears orange-red light with gray-black smoke”.

[0043] The environmental element recognition instruction is used to prompt the fire identification large language model to output an analysis of the environmental information in the scene image, such as image metadata such as time, geographic location, and device number, and environmental state data such as weather, wind speed, and rainfall; by introducing environmental elements, the fire identification large language model can assist in determining the possibility of fire occurrence, for example, the risk of fire is lower in a rainy environment, while the risk of fire significantly increases in a dry and windy environment. The environmental element recognition instruction can be set based on actual needs, such as “please retrieve the weather conditions of the area where the scene is located on the current day, and analyze it in combination with the picture”; the scene recognition result output by the fire identification large language model is, for example, “according to the geographic location and time information, the local weather is sunny with high wind speed and no rainfall; in combination with the picture, the flame and smoke are more likely to belong to real fire, rather than natural weather phenomena”.

[0044] The fire feature recognition instruction is used to prompt the fire identification large language model to output the extraction of typical fire features in the scene image and the comparison and differentiation with common interference factors, such as fire and smoke; common interference factors such as morning mist, cooking smoke, and light reflection; so as to ensure the reliability of fire feature recognition and reduce false positives due to environmental interference. The fire feature recognition instruction can be set based on actual needs, such as “please determine whether there are typical fire features (flame, smoke, etc.) in the picture, and explain the differences between them and common environmental interference factors (such as morning mist, cooking smoke, and light, etc.).”; the scene recognition result output by the fire identification large language model is, for example, “orange-red flames and black smoke rising upwards can be seen in the image; compared with morning mist, its shape is not uniform and its color is darker; compared with cooking smoke, its scale is larger and its diffusion speed is faster; compared with light, its edge is irregular and accompanied by smoke diffusion”.

[0045] In specific implementation, it can be started from a large scale and refined to a small scale, such as inputting a scene recognition instruction first, then inputting an environmental element recognition instruction, and then inputting a fire feature recognition instruction; then a final fire recognition result is obtained by inputting a summary instruction, which can be set based on actual needs, such as “please judge whether the image contains fire according to the foregoing analysis, and give reasons”. The fire recognition result output by the fire recognition large language model is “yes, the image contains fire, and the reasons are as follows: there are large-scale orange-red flames and black smoke, combined with sunny weather and wind speed conditions, which conforms to the typical performance of forest fire”.

[0046] In this embodiment, through the design of multiple rounds of dialogue, the fire recognition large language model can gradually construct the judgment logic of fire from different angles and dimensions, thereby improving the accuracy of the fire recognition result.

[0047] Further, in the third embodiment of the fire recognition method of the present application based on the first embodiment of the present application, the method further comprises the steps of: Step S14, obtaining a model training sample, wherein the model training sample comprises a scene image sample, a recognition instruction sample and an expected result sample; Step S15, obtaining an initial fire recognition large language model; Step S16, inputting the scene image sample and the recognition instruction sample into the initial fire recognition large language model to obtain a training recognition result output by the initial fire recognition large language model; Step S17, comparing the training recognition result with the expected result sample to obtain a result difference; Step S18, updating the initial fire recognition large language model according to the result difference to obtain the fire recognition large language model.

[0048] The model training sample is used to train the initial fire recognition large language model.

[0049] The model training sample is composed of multiple groups of training data, and the training data is composed of a triple (x, I, A exp ), wherein: x represents a scene image sample; I represents a recognition instruction sample designed for the scene image sample; A expindicates the expected result sample corresponding to the identification instruction sample; specifically, it can be annotated by an expert. When setting the training data, positive and negative samples can be set. The positive sample indicates the conclusion of the fire and the supporting reasons, such as "there is a fire, there are obvious orange-red flames and smoke in the picture, which conforms to the typical characteristics of fire." The negative sample indicates the conclusion of non-fire and the supporting reasons, such as "there is no fire, the red clouds in the picture belong to the sunset scene, and there is a lack of flame and smoke characteristics." The fine-tuning of the large model adopts the supervised fine-tuning (SFT) method. In order to ensure the effect of fine-tuning training, the model training samples need to cover the fire characteristics in different scenes, such as common sunlight, industrial steam and other interference factors. The scale of the model training sample should be large enough to ensure that the model can learn rich fire discrimination knowledge and can cope with interference factors in different environments. In terms of quality, ensure that the annotation of each image is correct and consistent, avoid label errors, and ensure the balance of positive and negative samples.

[0050] The specific training method can be set based on actual needs, such as SFT (Supervised Fine-Tuning, supervised fine-tuning). The basic principle is: combine visual input and natural language instructions to form a sample pair of identification instruction-expected result, and train the large model to generate answers that meet expectations under the condition of given scene image and identification instruction. The goal of training is to guide the model to learn the multi-modal association between scene images and identification instructions, so that it can not only identify fire, smoke and other fire characteristics, but also give reasonable explanations in combination with environmental interference factors, thereby realizing the comprehensive discrimination ability of "recognition + reasoning + explanation".

[0051] The result difference indicates the difference between the output of the fire identification large language model and the expected output, therefore, updating the fire identification large language model through the result difference can make the fire identification large language model closer to the desired expected output, minimizing the difference between the training identification result predicted by the model and the expected result sample provided by the expert; the update of the fire identification large language model can be realized by setting a loss function, which can be set based on actual needs.

[0052] Further, the step S18 includes the steps of: Step S181, judging whether there is a preset keyword in the training identification result or the expected result sample; Step S182, if there is a preset keyword in the training identification result or the expected result sample, updating the initial fire identification large language model according to the result difference.

[0053] In this embodiment, the preset keyword is set, and the preset keyword is related to the keywords in the fire field, such as the preset keyword set K = {fire, flame, smoke, non-fire, interference, …}; when it is determined whether the preset keyword exists in the training recognition result or the expected result sample, it is considered that the key factor related to the fire is involved, and therefore, the content needs to be focused on and constrained; the specific loss function can be represented as:

[0054] wherein, represents the probability of the i-th word generated by the model in the t-th round; represents the i-th word of the expected answer in the t-th round; II(·) is an indicator function, which is 1 only when the preset keyword belongs to the preset keyword set K, otherwise 0.

[0055] For example, the expected result sample is “There is a fire, and there are flames and smoke in the image, which conforms to the typical characteristics of forest fires”; The training recognition result is “There are bright light points in the picture, but no smoke is detected”.

[0056] At this time, “flame” and “smoke” both belong to the set K, so the loss function will impose strong constraints on the errors of these two word positions, and no constraints will be imposed on non-keywords such as “There are” and “bright light points”. In this way, the training can guide the model to correctly learn the generation of key semantics in the field, so as to ensure the reliability of the final answer in terms of professionalism and completeness.

[0057] When the initial fire identification large language model is trained to meet the training completion condition, a trained fire identification large language model is obtained; the training completion condition can be set based on actual needs, such as learning rate, batch size, and other hyperparameters, or training optimization target; the training optimization target can be set to accurately identify the core visual features of the flame and the smoke, such as color, shape, texture, and dynamics; understand the performance of “fire” in specific scenarios, such as the different characteristics of forest fires, urban fires, and industrial fires; distinguish between real fires and common interference materials in corresponding scenes, such as red clouds, warm light, water vapor, dust, and reflections, to reduce false positives; generate discrimination reasons that conform to the professional logic of fire identification, such as “Confirm fire: the flame in the image is orange-red, the smoke is dense, and it conforms to the characteristics of forest fires”; when the initial fire identification large language model meets the set training optimization target, it is considered to meet the training completion condition.

[0058] It can be understood that after the training is completed, in the application of the fire identification large language model, the fire identification large language model may make identification errors due to complex scenes or interference factors, so the fire identification large language model can be corrected and updated based on the actual detection situation in the application of the fire identification large language model; Specifically, it can include: Error judgment and identification: if the fire identification result is different from the real scene, the artificial judgment is wrong; Error cause analysis: determine the possible source of the fire identification large language model misjudgment such as light interference, similar shape, etc., and the related data of this time, such as scene image, identification instruction, fire identification result, as misjudgment sample.

[0059] Correction process: by comparing the difference between the fire identification result and the real scene, guiding the model to update the reasoning logic; It can be compared with the training process; Incremental training: misjudgment samples will be recorded and fed back to the training data of the model, through the incremental training mechanism, these misjudgment samples will be used to retrain the model, so that it can reduce similar errors in future judgment: Multi-round correction mechanism: in multi-round dialogue, if the model makes a misjudgment in a round of answer, the system will adjust in subsequent rounds. For example, if the model misjudges the existence of fire in the second round of answer, the subsequent answer can guide the model to reanalyze the previous answer according to the expert logic, so as to correct the error and enhance the understanding of the scene.

[0060] Automatic feedback loop: in order to ensure the continuous optimization of the system, all misjudgments and correction processes will be recorded in the system to form a feedback mechanism. Misjudgment samples are generated regularly to update and fine-tune the system, ensuring that the model improves its judgment accuracy in continuous operation.

[0061] Further, in the fourth embodiment of the fire identification method of the present application based on the first embodiment of the present application, the step S21 includes the following steps before the step S21: Step S25, acquiring the scene type of the fire identification light model application; Step S26, acquiring the model training sample corresponding to the scene type; Step S27, acquiring an initial fire identification light model; Step S28, training the initial fire identification light model through the model training sample to obtain the fire identification light model.

[0062] It can be understood that under different scene types, the state of fire occurrence is different, and the specific scene type can be set based on actual needs, such as forest, city, industry and oil field.

[0063] Since the edge end is deployed in the front-end system close to the fire monitoring site, the edge end is directly corresponding to the identification of the fire and the scene type. Therefore, in the embodiment, corresponding model training samples are set for different scene types, so that the fire identification light model under a specific scene can be trained for fire identification under the scene.

[0064] The model training samples can be set based on actual scene characteristics. For example: If the scene type is forest, forest fire images are collected, and the dynamic characteristics of flames and smoke are focused on. Unmanned aerial vehicles and ground high point monitoring cameras are used for data collection, covering images under different weather conditions.

[0065] If the scene type is city, city fire data mainly comes from building fires, traffic fires, etc. City monitoring cameras and traffic monitoring equipment are used to collect samples to ensure that images under different light and environmental interference are covered.

[0066] If the scene type is industry, the characteristics of fires in industrial environments are different from those in other scenes, mainly focusing on high-temperature equipment, chemical leaks, etc. Monitoring equipment in industrial areas is used to collect data on different equipment fires, focusing on oil fires, smoke, and other specific fire characteristics.

[0067] If the scene type is oil field, oil field fires usually occur in oil wells, oil storage tanks, etc., with strong flames and thick smoke. We collect fire images in these specific environments through special monitoring equipment and unmanned aerial vehicles, and the data covers the influence of special factors such as high temperature and oil vapor.

[0068] Data collection in different scenes can be done through different terminal devices to ensure the diversity and comprehensiveness of the data. For example, unmanned aerial vehicles are suitable for large areas such as forests and oil fields, and can provide fire images at different heights and angles to capture fire information in a wide area. Ground fixed cameras are suitable for collecting city and industrial fire data. These devices provide stable and local fire images to help the model more accurately identify local fire characteristics. Mobile inspection equipment such as fire inspection robots, robotic dogs, and vehicle-mounted monitoring systems can effectively supplement the blind area of fixed monitoring and are suitable for fire data acquisition in industrial facilities, complex urban structures, and other scenes, enhancing the system's ability to perceive real-time fire changes.

[0069] To improve the reliability of the samples, after collecting the image samples, the image samples can be enhanced, such as denoising, light adjustment, etc., and the image samples are standardized to improve the data quality and ensure that the model can effectively learn the fire characteristics in different environments during training.

[0070] After collecting the image samples, structured labeling is needed to form a labeled data pair. For example, for each image sample i, a label pair L is seti :

[0071] wherein y i ∈{0,1} represents whether a fire exists, 1 for existing and 0 for non-existing. S i ∈S represents an environmental scene label to which the image sample belongs, S={forest, city, industry, oilfield}.

[0072] For each image sample, in addition to labeling whether a fire exists, interference content images such as sunlight, smoke, etc. can also be labeled to avoid model misjudgment.

[0073] Further, the step S28 comprises the steps of: Step S281, image degradation enhancement is performed on the model training sample to obtain a degradation training sample; Step S282, the initial fire identification lightweight model is trained through the degradation training sample.

[0074] In actual monitoring scenarios, images can be affected by various environmental factors, such as unmanned aerial vehicle motion blur, direct sunlight overexposure, dust and haze interference, etc. These factors can affect the sensor visual imaging quality, interfere with the model recognition accuracy, increase the model false alarm rate, and reduce the recognition efficiency. Therefore, in order to improve the robustness and accuracy of the model, image degradation enhancement processing can be performed on the model training sample; the purpose of image degradation enhancement is to simulate the interference in the real scene and enhance the adaptability of the fire identification lightweight model to these interferences. For example, let the original image be I raw , the image after degradation enhancement be I aug , the degradation transformation function be τ(∙; θ aug ), then the degradation training sample after image degradation enhancement is:

[0075] θ aug is the parameter of degradation enhancement.

[0076] The specific model of image degradation enhancement can be set based on the type of interference in the scene; for example, a motion blur model is set for motion blur interference; the image stretching blur caused by the rapid movement of the unmanned aerial vehicle or the camera is simulated. Assuming that K motion is a linear motion kernel, and the image degradation model is:

[0077] wherein I' is the image after motion blur processing; K motion is a linear motion kernel with length l:

[0078] where θ is the moving direction, v is the moving speed, t is the shutter time, d is the target distance, w is the image width, and FOV is the field of view.

[0079] A fog interference model is set for fog interference; Based on the atmospheric scattering model, the field of view visibility reduction caused by fog, dust or smoke is simulated. Based on the atmospheric scattering model, the influence of fog can be simulated by the following formula:

[0080]

[0081] where t(x) is the transmittance, β∈[0.5,1.2] is the scattering coefficient, d(x) is the pixel depth, and A∈[0.8,1.0] is the atmospheric light intensity.

[0082] A sun overexposure model is set for strong light interference; To simulate strong light interference, a radial mask is used to generate a highlight area. Overexposure processing can be done by the following formula:

[0083] where M(x) is a radial mask that gradually changes from the center to the outside, η is the overexposure intensity (such as 60~100), and 255 is the maximum value of the pixel value, ensuring that the image part area produces saturated highlights.

[0084] Through these degradation enhancement models, we can effectively simulate the influence of various environmental factors on image quality, so as to train a more robust initial fire identification lightweight model and reduce false positives and false negatives in actual application.

[0085] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited to the action sequence described, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0086] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software on a general hardware platform as required, and of course can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as a ROM / RAM, a magnetic disk, or an optical disc) and includes a plurality of instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device) to execute the methods described in the various embodiments of the present application.

[0087] The present application also provides a fire identification system for implementing the above fire identification method, the fire identification system comprising a cloud server and an edge end; the cloud server is used for: obtaining a scene image uploaded by the edge end and inputting the scene image into a fire identification large language model, wherein the scene image is uploaded when the edge end detects a fire; inputting a plurality of rounds of identification instructions into the fire identification large language model to enable the fire identification large language model to output identification sub-results corresponding to the plurality of rounds of identification instructions according to the scene image, wherein the plurality of rounds of identification instructions are directed to different objects in the scene image; inputting a summary instruction into the fire identification large language model to enable the fire identification large language model to obtain a fire identification result by synthesizing a plurality of the identification sub-results; the edge end is used for: obtaining a scene image and inputting the scene image into a fire identification light model; obtaining an edge result corresponding to the scene image output by the fire identification light model; determining whether the edge result is a fire; if the edge result is a fire, uploading the scene image to the cloud server.

[0088] The present system identifies fire through a large language model, enabling the use of the multi-modal capabilities of the large language model to improve the accuracy of fire identification, and gradually obtaining identification results through a plurality of rounds of identification instructions, thereby enabling guided fire identification large language models to fully perceive the content of the scene image, thereby enhancing the reasoning depth and discrimination accuracy of the fire identification large language model. At the same time, by identifying fire in the cloud server and the edge end respectively, the cloud and the edge can be cooperatively detected to achieve dual improvement of fire identification accuracy and interpretability.

[0089] Further, the inputting of a plurality of rounds of identification instructions into the fire identification large language model comprises: inputting a scene recognition instruction into the fire identification large language model, so that the fire identification large language model outputs a scene recognition result according to the scene image; inputting an environmental element recognition instruction into the fire identification large language model, so that the fire identification large language model outputs an environmental element recognition result according to the scene image; inputting a fire feature recognition instruction into the fire identification large language model, so that the fire identification large language model outputs a fire feature recognition result according to the scene image.

[0090] Further, the method further comprises: obtaining a model training sample, wherein the model training sample comprises a scene image sample, a recognition instruction sample, and an expected result sample; obtaining an initial fire identification large language model; inputting the scene image sample and the recognition instruction sample into the initial fire identification large language model to obtain a training recognition result output by the initial fire identification large language model; comparing the training recognition result with the expected result sample to obtain a result difference; updating the initial fire identification large language model according to the result difference to obtain the fire identification large language model.

[0091] Further, the updating the initial fire identification large language model according to the result difference comprises: determining whether a preset keyword exists in the training recognition result or the expected result sample; if the preset keyword exists in the training recognition result or the expected result sample, updating the initial fire identification large language model according to the result difference.

[0092] Further, the inputting the scene image into the fire identification light model comprises: obtaining a scene type to which the fire identification light model is applied; obtaining a model training sample corresponding to the scene type; obtaining an initial fire identification light model; training the initial fire identification light model through the model training sample to obtain the fire identification light model.

[0093] Further, the training the initial fire identification light model through the model training sample comprises: performing image degradation enhancement on the model training sample to obtain a degraded training sample; training the initial fire identification light model through the degraded training sample.

[0094] Referring to Figure 3 In the hardware structure, the electronic device can include a communication module 10, a memory 20, a processor 30, and the like. In the electronic device, the processor 30 is connected with the memory 20 and the communication module 10, respectively, the memory 20 stores a computer program, the computer program is executed by the processor 30, and the computer program implements the steps of the above method embodiment when executed.

[0095] The communication module 10 can be connected with an external communication device through a network. The communication module 10 can receive a request sent by the external communication device, and also can send a request, an instruction and information to the external communication device. The external communication device can be other electronic devices, servers or Internet of Things devices, such as a television and the like.

[0096] The memory 20 can be used to store software programs and various data. The memory 20 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application program required by a function (such as obtaining a scene image uploaded by an edge terminal and inputting the scene image into a fire identification large language model), and the like; the data storage area can include a database, and the data storage area can store data or information created according to the use of the system, and the like. In addition, the memory 20 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device.

[0097] The processor 30 is the control center of the electronic device, connects each part of the entire electronic device through various interfaces and lines, executes the software programs and / or modules stored in the memory 20 and the data stored in the memory 20, and processes the data of the electronic device, so as to perform the overall monitoring of the electronic device. The processor 30 can include one or more processing units; optionally, the processor 30 can integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, the user interface and the application program, and the modem processor mainly processes the wireless communication. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 30.

[0098] Although Figure 3 Although not shown, the above-mentioned electronic device can also include a circuit control module for connecting with a power supply to ensure the normal work of other components. Those skilled in the art can understand that Figure 3 The electronic device structure shown in the above-mentioned electronic device does not constitute a limitation on the electronic device, and can include more or fewer components than shown in the figure, or combine certain components, or different component arrangements.

[0099] The present application also provides a computer readable storage medium having stored thereon a computer program. The computer readable storage medium can be a memory 20 in an electronic device, or at least one of a ROM (Read-Only Memory) / RAM (Random Access Memory), a magnetic disc, and an optical disc, and the computer readable storage medium includes a plurality of instructions to cause an end device (which can be a television, a car, a mobile phone, a computer, a server, a terminal, or a network device, etc.) having a processor to execute the method according to the embodiments of the present application. Figure 3 The computer readable storage medium can be a memory 20 in an electronic device, or at least one of a ROM (Read-Only Memory) / RAM (Random Access Memory), a magnetic disc, and an optical disc, and the computer readable storage medium includes a plurality of instructions to cause an end device (which can be a television, a car, a mobile phone, a computer, a server, a terminal, or a network device, etc.) having a processor to execute the method according to the embodiments of the present application.

[0100] In the present application, the terms "first", "second", "third", "fourth", "fifth" are only for descriptive purposes, and cannot be understood as indicating or implying relative importance. For those skilled in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.

[0101] In the description of the present application, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present application, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, different embodiments or examples described in the present application and the features of different embodiments or examples can be combined and combined by those skilled in the art without contradiction.

[0102] Although the embodiments of the present application have been shown and described above, the scope of protection of the present application is not limited thereto. It can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application, and those skilled in the art can make changes, modifications and replacements to the above embodiments within the scope of the present application, and these changes, modifications and replacements should be covered within the scope of protection of the present application. Therefore, the scope of protection of the present application should be subject to the scope of protection of the claims.

Claims

1. A fire detection method, characterized in that, The fire detection method, applied to cloud servers, includes: Acquire scene images uploaded from the edge terminal and input the scene images into the fire recognition language model, wherein the scene images are uploaded when a fire is detected at the edge terminal; Input multi-round recognition instructions to the fire recognition big language model, so that the fire recognition big language model outputs recognition sub-results corresponding to the multi-round recognition instructions based on the scene image, wherein the multi-round recognition instructions are for different objects in the scene image; Input a summary command into the fire identification language model so that the fire identification language model integrates multiple identification sub-results to obtain a fire identification result.

2. The fire detection method as described in claim 1, characterized in that, The input of multi-round recognition instructions to the fire recognition large language model includes: Input scene recognition instructions into the fire recognition big language model, so that the fire recognition big language model outputs scene recognition results based on the scene image; An environmental element recognition instruction is input into the fire recognition big language model, so that the fire recognition big language model outputs environmental element recognition results based on the scene image; Input fire feature recognition instructions into the fire recognition big language model so that the fire recognition big language model outputs fire feature recognition results based on the scene image.

3. The fire detection method as described in claim 1, characterized in that, The method further includes: Obtain model training samples, wherein the model training samples include scene image samples, recognition instruction samples, and expected result samples; Obtain an initial large language model for fire identification; The scene image samples and the recognition instruction samples are input into the initial fire recognition large language model to obtain the training recognition result output by the initial fire recognition large language model; The training recognition results are compared with the expected result samples to obtain the result differences; The initial fire identification language model is updated based on the differences in the results to obtain the fire identification language model.

4. The fire detection method as described in claim 3, characterized in that, The step of updating the initial fire identification large language model based on the difference in the results includes: Determine whether a preset keyword exists in the training recognition result or the expected result sample; If the training recognition results or the expected result samples contain preset keywords, the initial fire recognition big language model is updated based on the result differences.

5. The fire detection method as described in claim 1, characterized in that, Applied to the edge, the fire identification method includes: Acquire scene images and input the scene images into a lightweight fire identification model; Obtain the edge results corresponding to the scene image output by the lightweight fire recognition model; Determine whether the edge result indicates a fire has occurred; If the edge result indicates a fire has occurred, the scene image will be uploaded to the cloud server.

6. The fire detection method as described in claim 5, characterized in that, The process of inputting the scene image into the lightweight fire recognition model includes: Obtain the scenario type for the application of the lightweight fire identification model; Obtain model training samples corresponding to the scene type; Obtain an initial lightweight fire identification model; The initial lightweight fire identification model is trained using the model training samples to obtain the lightweight fire identification model.

7. The fire detection method as described in claim 6, characterized in that, The step of training the initial fire identification lightweight model using the model training samples includes: The training samples of the model are subjected to image degradation enhancement to obtain degraded training samples; The initial lightweight fire identification model is trained using the degraded training samples.

8. A fire detection system, characterized in that, The fire detection system includes a cloud server and an edge terminal; the cloud server is used for: Acquire scene images uploaded from the edge terminal and input the scene images into the fire recognition language model, wherein the scene images are uploaded when a fire is detected at the edge terminal; Input multi-round recognition instructions to the fire recognition big language model, so that the fire recognition big language model outputs recognition sub-results corresponding to the multi-round recognition instructions based on the scene image, wherein the multi-round recognition instructions are for different objects in the scene image; Input a summary command into the fire identification language model so that the fire identification language model integrates multiple identification sub-results to obtain a fire identification result; The edge end is used for: Acquire scene images and input the scene images into a lightweight fire identification model; Obtain the edge results corresponding to the scene image output by the lightweight fire recognition model; Determine whether the edge result indicates a fire has occurred; If the edge result indicates a fire has occurred, the scene image will be uploaded to the cloud server.

9. An electronic device, characterized in that, The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the fire identification method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the fire identification method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Hidden disaster hazard identification method, system and device, and storage medium

    CN119807817A

  • Fire monitoring method, fire monitoring server and fire monitoring system

    CN120412183A

  • Fire monitoring method and system and computer equipment

    CN120708155A