Fire-fighting early warning method, computer readable storage medium and mobile fire-fighting robot

By integrating smoke and vision sensors into a mobile firefighting robot and combining them with data fusion technology, personalized escape route warning voices are generated, solving the monitoring blind spots and false alarms of traditional smoke alarms, and achieving efficient fire identification and intelligent escape guidance.

CN121811566APending Publication Date: 2026-04-07DIGITAL HUAXIA (SHENZHEN) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Traditional fixed smoke detectors cannot effectively intervene or provide emergency guidance in complex scenarios. They have blind spots in monitoring, a high false alarm rate, and lack proactive response capabilities, making it difficult to meet the needs of intelligent and humanized fire protection.

Method used

The mobile firefighting robot, which integrates smoke sensors, vision sensors, and voice components, calculates the fire risk coefficient by fusing environmental smoke data and image data, and generates warning voices associated with escape routes, enabling dynamic detection and intelligent guidance throughout the house.

Benefits of technology

It significantly improves the early detection rate of fire hazards, reduces the false alarm rate, provides accurate fire identification and personalized escape route guidance, and buys precious escape time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121811566A_ABST
    Figure CN121811566A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of robots, and discloses a fire-fighting early warning method, a computer readable storage medium and a mobile fire-fighting robot. The fire-fighting early warning method comprises the following steps: acquiring environment smoke data of a target place acquired based on a smoke sensor and environment image data of the target place acquired based on a visual sensor; fusing the environment smoke data and the environment image data to obtain a fire risk coefficient; and if the fire risk coefficient exceeds a risk threshold, generating an early warning voice associated with the escape path based on the environment smoke data and the environment image data, and broadcasting the early warning voice through a voice assembly. By adopting the method, the fire risk can be found in time, so that the personnel can be dynamically guided to move towards the emergency exit, and precious gold escape time is gained for the initial stage of the fire.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of robots, and in particular to a fire-fighting early warning method, a computer readable storage medium and a mobile fire-fighting robot. BACKGROUND

[0002] At present, the mainstream fire-fighting alarm device mainly adopts a fixedly installed photoelectric or ion type smoke alarm, and its working principle is to trigger an alarm by detecting the change in smoke particle concentration in the air. This type of device is widely used in places such as residences, office buildings and nursing homes, and has become a basic fire-fighting and security configuration.

[0003] However, the traditional fixed smoke alarm has obvious limitations. First, since its installation position is fixed, the monitoring range is limited, and it is difficult to cover the corners of the room, partitioned areas or multi-story spaces, which can easily form a "monitoring dead angle", resulting in a local fire that cannot be discovered in time. Second, kitchen fumes, incense, water vapor and other non-fire aerosols often cause false alarms, affecting user experience, especially in sensitive environments such as homes and nursing homes. Third, the traditional alarm only has passive sensing function and lacks active response capability, which cannot effectively intervene or provide emergency guidance in the initial stage of fire, and is difficult to meet the intelligent and humanized fire-fighting needs in complex scenarios. SUMMARY

[0004] Therefore, the embodiments of the present application provide a fire-fighting early warning method, a computer readable storage medium and a mobile fire-fighting robot, which can effectively solve the technical problem that the traditional fixed smoke alarm cannot effectively intervene or provide emergency guidance in complex scenarios.

[0005] In a first aspect, the embodiments of the present application provide a fire-fighting early warning method applied to a mobile fire-fighting robot, wherein the mobile fire-fighting robot is integrated with a smoke sensor, a visual sensor and a voice component, and the method comprises: acquiring environmental smoke data of a target site based on the smoke sensor and environmental image data of the target site based on the visual sensor; fusing the environmental smoke data and the environmental image data to obtain a fire risk coefficient; if the fire risk coefficient exceeds a risk threshold, generating an early warning voice associated with an escape path based on the environmental smoke data and the environmental image data, and playing the early warning voice through the voice component.

[0006] In a second aspect, the embodiments of the present application provide a mobile fire-fighting robot, comprising: a smoke sensor configured to acquire environmental smoke data of a target site; a visual sensor configured to acquire environmental image data of the target site; a control component configured to acquire environmental smoke data of a target site based on the smoke sensor and environmental image data of the target site based on the visual sensor, fuse the environmental smoke data and the environmental image data to obtain a fire risk coefficient, and generate an escape route-associated early warning voice based on the environmental smoke data and the environmental image data if the fire risk coefficient exceeds a risk threshold value; a voice component configured to broadcast the early warning voice.

[0007] In a third aspect, an embodiment of the present application provides a computer readable storage medium having a computer program stored thereon, the computer program being executed by a controller to implement the following steps: acquire environmental smoke data of a target site based on the smoke sensor and environmental image data of the target site based on the visual sensor; fuse the environmental smoke data and the environmental image data to obtain a fire risk coefficient; generate an escape route-associated early warning voice based on the environmental smoke data and the environmental image data if the fire risk coefficient exceeds a risk threshold value, and broadcast the early warning voice through the voice component.

[0008] Embodiments of the present application have the following beneficial effects: First, by mounting the smoke sensor on a fire-fighting robot with autonomous movement capability, combined with its path planning capability, the target site (such as a home, a nursing home, etc.) can be dynamically detected in the preset period or emergency mode, effectively breaking through the spatial limitations of fixed point monitoring, significantly improving the early detection rate of fire hazards, and greatly enhancing the coverage rate of indoor environmental smoke detection.

[0009] Second, the environmental smoke data and the environmental image data are fused to jointly analyze whether there is an open fire, a high-temperature light-emitting body, or a smoke diffusion pattern, thereby distinguishing between real fires and daily interference sources. This multi-sensor cooperative judgment mechanism improves the accuracy of fire identification, making the false alarm rate significantly lower than that of traditional single sensor solutions.

[0010] Finally, after determining that the fire risk coefficient exceeds the threshold value, an early warning voice related to the escape route can be intelligently generated based on the current environmental data, and the voice component of the robot can broadcast it in real time. In addition, the robot can also dynamically guide personnel to move towards the safety exit based on its positioning and mapping capabilities, thereby gaining valuable golden escape time in the early stage of a fire. BRIEF DESCRIPTION OF DRAWINGS

[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some of the embodiments of the present application, and therefore should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can also be obtained without creative labor.

[0012] Figure 1 A hardware structure diagram of the mobile fire-fighting robot of the embodiments of the present application is shown; Figure 2 A software structure diagram of the mobile fire-fighting robot of the embodiments of the present application is shown; Figure 3 A flow chart of the fire-fighting early warning method of the embodiments of the present application is shown.

[0013] Main element symbol explanation: 1-biocular depth camera; 2-laser radar; 3-photoelectric smoke sensor; 4-AC / DC converter; 5-alarm buzzer; 6-lithium battery. DETAILED DESCRIPTION

[0014] The technical solutions of the embodiments of the present application will be described clearly and completely in the embodiments of the present application in combination with the drawings. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments.

[0015] The components of the embodiments of the present application generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of the present application.

[0016] In the following, the terms "include", "have", and their synonymous words used in various embodiments of the present application are only intended to represent specific features, numbers, steps, operations, elements, components, or combinations of the foregoing, and should not be understood as excluding the existence or adding the possibility of one or more other features, numbers, steps, operations, elements, components, or combinations of the foregoing. In addition, the terms "first", "second", "third", etc. are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.

[0017] Unless specifically defined, all other technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which various embodiments of the present application belong. The terminology used herein (e.g., the terminology used in the description portion and the claims) is for describing particular embodiments of the application and is not intended to be limiting in any way. Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which various embodiments of the present application belong. The terminology used herein, such as that found in the dictionary, will be interpreted as having a meaning that is the same as commonly understood in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless clearly defined in various embodiments of the present application.

[0018] Some embodiments of the present application are described in detail below with reference to the accompanying drawings. The following embodiments and features in the embodiments can be combined with each other without conflict.

[0019] The mobile fire-fighting robot is described below in connection with some specific embodiments.

[0020] Figure 1 A hardware structure diagram of the mobile fire-fighting robot according to an embodiment of the present application is shown. Exemplarily, the mobile fire-fighting robot at least includes a smoke sensor, a visual sensor, a control component, and a voice component.

[0021] The smoke sensor refers to a gas sensing device for detecting the concentration of suspended particulate matter in the air, specifically a photoelectric smoke sensor. The smoke sensor has a built-in light source and a photosensitive element to form a scattered light detection structure. When smoke enters the sampling cavity, the smoke particles scatter the light, and the photosensitive element receives the scattered light signal and converts it into an electrical signal output. The output signal includes the real-time smoke concentration value and its time rate of change (i.e., the concentration rise rate), which are used to reflect the different stage characteristics of the initial smoldering or rapid fire.

[0022] The visual sensor refers to a composite imaging device that can capture visible light images and depth information of the target site, including a binocular depth camera and a laser radar.

[0023] The binocular depth camera forms a stereo vision system with two spatially distributed cameras and combines infrared structured light or active light compensation modules to realize RGB image acquisition and three-dimensional point cloud reconstruction of the environment. It can synchronously acquire high-resolution color images and depth maps, and support accurate identification of flame areas, smoke diffusion patterns, and spatial obstacle distribution in low-light and complex texture environments.

[0024] The laser radar is used to provide spatial profile information to assist in completing room structure modeling and positioning navigation functions.

[0025] The control component refers to the core operation and decision unit integrated in the main control system of the humanoid intelligent robot. In the present application, it is used to execute the fire-fighting early warning method.

[0026] The voice component refers to a human-computer interaction output device for issuing voice warning information to on-site personnel, including an alarm buzzer and a speaker device. The alarm buzzer is used to issue a high-intensity sound and light warning signal (such as flashing light + high-frequency buzzing), to realize local emergency reminder; the speaker system is used to play the escape path associated warning voice generated by the control component, for example: “Please immediately evacuate from the living room to the gate, the right side passage is safe”.

[0027] Specifically, the smoke sensor is responsible for collecting environmental smoke data of the target site; the visual sensor is responsible for collecting environmental image data of the target site; the control component is responsible for obtaining the environmental smoke data of the target site based on the smoke sensor and the environmental image data of the target site based on the visual sensor; fusing the environmental smoke data and the environmental image data to obtain a fire risk coefficient; if the fire risk coefficient exceeds a risk threshold, generating an escape path associated warning voice based on the environmental smoke data and the environmental image data; the voice component is responsible for playing the warning voice. Optionally, the smoke sensor is arranged in the head and neck region of the fire warning robot through an adjustable support.

[0028] Figure 2 A software structure diagram of the mobile fire robot of the embodiment of the application is shown, and the application realizes unified communication scheduling through a three-layer architecture through a CAN bus, forms a technical closed loop of “active detection - intelligent decision - linkage response”, and specifically includes: The perception layer can be understood as the “sensory system” of the application, which is responsible for real-time collection of physical environment information of the target site, and provides stable working conditions for the front-end sensor, mainly including a sensor module and a DC-DC converter (a post-stage voltage stabilizing unit of an AC / DC converter).

[0029] Among them, the sensor module includes a photoelectric smoke sensor and a binocular depth camera, realizing dynamic smoke and image data collection under “active inspection”, breaking through the spatial limitation of traditional fixed alarms.

[0030] The DC-DC converter is connected to the output end of the internal lithium battery of the robot, and reduces the high-voltage direct-current power to low-voltage direct-current power suitable for the work of various sensors, such as photoelectric smoke sensors, binocular depth cameras, etc. Optionally, the DC-DC converter supports a dynamic voltage regulation mode, which can adjust the output power according to the task state (standby / inspection / alarm), to improve the energy efficiency ratio.

[0031] The decision layer can be understood as the “brain center” of the application, which undertakes key intelligent computing tasks such as data fusion, fire discrimination, and escape path generation, and is mainly composed of an ADC sampling circuit and a main control MCU (Microcontroller Unit), and runs in the robot main control system.

[0032] The ADC sampling circuit is connected to an analog signal output port of a smoke sensor, converts continuous smoke concentration electric signals into discrete digital quantities, and provides the host MCU for reading and processing, realizes lossless migration of analog signals to the digital domain, and guarantees the accuracy and timing consistency of environmental smoke data.

[0033] The host MCU is a high-performance embedded processor, which is the core operation platform of the decision layer of the application, and is used to execute the fire warning method in the application.

[0034] The execution layer can be understood as the effector of the application, which is responsible for converting the decision results into perceptible external actions or information output, and is mainly embodied as a voice component driving mechanism realized through an SPI interface.

[0035] The SPI (Serial Peripheral Interface) interface is a high-speed synchronous serial communication channel between the host MCU and the voice broadcast module, which is used to transmit two types of data, namely, warning text information (natural language prompts generated by the host MCU according to the escape path) and control instructions (triggering the text-to-speech processing of the voice synthesis chip), efficiently and reliably converting abstract decisions into human-understandable emergency guidance, and completing the last meter of safe transmission.

[0036] Optionally, the execution layer can further integrate a CAN bus interface for pushing remote alarm information to a home gateway and a cloud platform, realizing a three-level response mechanism of “local voice guidance + remote notification of family members / fire center”.

[0037] The fire warning method will be described below in combination with some specific embodiments.

[0038] Figure 3 A flowchart of the fire warning method of the embodiment of the application is shown. Exemplarily, the fire warning method includes the following steps: In step S302, environmental smoke data of a target site collected based on a smoke sensor and environmental image data of the target site collected based on a vision sensor are acquired.

[0039] The target site refers to a spatial range in which a mobile fire robot performs fire inspection and warning tasks, and specifically includes indoor closed or semi-closed building areas such as family residences, nursing home rooms, offices, corridors, and stairwells.

[0040] The environmental smoke data refers to a set of sensing signals collected continuously by a photoelectric smoke sensor integrated on the mobile fire robot, reflecting the concentration of suspended particulate matter in the air and its dynamic changes.

[0041] The environmental image data refers to the two-dimensional color image, infrared image or three-dimensional point cloud fusion data of the target site collected by the visual sensor (including a binocular depth camera and / or a laser radar auxiliary imaging system) carried by the mobile fire robot.

[0042] Specifically, the acquisition of the environmental smoke data mainly includes: The photoelectric smoke sensor is installed in the head and neck area of the robot, and the pitch angle is adjusted by ±30° through an adjustable support, so as to ensure that the surrounding air samples can be effectively sucked in under different postures (standing, bending, turning).

[0043] The sensor is internally provided with a labyrinth sampling cavity, and an infrared light-emitting diode is arranged as a light source. When there are smoke particles in the air, the light is scattered and captured by the lateral light-sensitive receiver. The receiver converts the received scattered light intensity into an analog voltage signal, which is converted into a digital quantity through an ADC sampling circuit.

[0044] The control component periodically reads the value and calculates two key parameters: the real-time smoke concentration value, which is expressed in percentage reading, reflecting the density of suspended particulate matter in the current air; and the smoke concentration change rate, which is the concentration increment per unit time, used to identify the difference between slow smoke release of smoldering fire and rapid smoke rising of sudden fire.

[0045] During the data acquisition process, the robot is in an autonomous inspection mode, moves at a low speed (0.3~0.5 m / s) along a preset path (such as living room→bedroom→kitchen→corridor), realizes continuous airflow monitoring in the whole house, and avoids the detection blind area caused by the position limitation of the traditional fixed alarm.

[0046] The acquisition process of the environmental image data mainly includes: The binocular depth camera is integrated in the central position of the face of the robot, has RGB color imaging and infrared active light supplement functions, and supports normal operation in a dark environment with illumination lower than 10 lux; the camera continuously collects visible light images of the target site at a rate of 15 frames per second (fps), with a resolution of 640×480, and synchronously outputs a depth map for three-dimensional space reconstruction.

[0047] The laser radar synchronously scans the distribution of surrounding obstacles, constructs a local point cloud map, and assists in positioning and navigation; the control component receives the original image frame in real time, and performs preprocessing operations: image denoising (using a bilateral filtering algorithm); white balance correction; normalizing the pixel value to the [0, 1] interval, and converting it into a floating-point tensor with a shape of (640, 480, 3); the preprocessed image data is cached in the local memory for subsequent calling by the flame and smoke identification module.

[0048] In an example, upon detecting the abnormal smoke concentration, the robot automatically adjusts the orientation so that the binocular camera is directed towards the inside of the kitchen. The image data shows that there is a grayish-white semi-transparent diffusion area above the cooking hob, which has different morphological characteristics from typical cooking fume, and no obvious flame light source is found. The image sequence is sent to the decision layer together with the smoke sensor data for multi-modal fusion analysis.

[0049] Optionally, all sensor data are time-stamped (with a precision of milliseconds), and are uniformly scheduled by the main control MCU; the smoke data and image frames are aligned in time windows (for example, within ±50 ms is considered as the same observation time), forming a spatiotemporally correlated data pair; if there is a communication delay or asynchronous collection, an interpolation method is used to complete the missing values, ensuring input consistency.

[0050] Through the above embodiments, the smoke sensor and the visual sensor are deeply integrated into the mobile robot platform, realizing the transition from "static point detection" to "dynamic global perception". Not only does it improve the ability to capture early fire signals, but it also provides a high-quality and high-timeliness raw data basis for subsequent multi-source information fusion.

[0051] Step S304, fusing the environmental smoke data and the environmental image data to obtain a fire risk coefficient.

[0052] The fire risk coefficient refers to a comprehensive evaluation index for quantitatively judging whether a target place is threatened by a real fire. It is a dimensionless value, and its value range is usually between [0, 1] or [0, 100], with a higher value indicating a higher possibility of fire.

[0053] In one of the embodiments, the environmental image data is subjected to flame region identification to obtain a flame position and a first confidence degree of the flame position; the environmental image data is subjected to smoke region identification to obtain a smoke position and a second confidence degree of the smoke position; the first confidence degree, the second confidence degree, the smoke concentration value, and the change rate of the smoke concentration value are fused to obtain the fire risk coefficient.

[0054] The first confidence degree refers to the confidence degree corresponding to the identification result of the target detection model that there is a flame in a certain image region in the environmental image data collected by the visual sensor, which is expressed as a probability value between 0 and 1.

[0055] The second confidence degree refers to the confidence degree corresponding to the identification result of the target detection model that there is smoke in a certain image region in the environmental image data collected by the visual sensor, which is also expressed as a probability value between 0 and 1.

[0056] Taking the calculation of the second confidence as an example, the environmental image data is converted into a normalized tensor and input into a target detection model, the target detection model comprising a plurality of different resolution convolutional layers; the environmental image data is converted into a plurality of feature maps of different scales through each convolutional layer; for each scale of the feature map, target prediction is performed on the feature map to generate a plurality of candidate bounding boxes, the candidate bounding boxes containing position coordinates, class probability and confidence score; for each candidate bounding box, the maximum class probability and the confidence score corresponding to the candidate bounding box are multiplied to obtain an initial confidence score of the candidate bounding box; based on the confidence threshold and the initial confidence score, a target bounding box is screened out from the candidate bounding boxes; the position coordinates and the initial confidence score in the target bounding box are taken as the smoke position and the second confidence of the smoke position, respectively.

[0057] The target detection model refers to an image analysis algorithm module based on a deep neural network, which is used to automatically identify and locate the position and category information of a specific target object (such as a flame or a smoke area) from environmental image data collected by a mobile fire-fighting robot.

[0058] Optionally, the target detection model adopts a convolutional neural network (CNN) architecture, preferably a single-stage detector such as YOLOv5, which contains a plurality of convolutional layers of different resolutions and supports multi-scale feature extraction; the model input is a normalized image tensor (for example, the size is 640x640x3, and the pixel value is normalized to the interval [0, 1]); the output is a group of candidate bounding boxes with semantic labels, each box containing position coordinates, class probability and confidence score; the model is specially trained using an indoor scene image dataset containing a large number of "flame" and "smoke" category labels, so that it has the ability to accurately identify under complex lighting, occlusion and similar interference (such as steam and dust); deployed on the main control MCU or edge computing unit of the mobile fire-fighting robot, supporting real-time inference.

[0059] The candidate bounding box is a set of rectangular boxes of potential target positions generated by the target detection model on different scale feature maps during the forward propagation process, which is used to represent the areas in the image where flames or smoke may exist.

[0060] The position coordinates refer to the two-dimensional spatial positioning information of the candidate bounding box or the target bounding box in the environmental image data, which is used to indicate the specific image area where the flame or the smoke is located.

[0061] The class probability refers to the probability distribution of the image content in a certain candidate bounding box belonging to a preset category (such as "flame", "smoke", "non-fire interference", "unknown", etc.) output by the target detection model.

[0062] The confidence score refers to the overall credibility of a certain candidate bounding box in measuring the existence of a real target (rather than noise or artifacts) in the region enclosed by the bounding box, regardless of the specific category.

[0063] The target bounding box refers to the optimal bounding box set representing the most credible flame or smoke region in the image, which is retained after screening among all candidate bounding boxes.

[0064] Specifically, after obtaining the original image frame from the camera, the first preprocessing is performed: linearly mapping the pixel value from the [0, 255] interval to the [0.0, 1.0] floating-point range; then converting the image into a normalized tensor with a shape of (1, 3, 640, 640) as the input of the target detection model.

[0065] The target detection model adopts a lightweight YOLOv5s architecture and is deployed on the edge AI inference engine of the robot host MCU. The model includes multiple convolution layers of different resolutions, forming a multi-scale feature extraction network, which outputs three scales of feature maps: 80x80 (for detecting small targets), 40x40 (for medium targets), and 20x20 (for large targets).

[0066] On each scale of feature map, the target detection model generates a number of candidate bounding boxes through the detection head, each box containing the following four pieces of information: position coordinates; category probability (a probability distribution vector indicating whether the region belongs to "smoke", "flame", "interference", or "background"); confidence score (reflecting the overall credibility of the existence of a real target in the bounding box, independent of the classification result); and all candidate boxes are preliminarily sorted by confidence.

[0067] The "maximum class probability" corresponding to each candidate bounding box is multiplied by the "confidence score" to obtain a comprehensive credibility indicator, i.e., the initial confidence score. Set a confidence threshold, and only keep the candidate boxes higher than this confidence threshold; perform non-maximum suppression on the retained bounding boxes to remove the overlapping degree. The final remaining bounding boxes are the target bounding boxes, representing the most credible smoke region in the image.

[0068] According to the screened target bounding box, the final output is completed, including: smoke position, which is the position coordinates of the target bounding box, mapped back to the original image coordinate system, for subsequent spatial analysis; and the second confidence, which is the initial confidence score of the target bounding box as the credibility indicator of the existence of smoke in this image recognition, transmitted to the fire risk fusion module to fuse the first confidence, the second confidence, the smoke concentration value, and the change rate of the smoke concentration value to obtain the fire risk coefficient.

[0069] Optionally, the fusing is performed through a weighted fusing formula (including a weight coefficient of the first confidence degree, a weight coefficient of the second confidence degree, a weight coefficient of the smoke concentration value, and a weight coefficient of the change rate of the smoke concentration value). It can be understood that the calculation of the first confidence degree is consistent with the calculation logic of the second confidence degree, which will not be described here.

[0070] Through the above embodiments, an end-to-end visual recognition link is constructed, and accurate positioning and credible quantification of visible smoke are realized. Compared with the traditional method of relying only on smoke sensors to trigger alarms, the introduction of the "second confidence degree" as a visual auxiliary criterion can effectively eliminate false alarm events caused by cooking fumes, humidifier steam, and the like.

[0071] In step S306, if the fire risk coefficient exceeds the risk threshold, an escape path-associated early warning voice is generated based on the environmental smoke data and the environmental image data, and the early warning voice is broadcast through a voice component.

[0072] The risk threshold refers to a preset critical value for determining whether to trigger a fire warning response, and its value range is usually in the interval [0, 1], corresponding to the determination standard of the fire risk coefficient. When the calculated fire risk coefficient is greater than or equal to the threshold, the system determines that there is a significant fire risk, and starts the early warning process; otherwise, it maintains the normal inspection state. The risk threshold can be dynamically adjusted according to the application scenario, for example, a lower threshold is set in a high-risk area (such as a warehouse or a power distribution room) to improve sensitivity, and the threshold is appropriately increased in an environment prone to false alarms (such as a kitchen) to enhance anti-interference capability.

[0073] The escape path refers to an optimal or suboptimal evacuation guide route from the current position to the safe exit, which is planned based on the current position of the robot, the spatial structure of the target site, and the distribution of the dangerous area. The escape path avoids the identified flame area, smoke dense area, and physical obstacles, and preferentially selects the path with the shortest distance, the least travel time, or the highest safety. The path is represented by a series of consecutive spatial coordinate points or grid cell sequences, serving as the basic data for generating voice prompts.

[0074] The early warning voice refers to the voice information associated with the escape path, which is broadcast by the mobile fire-fighting robot through its voice component after confirming that the fire risk exceeds the risk threshold. The early warning voice content includes but is not limited to voice instructions with directionality, location, and urgency, such as "please walk straight along the left corridor for ten meters and turn right to the safe exit" and "there is a fire ahead, please do not approach". The voice is generated by converting the text through a voice synthesis module, supporting multi-language and multi-pitch output.

[0075] In one of the embodiments, based on the environmental image data, room structure information of the target site is identified, the room structure information including a current position, an entrance and exit position, and an obstacle distribution; the target site is divided into a plurality of grid units; based on the position coordinates and the initial confidence score in the target bounding box, a smoke concentration value of each grid unit is determined; based on the current position, the entrance and exit position, the obstacle distribution, and the smoke concentration value of each grid unit, an escape path is generated; based on the room structure information, a plurality of reminder keywords on the escape path are determined, the reminder keywords including a room element, an escape direction, and an escape distance; the plurality of reminder keywords are fused to obtain a warning text in an order of the room elements in the escape path, and the warning text is converted into a warning voice.

[0076] The room structure information refers to a set of internal space layout features of the target site extracted by semantic segmentation and space modeling on the environmental image data, including but not limited to: wall contour, room type (such as bedroom, corridor, stairwell), door position, passable area, fixed facility distribution, etc. The room structure information is used to construct a local map to support subsequent positioning, path planning, and obstacle avoidance decision-making.

[0077] The entrance and exit position refers to the specific spatial coordinates or relative orientation of the passage opening that allows personnel to enter and exit the target site, including the main door, side door, emergency exit, fire door, etc. In the present application, the entrance and exit position is one of the candidate points for the end of the escape path planning, and the closest safe exit that is not blocked by smoke or fire is preferentially selected as the target end. This information can be obtained by pre-importing architectural drawings or dynamically updated by real-time identification of doorplate signs, green emergency light signs, etc. through visual sensors.

[0078] The obstacle distribution refers to the spatial distribution of static or dynamic objects that affect personnel passage or robot movement within the target site, including furniture, equipment, stacked items, collapsed structures, and temporarily stopped crowds, etc. The obstacle distribution is perceived and labeled in three dimensions by visual sensors combined with depth cameras or laser radars, and is mapped into a local map for path planning to avoid impassable areas. For dynamic obstacles, continuous tracking and dynamic adjustment of the escape path are supported.

[0079] The grid unit refers to the division of the spatial plane of the target site into a plurality of regular small area units (such as squares or rectangles), each unit having a unique coordinate index for digital expression of the space state. Each grid unit in the present application can be associated with a risk level, which is evaluated based on factors such as the flame position, smoke concentration prediction value, and presence or absence of obstacles covered by the grid unit.

[0080] The reminding keyword refers to the core semantic unit used to generate the pre-warning voice, the key guiding information segment extracted from the path planning result, including room elements, escape direction, escape distance, etc. For example, "corridor", "left turn", "five meters ahead" are typical reminding keywords. These keywords are combined in logical order according to the escape path to form coherent natural language sentences, improving the intelligibility and practicality of voice prompts.

[0081] The room element refers to the spatial component with clear function or identifying feature appearing on the escape path, used to help personnel locate the environment they are in and identify the travel node. Typical room elements include: corridor, stairwell, elevator hall, safety door, fire extinguisher box, corner, etc. Mentioning room elements in voice prompts can enhance spatial cognition, for example: "After passing the fire extinguisher box at the end of the corridor, please turn right".

[0082] The escape direction refers to the direction of travel required to reach the next path node from the current location, described in relative terms such as "straight", "left turn", "right turn", "diagonal forward", etc. In complex spaces, absolute directions (such as "north direction") can also be used for supplementary explanation. The escape direction is calculated from the geometric relationship between path points and calibrated with the robot's attitude angle to ensure accurate indication.

[0083] The escape distance refers to the straight-line or path distance between the current position and the next key node (such as a turning point, doorway) or the final safe exit, usually expressed in meters (m). It is reflected in voice prompts as "ten meters ahead", "about 20 steps to the exit", etc. This distance information is calculated by the path planning module based on grid cell size or actual coordinate difference, supporting dynamic refresh to reflect the latest path changes.

[0084] The pre-warning text refers to the intermediate text expression of natural language form converted from the escape path information, generated by the fusion of multiple reminding keywords arranged in order. For example: "Please walk straight along the corridor for 15 meters, turn left into the stairwell, and evacuate downward". The pre-warning text needs to meet the requirements of grammatical correctness, semantic clarity, and emphasis, and then input into the voice synthesis engine to convert into playable pre-warning voice signals.

[0085] Specifically, the room structure information is identified, including: after the mobile fire robot enters the area to be monitored and completes the preliminary environmental scanning, the control component calls its integrated visual sensor (such as an RGB camera or a depth camera) to continuously collect environmental image data. Based on these image data, the system jointly processes through a semantic segmentation network algorithm to build a local spatial cognition model, thereby extracting key room structure information. The room structure information includes: current position, indicating the real-time coordinate position of the robot itself in the target site; entrance and exit position, identifying visual clues such as doorways, emergency exit signs, and green safety lights, locating all passable exit positions; obstacle distribution, detecting fixed facilities (such as tables, chairs, and cabinets), wall corners, and dynamic obstacles (such as pedestrians and mobile devices), labeling their geometric contours and relative positions in space. The above information is integrated into a two-dimensional topological graph or grid map with semantic labels, serving as the basic input for subsequent path planning.

[0086] The grid cells are divided and the smoke risk is mapped, including: in order to achieve fine risk assessment and path search, the system divides the planar space of the target site into multiple regular grid cells, each of which represents a square region with a side length of 0.5 meters (which can be adjusted to 0.25-1 meters according to the accuracy requirement). The entire monitoring area is thus converted into an MxN matrix form map. Further, using the identified target bounding box (i.e. the filtered smoke / flame area detection result), its spatial position is projected onto the grid map. For each grid cell containing the center point or covering area of the target bounding box, the system assigns its corresponding equivalent smoke concentration value according to the following strategy. Through this process, each grid cell is assigned a value reflecting the local fire risk level, forming a "dynamic risk heat map".

[0087] After obtaining the map containing the current position, entrance and exit position, obstacle distribution, and smoke concentration value of each grid cell, the system starts the path planning module and performs optimal path search. Specifically, the grid where the user's current position is located is taken as the starting node, and the grids corresponding to all unblocked entrances and exits are taken as the terminal node set. Through the search cost function, an escape path that considers the shortest distance, lowest exposure risk, and passable feasibility is output, represented as a series of ordered grid coordinate sequences.

[0088] Optionally, the search cost function is f(n) = g(n) + h(n) + K x r(n). Wherein, g(n) represents the actual path cost that has occurred from the starting position to the current position, usually measured in the number of grid cells passed or the cumulative geometric distance. h(n) is a heuristic function used to estimate the potential shortest distance from the current position to the nearest available exit, which can use Euclidean distance or Manhattan distance to ensure that the search direction is directional. r(n) is the sum of the equivalent smoke concentration values corresponding to all grid cells traversed from the starting position to the nearest available exit via the current path, reflecting the overall fire exposure level of the path. K is an adjustable weight coefficient used to balance the priority between path length and safety risk, which can be dynamically configured according to application scenarios (such as nursing homes focusing on safety). f(n) is the total evaluation cost of the escape path.

[0089] Understandably, through this search cost function, high smoke aggregation areas can be actively avoided during path planning, and passages away from flames and smoke are preferred, thereby generating an optimal escape route that balances travel efficiency and life safety. This method is particularly suitable for complex building structures with multiple exits, obstacle obstructions, and non-uniform smoke diffusion, significantly improving the quality and timeliness of emergency response.

[0090] Extracting the reminder keywords and generating the warning text includes: first, based on the room structure information and path trajectory analysis, determining several key nodes on the path, and extracting reminder keywords from them. Then, logically sorting and grammatically connecting these keywords in the order of path advancement to generate a coherent warning text. Optionally, it can be achieved by matching a pre-set template, or it can be more flexible in expression by combining a natural language generation model.

[0091] Converting and broadcasting the warning voice includes: inputting the generated warning text into the onboard voice synthesis engine to convert it into clear and distinguishable audio signals. The engine supports multiple voice styles (male / female), speech speed adjustment, and multi-language switching (such as Chinese, English, Cantonese, etc.), to meet the needs of different user groups. After conversion is complete, the control component drives the voice component (speaker) of the robot's head to play the warning voice externally, with the volume automatically gain-adjusted according to the ambient environmental noise to ensure that information can still be effectively conveyed in noisy environments.

[0092] Through deep fusion of visual perception, sensor data and spatial reasoning ability, the transition from "passive alarm" to "active guidance" is realized. Compared with the traditional way of only issuing a buzzing alarm, the present application has the following significant advantages: accuracy: based on the actual space structure and real-time fire distribution, personalized paths are generated to avoid blind guidance; intelligence: dynamically update the risk map to support path re-planning to cope with sudden changes (such as new fire points); humanization: use natural language broadcast to reduce cognitive load, especially for vulnerable groups such as the elderly and children; robustness: fusion of multi-source information reduces misjudgment and improves the stability and reliability of the system in complex environments.

[0093] It can be understood that the device of the embodiment corresponds to the fire warning method of the above-mentioned embodiment, and the optional items in the above-mentioned embodiment are also applicable to the present embodiment, so the description is not repeated here.

[0094] The present application also provides a terminal device, which exemplarily comprises a processor and a memory, wherein the memory stores a computer program, and the processor executes the computer program to make the terminal device execute the functions of each module in the above-mentioned fire warning method or the above-mentioned fire warning device.

[0095] The processor can be an integrated circuit chip with a signal processing capability. The processor can be a general-purpose processor, including a central processing unit (CPU), a graphics processing unit (GPU), and a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or at least one of them. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc., which can implement or execute the disclosed methods, steps, and logic block diagrams in the embodiments of the present application.

[0096] The memory can be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), and the like. Among them, the memory is used to store a computer program, and the processor can execute the computer program correspondingly after receiving an execution instruction.

[0097] The application further provides a computer readable storage medium for storing the computer program used in the terminal device. For example, the computer readable storage medium can include, but is not limited to, a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0098] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can also be implemented by other manners. The apparatus embodiments described above are only schematic, for example, the flow charts and structural diagrams in the drawings show the possible implementation architectures, functions and operations of the apparatus, method and computer program product according to the embodiments of the present application. In this regard, each block in the flow chart or structural diagram can represent a module, a program segment or a part of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that, in alternative implementation manners, the functions noted in the blocks can also occur in different orders from those noted in the drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and sometimes they can be executed in reverse order, depending on the functions involved. It should also be noted that each block in the structural diagram and / or flow chart, and the combination of blocks in the structural diagram and / or flow chart, can be implemented by a dedicated hardware-based system for executing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0099] In addition, each functional module or unit in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0100] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a smart phone, a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the various embodiments of the method of the present application.

[0101] The above merely describes specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered within the protection scope of the present application.

Claims

1. A fire early warning method, characterized in that, Applied to a mobile firefighting robot, the mobile firefighting robot integrating a smoke sensor, a vision sensor, and a voice component, the method includes: Acquire environmental smoke data of the target location based on the smoke sensor, and environmental image data of the target location based on the vision sensor; By fusing the environmental smoke data and the environmental image data, a fire risk coefficient is obtained; If the fire risk coefficient exceeds the risk threshold, an early warning voice message associated with the escape route is generated based on the environmental smoke data and the environmental image data, and the early warning voice message is broadcast through the voice component.

2. The method according to claim 1, characterized in that, The environmental smoke data includes the smoke concentration value output by the smoke sensor and the rate of change of the smoke concentration value; The process of fusing the environmental smoke data and the environmental image data to obtain the fire risk coefficient includes: The environmental image data is used to identify the flame region, thereby obtaining the flame location and a first confidence level of the flame location; The environmental image data is used to identify smoke regions to obtain the smoke location and a second confidence level of the smoke location; The fire risk coefficient is obtained by fusing the first confidence level, the second confidence level, the smoke concentration value, and the rate of change of the smoke concentration value.

3. The method according to claim 2, characterized in that, The step of identifying smoke regions from the environmental image data to obtain the smoke location and a second confidence level for the smoke location includes: The environmental image data is converted into a normalized tensor and input into the target detection model, which includes multiple convolutional layers with different resolutions. The environmental image data is converted into multiple feature maps of different scales through each of the convolutional layers; For each scale of feature map, target prediction is performed on the target feature map to generate multiple candidate bounding boxes, which include location coordinates, class probability and confidence score; Based on the location coordinates, category probability, and confidence score in the candidate bounding box, the smoke location and a second confidence score for the smoke location are determined.

4. The method according to claim 3, characterized in that, Based on the location coordinates, class probability, and confidence score in the candidate bounding box, the smoke location is determined, along with a second confidence score for the smoke location, including: For each candidate bounding box, the maximum class probability and confidence score corresponding to the candidate bounding box are multiplied to obtain the initial confidence score of the candidate bounding box. Based on the confidence threshold and the initial confidence score, the target bounding box is selected from the candidate bounding boxes; The position coordinates and initial confidence score within the target bounding box are used as the smoke location and the second confidence score of the smoke location, respectively.

5. The method according to claim 4, characterized in that, The generation of an escape route-related warning voice based on the environmental smoke data and the environmental image data includes: Based on the environmental image data, the room structure information of the target location is identified, including the current location, the location of the entrance and exit, and the distribution of obstacles. An escape route is generated based on the room structure information and the target bounding box, and the escape route is converted into a warning voice.

6. The method according to claim 5, characterized in that, The process of generating an escape path based on the room structure information and the target bounding box includes: The target location is divided into multiple grid units; Based on the position coordinates and initial confidence score within the target bounding box, determine the corresponding smoke concentration value for each grid cell; An escape route is generated based on the current location, entrance / exit locations, obstacle distribution, and the smoke concentration value of each grid cell.

7. The method according to claim 5, characterized in that, The process of converting the escape route into a warning voice message includes: Based on the room structure information, multiple reminder keywords are determined on the escape route. The reminder keywords include room elements, escape direction, and escape distance. According to the order of room elements in the escape route, multiple reminder keywords are merged to obtain warning text, and the warning text is converted into warning voice.

8. A mobile firefighting robot, characterized in that, include: smoke Sensors are used to collect environmental smoke data at the target location; A visual sensor is used to acquire environmental image data of the target location; A control component is used to acquire environmental smoke data of the target location based on the smoke sensor, and environmental image data of the target location based on the vision sensor. By fusing the environmental smoke data and the environmental image data, a fire risk coefficient is obtained; If the fire risk coefficient exceeds the risk threshold, an early warning voice message associated with the escape route is generated based on the environmental smoke data and the environmental image data. A voice component is used to broadcast the warning voice message.

9. The robot according to claim 8, characterized in that, The smoke sensor is mounted on the head and neck area of ​​the fire alarm robot via an adjustable bracket.

10. A computer-readable storage medium, characterized in that, It stores a computer program, which, when executed on a processor, implements the fire early warning method according to any one of claims 1-7.