Fire rescue intelligent auxiliary system and method based on multi-mode perception

By adjusting voice volume, simplifying the interface, and generating directional escape instructions through a multimodal perception system, the problem of unclear information reception and path failure for firefighters in high-noise environments has been solved, achieving efficient and safe fire rescue assistance.

CN121661767APending Publication Date: 2026-03-13BEIJING DONGFANG HAILONG FIRE PROTECTION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing fire rescue auxiliary systems suffer from problems such as fixed voice alarm volume that is easily masked in high-noise environments, making it difficult for firefighters to clearly receive critical information; redundant parameters on the interface distract firefighters when they move; escape direction indicators are not dynamic and paths are prone to failure; and multimodal perception data lacks collaborative processing, resulting in low emergency decision-making efficiency and high safety risks.

Method used

The helmet uses a microphone to capture ambient noise levels and adjusts the voice output volume; it determines movement status based on displacement rate and simplifies the interface; it uses a helmet camera to identify the location of dense smoke and generate directional escape instructions; it integrates an alarm unit to simultaneously output voice and a simplified interface, and optimizes interaction with touch-sensitive hot zones.

Benefits of technology

Ensuring clear voice transmission in high-noise environments, focusing on key information during rapid movement, dynamically planning safe passages, improving emergency decision-making efficiency, and reducing safety risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121661767A_ABST
    Figure CN121661767A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of multi-modal data fusion and emergency decision, in particular to a fire rescue intelligent auxiliary system and method based on multi-modal perception. According to the invention, the data acquisition unit obtains the intensity of environmental noise, the displacement rate of firemen, the distribution position of dense smoke, the direction of a safety channel and the remaining time of an air respirator through helmet equipment, and the multi-mode coupling processing unit dynamically adjusts the voice volume according to the intensity of noise when triggering high-priority alarm in the remaining time. The method comprises the steps of activating interface simplification based on a displacement rate, judging dense smoke front blocking through space calculation, planning an optimal channel in combination with a path cost model, generating a composite semantic instruction bound with remaining time and orientation keywords, synchronously broadcasting a voice instruction and outputting a simplified interface by an integrated alarm unit, and setting a touch control hot area on the simplified interface. Freezing refreshing and direction identification flickering are supported, and fire rescue emergency decision-making efficiency and safety are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multimodal data fusion and emergency decision-making technology, and more specifically, to an intelligent auxiliary system and method for fire rescue based on multimodal perception. Background Technology

[0002] Multimodal data fusion and emergency decision-making is an important technology, specifically applied to the safety protection and escape guidance of firefighters in high-risk emergency scenarios such as fire scenes. The core is to integrate multi-dimensional perception data and dynamically adapt to the needs of emergency scenarios to achieve clear transmission of voice commands, focused presentation of key information, and accurate escape guidance, thus meeting the stringent requirements of fire rescue for real-time performance, reliability, and safety. Current fire rescue auxiliary technologies have significant limitations, making it difficult to meet the emergency assistance needs in high-risk scenarios. Existing systems lack a dynamic adaptation mechanism for environmental noise and voice output. The fixed volume of voice alarms is easily masked in the high-noise environment of a fire, preventing firefighters from clearly receiving critical alarm information. The display interface is not optimized based on the firefighter's movement status, resulting in excessive redundant parameters during rapid movement, which distracts attention and delays in obtaining core safety parameters such as the remaining time of the air respirator. Escape direction guidance is based on fixed orientation and is not linked to the distribution of dense smoke and the firefighter's real-time orientation, making the path invalid due to frontal obstruction by dense smoke. Furthermore, there is no dynamic planning of optimal safety passages, lacking adaptation to dynamic changes in the fire environment. At the same time, multimodal perception data lacks collaborative coupling processing, and voice commands, display interfaces, and escape guidance are not synchronized, failing to form a closed-loop assistance. This leads to low emergency decision-making efficiency and high safety risks for firefighters, failing to meet the core needs of rapid response and accurate evacuation in high-risk scenarios such as fires. To address this issue, we provide a fire rescue intelligent auxiliary system and method based on multimodal perception. Summary of the Invention

[0003] The purpose of this invention is to provide a fire rescue intelligent auxiliary system and method based on multimodal perception to solve the problems mentioned in the background art.

[0004] To achieve the above objectives, a fire rescue intelligent auxiliary system based on multimodal perception is provided, including: The data acquisition unit is used to acquire ambient noise intensity values ​​through the helmet microphone, acquire firefighter displacement rate through the positioning module, acquire on-site images through the helmet camera and identify the distribution location of dense smoke and the location of safety passages, and acquire the remaining time of the air respirator through the sensor. The multimodal coupling processing unit performs the following operations when a high-priority alarm is triggered during the remaining time of the air respirator: First, the voice output volume is adjusted based on the ambient noise intensity value. This includes comparing the current noise intensity value with a preset noise threshold. If the current noise intensity value exceeds the preset noise threshold, a target volume equal to the current noise intensity value plus a compensation amount is generated. The current movement state is determined based on the firefighter's displacement rate. When in the first movement state, the interface simplification engine is activated. The interface simplification engine hides non-emergency parameter data fields, retaining only the remaining time value and escape direction indicator. The remaining time value is fixed in a magnified form in the top center area of ​​the display interface. At the same time, the escape direction indicator is superimposed on the preset coordinate position of the real-time camera image. Finally, directional escape instructions are generated based on the location of the dense smoke distribution and the direction of the safety passage. This includes analyzing the relative relationship between the location of the dense smoke distribution and the direction the firefighter is facing. If it is determined that there is a frontal obstruction, the nearest safety passage is automatically associated, and the direction keyword is bound to the remaining time value to generate a compound semantic structure. The integrated alarm unit broadcasts voice content containing complex semantic structures at the target volume, and simultaneously outputs a simplified display interface generated by the interface simplification engine.

[0005] The second objective of this invention is to provide a method for implementing a multimodal perception-based intelligent auxiliary system for fire rescue, comprising the following steps: S1. Real-time capture of ambient sound wave signals through helmet microphone and conversion into noise intensity value; continuous acquisition of firefighter displacement rate through positioning module; acquisition of on-site images through helmet camera and identification of smoke distribution location and safety passage location by pre-trained model; and monitoring of remaining time through breathing apparatus sensor. S2. When a high-priority alarm is triggered by the remaining time, the noise intensity value is input to the threshold comparator and compared with the preset threshold in real time. If it exceeds the threshold, the compensation mechanism is activated. The compensation amount is matched by a lookup table and the programmable amplifier is driven to increase the amplitude of the voice signal to the target level. At the same time, the first movement status flag is activated based on the displacement rate change gradient and mean, triggering the video memory control module to close the non-emergency parameter layer. The font engine is called to enlarge the remaining time value to the top center of the interface according to the preset ratio. At the same time, the escape direction indicator is superimposed on the preset coordinate position of the camera video stream by the augmented reality fusion module as a transparent vector layer. The gyroscope orientation data and the coordinates of the smoke distribution position are spatially calculated. The frontal obstruction event is determined by the cosine angle model. The optimal channel cost is generated by the path cost calculation model, which integrates distance, curvature and smoke concentration gradient. Eight-directional keywords are output and bound to the remaining time to generate a composite semantic structure. S3. Broadcast voice commands with complex semantic structures at the compensated target volume, and simultaneously output a simplified display interface. The voice content integrates the remaining time and location keywords, and the interface highlights and enlarges the remaining time and dynamic direction indicators. S4. Set a touch hotspot in the simplified interface. Respond to a single click operation to freeze the interface refresh and activate the directional indicator pulse flashing. After a specified duration, the data update will be automatically restored.

[0006] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention acquires environmental noise, displacement rate, smoke location, and remaining time for emergency breathing apparatus from multiple dimensions through a data acquisition unit. When a high-priority alarm occurs on the emergency breathing apparatus, the multimodal coupling processing unit dynamically adjusts the voice volume according to noise intensity, matching compensation amounts to prevent voice from being masked, thus solving the problem of unclear command reception under high noise conditions. The interface is simplified based on displacement rate activation, hiding non-emergency parameters, amplifying the remaining time, and overlaying directional indicators with AR to address information acquisition delays during rapid movement. Frontal obstruction is determined through gyroscope and smoke location spatial calculation, and the optimal channel is selected using a path cost model to generate directional commands, avoiding path failure. The integrated alarm unit synchronizes voice and interface, and the touch hotspot optimizes interaction, achieving multimodal data collaboration and dynamic scene adaptation. This improves the efficiency of firefighters' emergency decision-making, reduces safety risks, and meets the needs of high-risk scenarios. Attached Figure Description

[0007] Figure 1 This is an overall block diagram of the present invention; Figure 2 This is the overall flowchart of the present invention.

[0008] The meanings of the labels in the diagram are as follows: 1. Data acquisition unit; 2. Multimodal coupling processing unit; 3. Integrated alarm unit. Detailed Implementation

[0009] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0010] This invention provides an intelligent auxiliary system for fire rescue based on multimodal perception. Please refer to [link / reference]. Figure 1 As shown, it includes: Data acquisition unit 1 is used to acquire ambient noise intensity values ​​through the helmet microphone, acquire firefighter displacement rate through the positioning module, acquire on-site images through the helmet camera and identify the distribution location of dense smoke and the location of safety passages, and acquire the remaining time of the air respirator through the sensor. When the multimodal coupling processing unit 2 triggers a high-priority alarm during the remaining time of the air respirator, it performs the following operations: First, the voice output volume is adjusted based on the ambient noise intensity value. This includes comparing the current noise intensity value with a preset noise threshold. If the current noise intensity value exceeds the preset noise threshold, a target volume equal to the current noise intensity value plus a compensation amount is generated. The current movement state is determined based on the firefighter's displacement rate. When in the first movement state, the interface simplification engine is activated. The interface simplification engine hides non-emergency parameter data fields, retaining only the remaining time value and escape direction indicator. The remaining time value is fixed in a magnified form in the top center area of ​​the display interface. At the same time, the escape direction indicator is superimposed on the preset coordinate position of the real-time camera image. Finally, directional escape instructions are generated based on the location of the dense smoke distribution and the direction of the safety passage. This includes analyzing the relative relationship between the location of the dense smoke distribution and the direction the firefighter is facing. If it is determined that there is a frontal obstruction, the nearest safety passage is automatically associated, and the direction keyword is bound to the remaining time value to generate a compound semantic structure. The integrated alarm unit 3 broadcasts voice content containing complex semantic structures at the target volume, and simultaneously outputs a simplified display interface generated by the interface simplification engine.

[0011] The operation of comparing the current noise intensity value with the preset noise threshold specifically includes: The raw sound wave signal collected by the helmet microphone is processed in frames by a real-time audio processing circuit. Each frame of the signal is converted into a spectral energy distribution by a fast Fourier transform, and then the ambient noise intensity value is calculated by weighted integration. The ambient noise intensity value is input into a threshold comparator and differentially calculated with a preset noise threshold. When the differential value is positive, a voice compensation flag signal is triggered. The threshold comparator uses hardware logic gate circuits to realize real-time judgment, and the output is connected to the enable port of the target volume generation module.

[0012] The operation of generating a target volume equal to the current noise intensity value plus a compensation amount specifically includes: When the threshold comparator outputs a voice compensation flag signal, the digital gain controller is activated. The digital gain controller reads the current noise intensity value and calls the compensation lookup table. The compensation lookup table stores fixed compensation values ​​corresponding to different noise intensity ranges. The fixed compensation values ​​are loaded onto the gain control pin of the programmable amplifier through the digital-to-analog converter, so that the amplitude of the alarm voice signal output by the voice synthesis chip is increased to the compensated level. At the same time, the amplitude limiting circuit is used to prevent signal overload distortion, and finally the speaker is driven to output penetrating alarm voice.

[0013] The operation of determining the current movement status based on the firefighter's displacement rate specifically includes: The displacement data stream output by the positioning module is input into the motion state classifier. The motion state classifier calculates the displacement rate change gradient through a sliding time window. When the displacement rate change gradient exceeds the dynamic threshold and the average rate is greater than the state judgment benchmark value within three consecutive time windows, the first motion state flag is activated. The start signal of the interface simplification engine is electrically interlocked with the flag signal to ensure that the simplified interface is forcibly enabled in the rapid movement state.

[0014] The interface simplification engine's operation of hiding non-urgent parameter data fields specifically includes: A priority mapping database is established, marking the remaining time of the air respirator and the escape direction indicator as the highest priority parameters, and the battery voltage and altitude data as the lowest priority parameters. When the first movement status flag is activated, the graphics rendering interface is called to close the memory address of the display layer corresponding to the lowest priority parameter, while the highest priority parameter is written to the video memory buffer. The escape direction indicator is dynamically generated by the vector graphics engine, and its shape is adaptively switched to a directional arrow or a path light strip according to the orientation of the safety passage.

[0015] The operation of fixing the remaining time value in a magnified form at the top center of the display interface specifically includes: The font scaling engine is called to parse the character width of the remaining time value, calculate the maximum readable size according to the screen resolution, generate an enlarged font with the rule of occupying 30% of the top area of ​​the interface, and at the same time enable the anti-aliasing rendering pipeline to eliminate pixelated edges. The operation of overlaying the escape direction indicator onto the real-time camera footage is achieved through an augmented reality fusion module. This includes acquiring frame buffer data from the camera video stream, creating a transparent layer at preset coordinates, overlaying the vector direction indicator onto the video stream in alpha blending mode, and adjusting the overlay angle in real time by matching the camera attitude data.

[0016] The specific procedures for analyzing the relative position of dense smoke distribution and the direction firefighters are facing include: The orientation data output from the helmet gyroscope is input into the spatial relationship solver. The solver transforms the coordinates of the smoke distribution location obtained from image recognition to the helmet coordinate system, calculates the angle between the center point of the smoke area and the facing direction axis, and establishes an obstruction determination model based on the cosine value of the angle. When the cosine value is greater than the orientation correlation threshold, it is determined that the smoke is located in the frontal field of view and a path obstruction event signal is generated.

[0017] The operation of automatically associating the location of the nearest safe passage specifically includes: After receiving the path obstruction event signal, the path decision-maker retrieves the set of safe passage locations output by image recognition, evaluates the accessibility of each passage through the path cost calculation model, and generates a cost value by combining the straight-line distance of the passage, the curvature of the path, and the gradient of the smoke diffusion concentration. The passage with the lowest cost value is selected as the optimal escape path. The semantic binding engine reads the direction angle data of the optimal path, quantifies it into eight-directional keywords, and embeds the directional placeholders of the voice template to generate a composite semantic structure.

[0018] The operation of the path cost calculation model further includes: A 3D environmental topology map is constructed, mapping the distribution location of dense smoke identified in the image to the orientation of safety passages as grid nodes. The shortest path distance from the firefighter's current location to each passage node is calculated using the Dijkstra algorithm. A dense smoke diffusion simulator is introduced to predict the smoke coverage range within a specified time period in the future. The path accessibility coefficient is dynamically adjusted. The final cost is generated by weighted summation of path length weight, accessibility coefficient weight, and path turning penalty coefficient. The output is connected to the optimal path selector.

[0019] Further explanation is needed: After the helmet-mounted voice communication system completes the acquisition of the original audio signal, to ensure that operators can clearly receive voice commands in complex environments and to prevent environmental noise from obscuring key information, it is necessary to accurately determine whether to activate the voice compensation mechanism by detecting and comparing real-time noise intensity with a threshold. The core of this process is to quantify and compare the real-time environmental noise intensity with a preset threshold, and trigger the compensation signal through a rapid hardware-level response to ensure the clarity and real-time performance of voice communication. The specific implementation method is as follows: First, the raw sound wave signal collected by the helmet microphone is processed by a real-time audio processing circuit, which integrates signal amplification, filtering, and analog-to-digital conversion. This circuit can complete signal processing in milliseconds, adapting to the rapidly changing noise characteristics of the work environment. The raw sound wave signal is an analog ambient sound signal collected by the helmet's built-in omnidirectional microphone, including work noise, background noise, and possible speech signals. The sampling frequency is set to 16kHz to ensure coverage of the human ear's audible range. Frame processing divides the continuous analog sound wave signal into discrete frames of fixed length to avoid processing delays caused by excessively long signals. The frame length is set to 20 milliseconds, with a 5-millisecond overlap between frames (25% overlap rate), ensuring signal continuity. To improve both accuracy and reduce inter-frame information loss, the real-time audio processing circuit uses a built-in sample-and-hold circuit to extract the original signal at 20-millisecond intervals. Each frame is amplified to a suitable amplitude by a preamplifier to avoid weak signals or saturation distortion. A low-pass filter (cutoff frequency 8kHz) then filters out high-frequency interference, resulting in a clean single-frame analog signal. The processed analog signal is then converted into a spectral energy distribution using a Fast Fourier Transform (FFT). The FFT converts a time-domain (time dimension) sound wave signal into a frequency-domain (frequency dimension) energy distribution, visually representing the energy intensity of noise at different frequencies. In operational scenarios, interference noise in speech is mostly concentrated in the 500Hz-4kHz frequency band. Through FFT... The system can accurately extract energy information for this frequency band. In practice, the FFT processing unit built into the real-time audio processing circuit first converts the single-frame analog signal into a 16-bit digital signal via an analog-to-digital converter (ADC), and then performs a 1024-point FFT operation on the digital signal to obtain the energy values ​​corresponding to 512 frequency points within the 0-8kHz frequency band, forming a spectral energy distribution map of the signal frame. The energy value of each frequency point in the map is expressed in decibels, intuitively reflecting the strength of the corresponding frequency noise. Subsequently, the ambient noise intensity value is calculated through weighted integration. The weighted integration takes into account the differences in human ear sensitivity to different frequencies of sound, assigning differentiated weights to the energy values ​​of different frequency bands, and then accumulating the total energy to avoid the problem of simply calculating an average. To address the issue of misjudgment of noise intensity caused by energy, the system first divides the frequency spectrum into three bands: low frequency (below 500Hz), mid frequency (500Hz-4kHz), and high frequency (above 4kHz). The mid frequency band (speech-sensitive band) is weighted at 1.0, the low frequency band at 0.3, and the high frequency band at 0.5 (the weights are calibrated through extensive human hearing experiments). Then, the energy distribution spectrum obtained by the Fast Fourier Transform is used to extract the energy value of each frequency point according to the frequency band. These values ​​are then multiplied by their corresponding weights and summed to obtain the weighted energy value of a single frame signal. Finally, the average of the weighted energy values ​​from 10 consecutive frames is taken as the ambient noise intensity value at the current moment, ensuring stable results and avoiding false triggering caused by abnormal signals in a single frame.

[0020] The calculated ambient noise intensity value is input into a threshold comparator and differentially analyzed with a preset noise threshold. The threshold comparator is a hardware module used to quantify and compare the magnitudes of two electrical signals. Its core function is to calculate the difference between the electrical signal corresponding to the ambient noise intensity and the electrical signal corresponding to the preset threshold to determine whether the noise exceeds the limit. The preset noise threshold is a critical value set based on the voice communication requirements of the work scenario. By statistically analyzing common noise intensities in different work environments and considering the minimum signal-to-noise ratio at which the human ear can clearly distinguish speech, the preset noise threshold is set to 85dB. This value is stored in the comparator's built-in register and can be manually adjusted according to the actual scenario via the system interface. In practice, the ambient noise intensity value is first processed by a digital-to-analog converter. The noise is converted by the DAC into an analog voltage signal proportional to the intensity. The threshold comparator reads the reference voltage signal corresponding to the preset noise threshold from the register, and then performs a differential operation, subtracting the reference voltage signal from the voltage signal corresponding to the ambient noise intensity to obtain the difference value. The differential operation is implemented by the operational amplifier built into the comparator, with a response time of less than or equal to 1 millisecond, meeting real-time requirements. When the difference value is positive, it indicates that the ambient noise intensity exceeds the preset threshold. At this time, the threshold comparator triggers the voice compensation flag signal. A positive difference value means that the voltage signal corresponding to the ambient noise intensity is greater than the reference voltage signal, which means that the current noise may affect the speech clarity, and the compensation mechanism needs to be activated. The voice compensation flag signal is a high-level digital signal. As the control signal that triggers subsequent voice compensation functions, a low level (0V) will not trigger compensation. The threshold comparator uses hardware logic gates for real-time judgment, unlike software processing. The hardware logic gates consist of AND gates, OR gates, NOT gates, and comparators. They do not require processor computation and can directly perform signal comparison and judgment through circuit logic, with a response time controlled in the microsecond range. This ensures instantaneous compensation triggering when noise suddenly increases, preventing voice information loss. In specific implementation, the hardware logic gates determine whether the voltage signal after differential operation is greater than 0V. If it is greater than 0V, a high-level voice compensation flag signal is output; if it is less than or equal to 0V, a low-level output is maintained. The output of the threshold comparator is directly connected to... The enable port of the target volume generation module is its control interface. It is used to receive signals to start or stop the compensation function. When a high-level voice compensation flag signal is received, the enable port is activated, and the target volume generation module immediately starts the voice compensation algorithm. When no compensation flag signal (low level) is received, the enable port is closed, and the target volume generation module outputs voice at the default volume. This ensures that unnecessary volume overload is not generated when the noise level is within acceptable limits, thus ensuring clear communication while avoiding discomfort to the human ear caused by overcompensation. It enables real-time monitoring and rapid compensation response of environmental noise, ensuring that operators can still clearly receive voice commands in complex noisy environments, thereby improving operational safety and communication efficiency.

[0021] After the threshold comparator outputs the voice compensation flag signal, a target volume for the superimposed compensation amount needs to be generated through digital gain control and signal amplification links. The core is to match the corresponding compensation amount according to the current ambient noise intensity, amplify the alarm voice signal to a level that can penetrate noise, and at the same time, prevent signal distortion through hardware protection to ensure that operators can still clearly receive alarm information in high-noise environments. The specific implementation method is as follows: When the threshold comparator outputs a high-level voice compensation flag signal, this signal directly triggers the digital gain controller (DGB). The DGB is a hardware module integrating digital signal processing and gain adjustment logic. Its core function is to dynamically adjust the amplification of the voice signal based on noise intensity, with a response time of less than or equal to 2 milliseconds. This adapts to the real-time interaction requirements of noise and alarm signals in operational scenarios. After startup, the DGB reads the current ambient noise intensity value (in decibels) from the result register of the real-time audio processing circuit via the system's internal data bus. Simultaneously, it requests a lookup table for the compensation amount. The compensation amount lookup table is a structured data table pre-stored in the controller's built-in read-only memory (ROM). Its core principle is to divide noise intensity into intervals and assign a fixed compensation value to each interval. The compensation amount is set based on a clear reception standard of a voice signal-to-noise ratio (SNR) greater than or equal to 15dB, meaning the compensated voice volume must be 15dB higher than the current noise level. For example, for a noise intensity range of 85-90dB, the fixed compensation amount is 8dB; for 90-95dB... The B range corresponds to a compensation of 12dB, and the 95-100dB range corresponds to a compensation of 15dB. Both the range division and the compensation amount have been calibrated through extensive real-world testing in various operational scenarios to ensure that the compensated speech can penetrate noise without damaging hearing due to excessive volume. When invoked, the digital gain controller first determines the range to which the current noise intensity value belongs, then retrieves the corresponding fixed compensation value from a lookup table to complete the compensation matching. The fixed compensation value output from the lookup table is a digital signal, which needs to be converted into an analog voltage signal by a digital-to-analog converter (DAC) before being applied to the gain control pin of the programmable amplifier. The DAC is a hardware component that converts digital signals into analog voltage signals proportional to the value. Here, a 12-bit resolution DAC is selected to ensure that the compensation conversion accuracy is less than or equal to 0.1dB, avoiding insufficient or excessive compensation due to conversion errors. In practice, the DAC receives the compensation digital signal output from the digital gain controller and converts it into the corresponding analog voltage; for example, 12dB compensation corresponds to 1.2V, and 8dB compensation corresponds to 0.The 8V voltage and its mapping relationship with the compensation amount are preset through the DAC calibration program, with a calibration cycle of every 3 months to ensure long-term accuracy. The programmable amplifier (PFA) is an amplifier whose gain can be adjusted via an external voltage signal. Its gain control pin is the key interface for receiving external voltage and adjusting the amplification factor. The higher the voltage, the greater the amplification factor. After directly loading the analog voltage signal output from the DAC onto this pin, the PFA's amplification factor is set to a value matching the compensation amount, preparing for subsequent voice signal amplification. At this time, the original alarm voice signal output from the voice synthesis chip is input to the PFA, and after amplification, its amplitude is increased to the compensated level. The voice synthesis chip is a dedicated chip that pre-stores alarm voice data and can output standard analog voice signals. Its original output signal amplitude typically corresponds to a volume of 30dB. In high-noise environments, it needs to be amplified by the programmable amplifier. For example, if the current compensation amount is 12dB, the amplifier amplifies the original signal by 12dB, making the output signal amplitude correspond to 42dB. After adding 92dB of ambient noise, the actual received voice signal-to-noise ratio is 15dB, meeting the requirements for clear reception.

[0022] To prevent the amplified signal from exceeding the circuit's maximum output range due to excessive amplitude, a limiting circuit is connected in series at the output of the programmable amplifier. This limiting circuit is a hardware protection circuit composed of a bidirectional Zener diode and an operational amplifier. Its preset maximum allowable amplitude is ±1.2V, corresponding to a voice volume of 60dB, which is the upper limit of human hearing safety. When the amplified signal amplitude exceeds this threshold, the limiting circuit automatically clips the excess portion, retaining only the signal within ±1.2V. If the signal amplitude is within the threshold range, it passes directly. This ensures that the voice signal is not overloaded or distorted, and also prevents excessive volume from damaging the worker's hearing. Finally, the compensated voice signal after amplitude limiting is further amplified by a power amplifier to a power level capable of driving a speaker, ultimately driving the helmet's built-in full-range speaker to output a penetrating alarm. The voice and power amplifier circuit amplifies small signals to sufficient power to drive the speaker. A Class D power amplifier with an efficiency of ≥85% is selected here to avoid affecting helmet safety due to heat generation. The speaker is a 40mm diameter full-range unit with a frequency response range of 200Hz-8kHz. The sound efficiency of the 1-4kHz audio band is optimized to ensure the recognizability of the alarm voice. The output penetrating alarm voice can cover a range of 3-5 meters in the working environment. In a 92dB noise environment, the voice recognition rate of the operator is ≥95%, which fully meets the alarm requirements of high-noise working scenarios. It ensures that the alarm voice penetrates the ambient noise and avoids signal distortion and hearing damage through multiple hardware protections, achieving clear and safe transmission of alarm voice in high-noise working scenarios.

[0023] In the interface interaction system of firefighters' work equipment, when firefighters are moving rapidly, the redundant information in the complex interface increases the operational burden and may even delay the reception of critical instructions. Therefore, it is necessary to accurately determine their movement status based on displacement rate and force the activation of a simplified interface through hardware-level linkage. This process is based on real-time analysis of positioning data, combined with dynamic thresholds and electrical interlocks, to ensure the accuracy of status determination and the mandatory switching of the interface. The specific implementation method is as follows: First, the displacement data stream output by the positioning module is continuously input into the motion state classifier. The positioning module is a dual-mode positioning component integrated into a firefighter's helmet or tactical vest, with a sampling frequency of 1Hz, outputting one data point per second. This satisfies the accuracy requirements of displacement analysis while avoiding data redundancy. The displacement data stream is structured data containing timestamps and latitude / longitude coordinates. The motion state classifier is an embedded hardware module whose core function is to analyze the displacement rate and determine the motion state through a sliding time window. After receiving the displacement data stream, it first verifies the validity of the coordinate data before initiating the sliding time window processing. The sliding time window is an analysis window that extracts continuous data at a fixed length. The window length is set to 3 seconds, covering a movement distance of approximately 3 steps for the firefighter, reflecting the actual movement trend. The step size is set to 1 second, and the window data is updated every second to ensure real-time performance. That is, each window contains 3 displacement data points from the most recent 3 seconds. The motion state classifier calculates the gradient of displacement rate change through the sliding time window. The specific calculation is as follows: 计算窗口内的平均位移速率,先将窗口内首尾两条数据的经纬度坐标转换为地面直线距离,例如窗口内首条坐标(116.325°E,39.938°N)、尾条坐标(116 .326 °E,39.939°N),计算得直线距离约15米,再用该距离除以3秒,得到平均速率5米 / 秒,计算位移速率变化梯度,即当前窗口的平均速率与前一个窗口的平均速率的差值绝对值, 例如当前窗口速率5米 / 秒,前一窗口速率3米 / 秒,梯度为2米 / 秒2,该梯度值反映速率变化的剧烈程度,梯度越大,说明消防员移动速度提升越快,越可能处于紧急快速移动状态。

[0024] 当连续三个时间窗内位移速率变化梯度均超过动态阈值且速率均值大于状态判定基准值时,移动状态分类器激活第一移动状态标志,动态阈值是根据消防员日常训练与实战数据设 定的速率变化临界值,通过统计100名消防员的快速移动数据,确定梯度动态阈值为0.8米 / 秒2,该值能区分快速加速移动与正常行走的缓慢速率变化,状态判定基准值是界定快 速移动与正常移动的速率临界值,设为1.5米 / 秒,消防员正常行走速率约1米 / 秒,快速移动或冲锋速率通常大于等于1.5米 / 秒,具体判定时,分类器会连续监测三个滑动时间窗 的计算结果,若第一个窗口梯度0.9米 / 秒2、速率1.6米 / 秒(双条件满足),第二个窗口梯度1.0米 / 秒2、速率1.8米 / 秒(双条件满足),第三个窗口梯度0.9米 / 秒2、速率2.0 米 / 秒(双条件满足),则判定为连续三个窗口满足条件,此时分类器的状态寄存器会输出高电平(3.3V)的第一移动状态标志信号,该标志代表消防员当前处于紧急快速移动状态 ,需强制启用简化界面,界面简化引擎的启动信号与第一移动状态标志信号通过电气互锁实现强制联动,界面简化引擎是控制消防员装备显示界面的硬件模块,其核心功能是隐藏 冗余功能入口,仅保留核心功能显示,减少界面操作步骤与视觉干扰,电气互锁是通过硬件逻辑门电路(与门)实现的强制关联机制,确保只有当第一移动状态标志激活时,界面简 化引擎才能启动,避免手动切换延迟或遗漏,具体实施时,与门电路的两个输入端分别连接第一移动状态标志信号与界面供电信号,当第一移动状态标志为高电平(快速移动)且界 面供电正常(设备未断电)时,与门输出高电平信号,该信号直接接入界面简化引擎的使能端口,引擎接收到使能信号后,立即调用预存的简化界面配置文件,在50毫秒内完成界面 切换,若第一移动状态标志为低电平(非快速移动),则与门输出低电平,引擎使能端口无信号,界面保持正常显示模式,这种电气互锁机制无需软件干预,完全通过硬件逻辑实现, 响应时间小于等于100毫秒,确保消防员在快速移动启动瞬间,界面已同步简化,不会因软件延迟影响操作效率,既能精准识别消防员的快速移动状态,又能强制启用适配快速操作 的简化界面,有效平衡了作业安全性与操作便捷性,为消防员在紧急场景下的高效设备交互提供了可靠保障。

[0025] After the simplified engine is activated by the first movement status indicator via electrical interlock, the core operation of the engine is to hide non-emergency data fields according to parameter priority. When firefighters move quickly, they only need to focus on the core parameters that are directly related to life safety. Redundant low-priority parameters will distract attention and prolong the time to obtain critical information. Therefore, it is necessary to first establish a parameter priority system to clarify the display rules, and then achieve precise display switching through memory address control and video memory writing. At the same time, it is necessary to adapt to the dynamic shape requirements of the escape direction indicator. The specific implementation method is as follows: First, a priority mapping database is established. This database is a structured data storage module that stores all interface display parameters and their priority levels. It uses a lightweight SQLite database to adapt to the low storage and fast read requirements of embedded devices. The database table structure contains six core fields: parameter ID, parameter name, priority level, display layer number, memory address offset, and refresh rate. The priority level is divided into three levels: highest priority (P1), medium priority (P2, reserved for temporary calls in special scenarios), and lowest priority (P3). The core logic of marking priorities is based on the immediate safety impact of firefighters in actual combat scenarios. The remaining time of the air respirator directly determines the firefighter's safe stay limit in the fire scene. If the remaining time is less than 5 minutes, immediate evacuation is required. This is a life safety parameter that is obtained without delay. The escape direction indicator provides key information for emergency evacuation. Orientation guidance, if obtained incorrectly or delayed, will cause deviations in the escape route. Both are marked with the highest priority (P1). Battery voltage only reflects the power supply status of the device (the device has a built-in low-voltage alarm function, and will actively pop up a window to remind when the voltage is too low, so continuous monitoring is not required when rapid movement is not necessary). Altitude data is mostly used for vertical fire spread analysis (an auxiliary decision-making parameter during non-emergency movement, which does not directly affect immediate escape). Neither of these will affect the immediate safety of firefighters, so they are marked with the lowest priority (P3). The database is initially configured before the device leaves the factory. Subsequently, it can be connected to the management terminal via USB interface to adjust the priorities according to different fire scene scenarios to ensure adaptation to actual needs. The interface is simplified. When the engine starts, it will read the priority of all parameters, the corresponding display layer number, and memory address information through the database interface and store them in the engine's temporary cache area to prepare data for subsequent display control.

[0026] When the first movement state flag is activated (outputs a high-level signal), the interface simplification engine immediately calls the graphics rendering interface to execute display control operations. The graphics rendering interface is a software interface connecting the engine and the display hardware, supporting the enabling / disabling of independent display layers with different parameters. Each parameter corresponds to a dedicated display layer, and each layer has a fixed contiguous address space in the device memory. By disabling read and write permissions for this address space, layer data can be prevented from being transferred to the display hardware, thus hiding the parameter. In practice, the engine first retrieves the display layer number of the lowest priority (P3) parameter from the temporary buffer, calls the layer disable function of the graphics rendering interface, and... Upon receiving the layer number, the interface sends an address lock signal to the device memory controller, preventing the CPU from writing or reading data from the memory address ranges corresponding to layers 4 and 5. At this time, the display data for these two layers cannot be updated, and the existing display content will be cleared during the next frame refresh of the display hardware. The battery voltage and height data fields will no longer appear on the interface. While hiding low-priority parameters, the interface simplification engine writes the display data for the highest-priority (P1) parameter to the video memory buffer. The video memory buffer is a high-speed SRAM area used by the display hardware to temporarily store image data to be displayed. The buffer is configured according to pixel coordinates - RGB color values ​​and transparency. The data is stored in a formatted manner and divided into fixed display areas. In specific operations, the engine first reads the real-time data of the remaining time of the air respirator, calls the built-in font rendering module to generate bold white font size 24, and then writes the font image data to the (150,60)-(350,110) pixel coordinate area of ​​the buffer through the video memory write function. The upper half of the interface is centered to avoid eye deviation. For the escape direction indicator, its display data is dynamically generated by the vector graphics engine. The vector graphics engine is a rendering module that generates graphics based on mathematical paths. Unlike bitmaps, vector graphics can adapt to different display resolutions without distortion. The engine will first receive the output of the positioning module. The system calculates the angle between the escape direction and the firefighter's current orientation based on the safety passage's location data. If the angle is less than or equal to 30° (the safety passage is directly in front or near), a path light strip is generated (the light strip is a yellow gradient line, 12 pixels wide, extending to the bottom edge of the interface, with the safety passage in white font size 4 at the end). If the angle is greater than 30° (the safety passage is to the side or behind), a directional arrow is generated (the arrow is a red solid triangle, 35 pixels on each side, pointing in the same direction as the safety passage, with a blinking effect at 100 milliseconds per second). The generated vector graphic data is written to the buffer at (200, 130) - (400, ...).A 250-pixel coordinate area, located below the remaining time on the breathing apparatus, forms a vertical visual flow guiding core parameters and directions. This ensures firefighters can continuously obtain critical information without shifting their gaze, enabling precise concealment of non-emergency parameters and focused display of core parameters. Furthermore, all memory control and video memory write operations are completed within 150 milliseconds, avoiding delays in decision-making due to interface switching latency, and providing clear interface support for safe operations in emergency scenarios.

[0027] After hiding the low-priority parameters, the interface simplification engine needs to further enhance the visual prominence and scene adaptability of the highest-priority parameters. The remaining time of the breathing apparatus, as a core parameter directly determining the firefighter's escape window, needs to be fixed in a magnified form in the top center of the interface, where it is most easily visible, ensuring quick access without deliberate searching. If the escape direction indicator is displayed as a standalone graphic, it is easily detached from the complex background of the fire scene, leading to misjudgment of direction. Therefore, it needs to be dynamically overlaid with real-time camera footage through an augmented reality fusion module, while also adapting to changes in camera posture to adjust the angle. The specific implementation method is as follows: First, the interface simplification engine calls the font scaling engine to parse the character width of the remaining time value of the air respirator and calculates the maximum readable size based on the screen resolution. The font scaling engine is a software module used to dynamically adjust the character size and adapt to different display areas. It supports automatically calculating the scaling factor based on the target display area to avoid characters being cropped due to being too large or blurred due to being too small. The character width is the pixel width of a single character in the horizontal direction. In practice, the engine first obtains the displayed text of the current remaining time, calls the character measurement function of the font scaling engine to calculate the original width of each character, and calculates the total original width as 104 pixels. Then, it reads the screen resolution parameters. Taking a helmet HUD as an example... The common resolution is 1280×720. The top area of ​​the interface is defined as a full-screen width of 1280 pixels and a height of 30% of the total screen height, which is 216 pixels. This area is the range where firefighters' eyes naturally focus when moving quickly, and they can capture the image without having to look down or up. Then, the usable display area in the top center is calculated. To avoid characters being too close to the screen edge and causing cropping, 140 pixels of blank space are reserved on the left and right. The usable width is 1280-280=1000 pixels, and the usable height is 80% of 216 pixels, which is 173 pixels. A 20% blank space is reserved to avoid characters crowding up and down. Based on the ratio of the total original bit width to the usable width, the scaling factor is calculated (1000÷104≈9).6) Scale up the bit width and height of each character according to the calculated scaling factor, ensuring that the total width of all characters after scaling is less than or equal to 1000 pixels and the height is less than or equal to 173 pixels. This fills the available area without the risk of cropping, achieving the maximum readable size. After determining the character size, generate the scaled font according to the rule of occupying 30% of the top area of ​​the interface. At the same time, enable the anti-aliasing rendering pipeline to eliminate pixelated edges. The anti-aliasing rendering pipeline is a graphics processing flow that eliminates jagged distortion at the edges of characters through multisampling and pixel blending technology, avoiding stepped edges due to pixel dispersion after scaling, which affects readability. In specific implementation, the font scaling engine generates bitmap data of the characters according to the calculated scaling factor and calls the multisampling anti-aliasing function of the anti-aliasing rendering pipeline. This function samples the color of the four sampling points around each pixel. If some sampling points are located inside the character and some are located outside, the character color (preset is pure white, RGB value 255, 255, 255) and the background color (preset is semi-transparent red, RGB value 255, 0, 0, alpha value 12) are mixed proportionally. 8) To create a gradient transition at the edges and completely eliminate jagged edges, the engine adjusts the character spacing and line height to ensure that the characters occupy 30% of the top area. The horizontal character spacing is set to 20 pixels, and the characters are vertically centered in an area 216 pixels high at the top. The top of the characters is 23 pixels from the top of the screen, and the bottom is 23 pixels from the top. The total area of ​​the character area is approximately 1000 × 170 = 170,000 pixels, and the total area of ​​the top area is 1280 × 216 = 276,480 pixels. The 30% top area refers to 30% of the total screen height. The character area occupies the majority of the area within this 30%. For example, if the character height is 173 pixels and the top area height is 216 pixels, the character area occupies 80% of the top area height and 78% of the top area width (1000 ÷ 1280). Visually, the character area is prominent and not crowded within the top 30% area. After rendering, the enlarged character data is written to the top center coordinate area of ​​the video memory buffer to ensure that this area is prioritized for displaying the remaining time each time the screen refreshes, and is not covered by other elements.

[0028] Meanwhile, the AR fusion operation of the escape direction indicator is implemented through the augmented reality fusion module. The first step is to acquire the frame buffer data of the camera video stream. The augmented reality fusion module is a hardware and software collaborative module that synthesizes the virtual indicator with the real camera footage in real time. Its core is low-latency image data processing to avoid the indicator and the screen being out of sync. The frame buffer data is the temporary storage data in memory for each frame of image captured by the camera sensor. The format is usually YUV420, balancing image quality and data volume, and adapting to the processing capabilities of embedded devices. In specific implementation, the module subscribes to the real-time frame data push from the camera through the built-in camera driver interface. The camera sampling frequency is set to 30 frames / second to meet the smoothness of human vision. The driver stores each frame of YUV data in a preset frame buffer area. The module triggers a hardware interrupt every time the camera generates a frame through a frame ready interrupt signal, quickly reads the frame buffer data, and calls the format conversion function to convert the YUV data. The data is converted to RGB format for easy fusion with the RGB data of the virtual indicator. It is then stored in the temporary frame buffer of the fusion module. The entire data acquisition process has a latency controlled within 33 milliseconds to ensure no significant lag between the indicator and the image. A transparent layer is created at preset coordinates, and the vector direction indicator is overlaid onto the video stream using alpha blending mode. The transparent layer is a virtual layer with the same size as the frame, and its display intensity is controlled by the alpha value (transparency parameter, 0 for completely transparent, 255 for completely opaque). Here, the alpha value is set to 128 (50% transparency) to ensure the indicator is prominent without obscuring key information in the background. The preset coordinates are selected in the lower center of the frame, which is the natural point where the firefighter's line of sight naturally falls as they move down from the top of the remaining time, eliminating the need for significant line-of-sight adjustments. Alpha blending mode is a rendering technique that uses pixel-level color calculations to achieve a natural fusion of the virtual indicator and the background. The specific calculation logic is as follows: The blended pixel color = (indicator color × alpha value + background color × (255 - alpha value)) ÷ 255; for example, if the indicator is red (RGB255,0,0) and the corresponding background pixel is fire gray (RGB128,128,128), the blended color is (255×128 + 128×127) ÷ 255 = 191,63,63, presenting the effect of a red indicator superimposed on a gray background, clearly identifying the indicator without obscuring background details. The indicator shape (directional arrow or path light band) generated by the vector graphics engine is written to the preset coordinates of the transparent layer in real time. Then, through the layer composition function of the blending module, the transparent layer and the RGB image of the temporary frame buffer are blended pixel by pixel to generate composite frame data with the indicator. Finally, the blending module needs to perform real-time... The indicator's superimposed angle is adjusted to match the camera's attitude data, ensuring that the indicator always points to the safety passage. The camera attitude data is real-time rotation angle and tilt status data collected by the device's built-in inertial measurement unit, with an update frequency of 100Hz. It can accurately reflect the camera's horizontal rotation (yaw angle) and vertical tilt (pitch angle) status. The fixed geographical location of the safety passage is provided by the positioning module. The fusion module inputs the attitude data and orientation data into the angle calculation function to calculate the angle that the indicator needs to be adjusted. For example, if the camera rotates 30° to the right (yaw angle +30°), the indicator arrow needs to rotate 30° to the left to keep pointing due east. If the camera tilts down 15° (pitch angle -15°), the indicator arrow needs to tilt up 15° to avoid the indicator pointing to the ground instead of the safety passage due to the downward view of the screen.

[0029] Angle adjustment commands are sent to the vector graphics engine in real time. The engine regenerates the rotated indicator graphic, writes it to the transparent layer, and re-executes the fusion to ensure that the indicator in each frame of the composite image accurately matches the location of the safety passage. There will be no directional deviation due to camera rotation, and the escape direction indicator can be deeply integrated with the real scene. This completely solves the problems of delay in obtaining key information and disconnection in direction judgment when moving quickly, providing efficient interface support for firefighters' emergency escape.

[0030] While the augmented reality fusion module provides firefighters with dynamic escape directional guidance, the random distribution of dense smoke in a fire scene may obstruct the original directional path. If only directional guidance is relied upon while ignoring smoke obstruction, firefighters may be easily led into dangerous areas. Therefore, it is necessary to conduct spatial correlation analysis between the helmet gyroscope's attitude data and the location of dense smoke distribution to accurately determine whether the dense smoke is in the frontal field of vision and whether it obstructs the escape path, providing a basis for real-time adjustment of the escape direction. The specific implementation method is as follows: The orientation data output from the helmet gyroscope is continuously input into the spatial relationship solver. The solver transforms the smoke distribution coordinates obtained from image recognition to a unified helmet coordinate system, ensuring comparability within the same spatial framework. The helmet gyroscope is a MEMS (Micro-Electro-Mechanical Systems) inertial sensor integrated into the helmet liner. Its core function is to collect real-time head posture data of the firefighter, with a sampling frequency set to 100Hz (balancing accuracy and low power consumption). The output data includes a timestamp, yaw angle, and pitch angle. The yaw angle is the horizontal rotation angle around the helmet's vertical axis (z-axis) (0° corresponds to due north, 90° to due east), reflecting the firefighter's horizontal orientation. The pitch angle is the vertical rotation angle around the helmet's horizontal axis (y-axis) (0° corresponds to horizontal forward, +30° corresponds to looking up, -30° corresponds to looking down), reflecting the vertical orientation. The data is transmitted in real-time to the spatial relationship solver via an I2C bus. The processor is an embedded computing module based on the ARM Cortex-M7 core. It pre-stores the calibration parameters of the camera and helmet. Its core function is to realize coordinate transformation between different coordinate systems. The image recognition module (deployed in the edge computing unit of the helmet, using a lightweight YOLOv5 model, with an inference speed of greater than or equal to 30 frames / second) first performs smoke detection on the real-time image captured by the camera. By identifying areas with low grayscale values, blurred edges, and dynamic changes in the image, it marks them as smoke areas and outputs the two-dimensional pixel coordinates of the area in the camera coordinate system. Then, combined with the camera's intrinsic parameters (focal length 8mm, pixel size 1.12μm, also factory calibrated), it converts the pixel coordinates into three-dimensional coordinates in the camera coordinate system (the origin of the camera coordinate system is the lens optical center, the x-axis points forward, the y-axis points to the right, and the z-axis points upward. The three-dimensional coordinates are as follows: (3.0m, 0.5m, 0.2m) indicates that the smoke is 3 meters in front of the camera, 0.5 meters to the right, and 0 meters above the camera.(At a distance of 2 meters), the spatial relationship solver then calls the pre-stored calibration parameters to transform the three-dimensional coordinates of the smoke in the camera coordinate system to the helmet coordinate system. This coordinate system has the helmet's center of gravity as its origin, with the x-axis along the direction the firefighter is facing (forward), the y-axis horizontally to the right, and the z-axis vertically upward. The transformation process is achieved through translation and rotation. First, the offset distance is subtracted from the three-dimensional coordinates of the camera coordinate system to eliminate the positional deviation between the camera and the helmet. Then, the coordinates are rotated and corrected according to the relative angle (5° downward from the camera) (fine-tuning the vertical coordinates to ensure that the position of the smoke is consistent with the firefighter's viewpoint). Finally, the result is... The system calculates the three-dimensional coordinates of the smoke distribution location in the helmet coordinate system, establishing a spatial correlation between the smoke location and the firefighter's facing direction. After coordinate transformation, the spatial relationship solver calculates the angle between the center point of the smoke area and the axis of the firefighter's facing direction, and establishes a blocking judgment model based on the cosine value of the angle. The center point of the smoke area is the geometric center obtained by the image recognition module after extracting the contour of the smoke area and using a polygon fitting algorithm. First, all edge pixels of the smoke area are marked and fitted into an irregular polygon. Then, the average value of the coordinates of all vertices of the polygon is calculated as the center point coordinates, along with the facing direction axis. The x-axis of the helmet coordinate system represents the straight line in space directly in front of the firefighter, with a direction vector of (1,0,0) (x-axis pointing forward, y and z axes have no components). When calculating the angle between the two, the solver first constructs a spatial vector from the center point of the smoke to the helmet origin, i.e., the three-dimensional coordinates of the center point (2.95,0.52,0.25). Then, it calculates the spatial angle between this vector and the facing direction axis vector (1,0,0). Since the cosine of the angle directly reflects the proximity of the two vectors, it is used as the core indicator for obstruction determination, and an obstruction determination model is constructed. The model uses the cosine of the angle... The input is the value, and the output is whether it blocks the frontal path. A preset threshold is used for quantitative judgment to avoid subjective judgment errors. In the specific calculation, the solver uses the logic of vector dot product (the dot product of two vectors equals the product of their magnitudes multiplied by the cosine of the included angle). First, it calculates the dot product of the spatial vector and the axial vector (the result is the x-component of the spatial vector, 2.95). Then, it calculates the magnitudes of the two vectors separately (the magnitude of the spatial vector is the square root of the sum of the squares of its coordinates, and the magnitude of the axial vector is 1). Finally, it obtains the cosine of the included angle. The entire calculation process takes less than or equal to 1 millisecond, ensuring real-time response to changes in the fire scene.

[0031] When the cosine value of the included angle is greater than the azimuth correlation threshold, the spatial relationship solver determines that the dense smoke is located in the frontal field of vision and generates a path obstruction event signal, triggering an adjustment of the escape direction. The azimuth correlation threshold is a quantitative threshold set based on the actual frontal field of vision range of firefighters. By statistically analyzing the field of vision test data of 50 firefighters (when firefighters are standing naturally, the average effective horizontal frontal field of vision is about 60°, i.e., 30° to the left and right, and the vertical frontal field of vision is about 45°, i.e., 22.5° up and down), the cosine value corresponding to the 30° included angle in the horizontal direction is cos3. Since 0° ≈ 0.866, the azimuth correlation threshold is set to 0.866 (this can be fine-tuned via the management terminal based on individual firefighter field of vision differences, with an adjustment range of 0.8-0.9). The solver compares the real-time calculated cosine value of the angle with this threshold. If the cosine value is greater than 0.866, it means that the angle between the center point of the dense smoke area and the facing direction axis is less than 30°. This indicates that the dense smoke is located within the firefighter's horizontal 60° frontal field of vision and is relatively close (the center point's x-coordinate is 2.95m, which is within a dangerous distance of 3 meters), thus obstructing the current escape route. When the escape route indicator (originally pointing straight ahead, now needing to bypass the dense smoke) is activated, the solver immediately generates a path obstruction event signal. This signal is a 3.3V high-level digital signal, accompanied by structured information about the smoke location (helmet coordinates) and obstruction direction (front). This information is synchronously transmitted via the SPI bus to the vector graphics engine and interface simplification engine of the escape route indicator. Upon receiving the signal, the vector graphics engine immediately stops rendering the original frontal path light strip, recalculates the new direction to bypass the smoke, and generates a red flashing arrow (flashing frequency 2 times / second to enhance warning). The interface simplification engine then overlays a semi-transparent red warning text "Smoke Obstruction, Direction Adjusted" below the remaining time display area at the top (font size is half the remaining time font, without obscuring core parameters). If the cosine value is less than or equal to 0.866, it indicates that the smoke is to the side (e.g., a cosine value of 0.7 corresponds to an angle of 45°) or behind, and does not affect the frontal escape route. The solver maintains a low-level output, does not trigger any obstruction events, ensures the escape route indicator works normally, and avoids unnecessary direction adjustments that could interfere with the firefighter's judgment.

[0032] After the spatial relationship solver generates a path obstruction event signal, if only the original escape direction guidance is stopped without timely association of new safety passages, firefighters may easily become disoriented in dense smoke. Therefore, it is necessary to use a path decision-maker to quickly assess the accessibility of all available safety passages, select the optimal path, and convert it into intuitive semantic instructions to ensure that firefighters can instantly understand the escape direction. The specific implementation method is as follows: 首先,路径决策器接收路径阻挡事件信号后,立即调取图像识别输出的安全通道方位集合,路径决策器是集成路径评估与最优选择功能的嵌入式计算模块(基于A RMCortex -A7内核,支持多线程处理,确保复杂场景下响应时间小于等于100毫秒),其通过中断触发机制接收路径阻挡事件信号,信号中包含浓烟阻挡的方位(如正面)、消防员当前位置( 头盔定位模块输出的经纬度,如116.325°E,39.938°N)信息,接收到信号后,决策器立即暂停原路径规划线程,启动安全通道检索线程,安全通道方位集合是图像识别模块( 采用轻量化语义分割模型,对火场画面中的安全出口标识、门体轮廓等特征进行识别)实时输出的结构化数据集合,每个集合元素包含通道唯一ID、通道中心点经纬度、通道宽度( 如0.8米 / 1.2米,反映通行能力)、识别置信度(0-1,大于等于0.7判定为有效通道)、通道状态(如‘畅通’‘半堵塞’,基于轮廓完整性判断),数据存储于设备的临时数据库(采 用Redis内存数据库,支持毫秒级查询)中,更新频率与图像识别频率一致(30帧 / 秒),路径决策器通过数据库接口,按消防员当前位置半径50米的范围检索安全通道,同时过滤 识别置信度小于0.7、状态为半堵塞的无效通道,最终得到可用的安全通道方位集合,为后续可达性评估提供基础数据,随后,通过路径代价计算模型评估各通道的可达性,该模型 综合通道直线距离、路径曲率及浓烟扩散浓度梯度生成代价值,代价值越低,代表通道可达性越强、逃生风险越低,路径代价计算模型是基于火场逃生风险因子构建的多维度评估模型,每个评估维度均通过实战数据校准权重,确保结果贴合实际逃生需求,其一,通道直线距离是消防员当前位置到通道中心点的地面直线距离(基于WGS84坐标系的距离公式计 算,如C1距离25米,C2距离32米,C3距离40米),距离越远,逃生耗时越长、风险越高,因此权重设为0.4,计算时先将距离归一化(如50米对应最大归一化值1,25米对应0.5) ,再乘以权重得到距离代价(C1距离代价0.2,C20.256,C30.32),其二,路径曲率是消防员当前位置到通道的最优路径(基于三维占用栅格地图规划,避开浓烟、障碍物)的弯 曲程度,曲率越大,代表路径需多次转向,易延误逃生时间,权重设为0.3,计算时通过路径上相邻两个转向点的角度差求和(如C1路径仅1次转向,角度差30°,归一化后0.1;C2 路径3次转向,角度差总和90°,归一化后0.3),乘以权重得到曲率代价(C10.03,C20.09,C30.12),其三,浓烟扩散浓度梯度是路径上浓烟浓度的变化速率(基于图像识别 模块输出的浓烟浓度值,如C1路径浓烟浓度从当前位置的0.3mg / m3降至通道处的0.1mg / m3,梯度为-0.008mg / (m3·m);C2路径浓度从0.3升至0.5mg / m3,梯度为0. 006mg / (m3·m)),梯度为正值且越大,代表越靠近通道浓烟越浓,风险越高,权重设为0.3,计算时将浓度梯度归一化(正梯度最大1,负梯度最小0),乘以权重得到浓度代价(C10.0 6 ,C20.18,C30.21)。最终,每个通道的总代价值为三个维度代价之和(C1总代价0.2+0.03+0.06=0.29,C20.256+0.09+0.18=0.526,C30.32+0.12+0.21=0 .65),清晰地量化各通道的可达性差异。

[0033] After obtaining the cost value of each channel, the path decision-maker selects the channel with the lowest cost value as the optimal escape route (e.g., C1 has the lowest cost value of 0.29, so it is determined to be the optimal route). Then, the semantic binding engine converts the direction information of the optimal route into intuitive voice commands. The semantic binding engine is a software module responsible for converting spatial orientation data into natural language semantics. First, the engine reads the optimal path orientation angle data output by the path decision-maker. The orientation angle is the angle between the firefighter's current facing direction (output by the helmet gyroscope, e.g., facing due north, 0°) and the optimal channel direction (the azimuth angle from the firefighter's current position to the center point of channel C1, e.g., 45°). The relative orientation angle is calculated (45°, meaning the firefighter needs to turn 45° to the right towards channel C1). Then, the orientation angle is quantified into eight directional keywords. These eight directional keywords are standardized terms based on geographical orientation (north, northeast, east, southeast, south, southwest, west, northwest), corresponding to the orientation angles. The range includes: North (337.5°-22.5°), Northeast (22.5°-67.5°), and East (67.5°-112.5°), etc. 45° belongs to the Northeast direction, so it is quantified as the keyword Northeast. At the same time, the engine reads the straight-line distance of the optimal channel (25 meters) and generates supplementary information distance of 25 meters. Finally, the engine retrieves the pre-stored voice template and embeds the 25-meter Northeast distance into the {direction}{distance} placeholders in the template to generate a voice text with a complex semantic structure. This text is converted into a voice signal by the text-to-speech (TTS) engine and pushed to the helmet speaker (the volume is automatically adjusted based on the current ambient noise level to ensure penetration). At the same time, a semi-transparent prompt box of this text is superimposed on the simplified engine display area of ​​the interface (located below the escape direction indicator to avoid obscuring the core parameters), realizing dual guidance of voice and vision, ensuring that firefighters can quickly understand and execute escape instructions during emergency movement.

[0034] Based on the preliminary assessment of accessibility of the path cost calculation model by comprehensively considering distance, curvature, and smoke concentration, in order to further adapt to the complex scenarios of dynamic smoke diffusion and changing obstacle distribution in fire scenes, the model needs to achieve refined calculation of cost values ​​through 3D environment modeling, path algorithm optimization, and risk prediction simulation. The core is to first construct a realistic spatial carrier, and then quantify static distance, dynamic risk, and passage difficulty through algorithms and simulators, ultimately generating a cost value that truly reflects the escape risk. The specific implementation method is as follows: First, a 3D environmental topology map is constructed, and the distribution locations of dense smoke identified by image recognition and the orientation of safety passages are mapped as grid nodes. This provides a structured spatial carrier for subsequent path calculation. The 3D environmental topology map is a digital model that abstracts the fire scene operation space (based on the firefighter's current location and a 50-meter radius around it determined by the positioning module, covering a height range from the ground to the ceiling of 10 meters) into a 3D grid structure. Each grid unit (i.e., grid node) corresponds to a 1m × 1m × 1m voxel area in the real space. This size can accurately distinguish between obstacles and passages while avoiding computational delays caused by too many grids. Each node needs to store four types of core information: spatial coordinates (based on the helmet coordinate system, such as (x=5, y=3, z=2)), obstacle markers (0 for no obstacle, 1 for obstacle, generated by walls, equipment, etc. detected by the image recognition module), dense smoke concentration value (a quantized value of 0-1, 0 for no dense smoke, 1 for extremely dense smoke, derived from dense smoke area analysis by image recognition), and passage status (passable, cautious, impassable, determined by a combination of obstacle markers and dense smoke concentration). The construction process is divided into: Determine the map boundary by defining a 50m x 50m x 10m cube centered on the firefighter's current location, ensuring coverage of all possible safety passages and areas with dense smoke. Perform grid division, using an embedded map engine (optimized based on the OctoMap open-source framework and adapted to embedded device computing power) to divide the cube into 25,000 1m x 1m x 1m grid nodes. Each node is assigned a unique ID (e.g., Node_5_3_2 corresponds to coordinates (5,3,2)). Complete coordinate mapping by matching the smoke distribution location (3D coordinates in the helmet coordinate system, e.g., (6,4,2)) output by the image recognition module with the safety passage orientation (e.g., coordinates of the center point of passage C1, e.g., (15,8,1.5)) to the corresponding grid nodes. Nodes corresponding to the smoke locations are then assigned their smoke density values. The density value is set to the concentration output by image recognition (e.g., 0.8). The passage status is marked as "Cautionary Passage". The node corresponding to the safe passage location (passage width 1.2 meters, corresponding to two adjacent nodes (15,8,1.5) and (16,8,1.5)) is marked as the target node, and the passage status is set to "Passable". At the same time, the safe passage ID is recorded in the node attributes to ensure that the passage information can be associated later. After completing the construction of the 3D environment topology map, the shortest path distance from the firefighter's current position to each safe passage target node is calculated using the Dijkstra algorithm. The Dijkstra algorithm is a classic algorithm for finding the shortest path from the starting point to each node in a weighted graph. It is suitable for scenarios in fire scenes where the path weights (such as distance and risk) are non-negative. It can accurately avoid obstacles and nodes with high concentrations of dense smoke and output the actual passable shortest path. In specific implementation: The algorithm's starting and target nodes are determined. The starting node is the grid node corresponding to the firefighter's current position (obtained by matching the coordinates of the helmet positioning module with the map node coordinates, such as (5,3,2)). The target nodes are all nodes marked with safety passage IDs (e.g., C1 corresponds to (15,8,1.5) and (16,8,1.5), C2 corresponds to (20,5,1.5)). Path parameters are initialized by setting the initial distance of all nodes on the map to infinity (indicating unreachable), and the initial distance of the starting node to 0 (distance from itself is 0). A priority queue (based on...) is then used... Nodes are stored in ascending order of current distance (starting from smallest to largest). Initially, only the starting node is added. The shortest path is calculated iteratively. Each time, the node with the smallest distance is taken from the priority queue, and its 6 adjacent nodes (front, back, left, right, up, down, e.g., the adjacent nodes of (5,3,2) are (4,3,2), (6,3,2), (5,2,2), (5,4,2), (5,3,1), (5,3,3)) are traversed. For each adjacent node, if its obstacle is marked as 0 (no obstacle) and the smoke concentration is less than 0.9 (not completely blocked), then the path from the starting point through the current distance is calculated. The distance from a node to its neighboring node is calculated as (current node distance + 1 meter, due to the 1-meter spacing between grid nodes). If this distance is less than the initial distance of the neighboring node, its initial distance is updated, and it is added to a priority queue. This process is repeated until all nodes in the priority queue have been processed. At this point, the initial distance of each safety passage target node is the shortest path distance from the firefighter's current position to that passage (e.g., the shortest path distance for C1 is 25 meters, and for C2 it is 32 meters). The algorithm's execution time is controlled within 50 milliseconds to meet real-time requirements. To prevent the path from becoming impassable due to the spread of dense smoke, [further details are needed]. A dense smoke diffusion simulator is introduced to predict the smoke coverage area within a specified time period and dynamically adjust the path passability coefficient. The simulator is a prediction tool built on a simplified fluid dynamics model (a simplified lattice Boltzmann model adapted to the computing power of embedded devices). It can combine real-time environmental parameters of the fire scene to predict the diffusion trend of dense smoke over a future period. The specified time period is set to 30 seconds based on the firefighters' escape reaction speed and movement speed, predicting which grid nodes the dense smoke will cover in the next 30 seconds, ensuring that the path is not only feasible now but also safe in the next 30 seconds. In specific implementation: The simulator collects input parameters and reads real-time fire data through sensor interfaces, including the current smoke concentration at each grid node (from a 3D environmental topology map), ambient temperature (output from a helmet-mounted temperature sensor, e.g., 50℃), local wind speed (collected by a miniature anemometer, e.g., 0.5 m / s, northeast direction), and combustible material distribution (pre-stored locations of equipment and flammable materials in the map, e.g., cable tray areas marked as high diffusion sources). It then performs diffusion prediction, substituting the input parameters into a simplified diffusion equation and iterating 30 times at 1-second time steps to obtain the predicted smoke concentration for each grid node at each time step. For example, if the current smoke concentration at node (10,6,2) is 0.3, influenced by the northeast wind, it will rise to 0.5 after 10 seconds and to 0.8 after 30 seconds. The simulator also dynamically adjusts the path passability coefficient, an indicator quantifying the overall safety of the path (values ​​from 0 to 1, where 1 is completely safe and 0 is impassable). The calculation method is as follows: Extract all grid nodes contained in the shortest path obtained by the Dijkstra algorithm. First, calculate the current average smoke concentration of these nodes to obtain the initial passability coefficient (e.g., average concentration 0.3, initial coefficient 0.7). Then, combine this with the concentration prediction value after 30 seconds. If the concentration of a node after 30 seconds is greater than or equal to 0.8 (determined as high risk in the future), reduce the coefficient of that node by 0.3 (e.g., from 0.7 to 0.4). Finally, take the average of the corrected coefficients of all nodes as the final passability coefficient of the path (e.g., final passability coefficient of path C1 0.65, path C2 0.4), realizing dynamic adaptation to the future safety of the path. Finally, the final cost of the path cost calculation model is composed of path length weight, passability coefficient weight, and path transformation weight. The path length weight is generated by weighted summation of penalty coefficients, and the output is connected to the optimal path selector. The path length weight is a quantification term based on the shortest path distance. First, the shortest path distance of all passages is normalized (e.g., a normalization value of 1 corresponds to a maximum distance of 50 meters, and 0.5 corresponds to 25 meters), and then multiplied by a preset weight coefficient of 0.4 (this coefficient is calibrated through 100 fire simulation experiments to ensure that the impact of distance on cost is consistent with actual escape risk). For example, the C1 path distance is 25 meters, and after normalization, it is 0.5, so the length weight term is 0.5 × 0.4 = 0.2. The passage coefficient weight is a quantification term based on the final passage coefficient. It is directly multiplied by the weight coefficient of 0.4 (passage safety is crucial for escape, so the weight is comparable to that of distance). The passability coefficient for path C1 is 0.65, so this item is 0.65 × 0.4 = 0.26. The path turning penalty coefficient is a term that quantifies the impact of the number of path turns on escape efficiency. First, the number of turns in the shortest path is counted (a change in direction between adjacent nodes is considered one turn, such as from (5,3,2) to (6,3,2) and then to (6,4,2) is one turn). The number of turns is normalized (the maximum number of turns is 1 for 10 turns and 0.1 for 1 turn). Then, it is multiplied by a weighting coefficient of 0.2 (turning will delay escape time, but the impact is less than that of distance and safety). For example, if path C1 turns once, the normalized value is 0.1, so this item is 0.1 × 0.2 = 0.02. The three factors are added together to get the final cost value (path C1: 0.2 + 0.26 + 0.02). (02=0.48, C2 path is calculated to be 0.75). The output of the model adopts the SPI hardware communication interface to transmit the channel ID and the corresponding data of the final cost value of all safety channels to the optimal path selector in real time. The selector automatically selects the channel with the smallest cost value by comparing all cost values, marks it as the optimal escape path, and outputs the direction angle, distance and other information of the channel at the same time, providing a basis for subsequent semantic binding and directional guidance. The path cost calculation model can accurately quantify the comprehensive risk of fire escape path, avoid ignoring the risk of dynamic dense smoke due to static distance misjudgment, and ensure the convenience of the path through turning penalty. Finally, it provides a reliable quantitative basis for the selection of the optimal path and adapts to the escape decision-making needs in complex fire environment.

[0035] In this invention, the data acquisition unit obtains environmental noise intensity, firefighter displacement rate, smoke distribution location, safety passage orientation, and remaining time for the air respirator through the helmet device. When the high-priority alarm is triggered with the remaining time, the multimodal coupling processing unit dynamically adjusts the voice volume according to the noise intensity, activates the simplified interface based on the displacement rate, determines the frontal obstruction of smoke through spatial calculation, plans the optimal passage by combining the path cost model, generates a composite semantic command that binds the remaining time and orientation keywords, integrates the alarm unit to synchronously broadcast the voice command and output the simplified interface, sets up a touch hotspot for the simplified interface, supports freeze refresh and directional indicator flashing, and improves the efficiency and safety of fire rescue emergency decision-making.

[0036] The second objective of this invention is to provide a method for implementing a fire rescue intelligent auxiliary system based on multimodal perception, including any of the above-mentioned features, comprising the following steps: S1. Real-time capture of ambient sound wave signals through helmet microphone and conversion into noise intensity value; continuous acquisition of firefighter displacement rate through positioning module; acquisition of on-site images through helmet camera and identification of smoke distribution location and safety passage location by pre-trained model; and monitoring of remaining time through breathing apparatus sensor. S2. When a high-priority alarm is triggered by the remaining time, the noise intensity value is input to the threshold comparator and compared with the preset threshold in real time. If it exceeds the threshold, the compensation mechanism is activated. The compensation amount is matched by a lookup table and the programmable amplifier is driven to increase the amplitude of the voice signal to the target level. At the same time, the first movement status flag is activated based on the displacement rate change gradient and mean, triggering the video memory control module to close the non-emergency parameter layer. The font engine is called to enlarge the remaining time value to the top center of the interface according to the preset ratio. At the same time, the escape direction indicator is superimposed on the preset coordinate position of the camera video stream by the augmented reality fusion module as a transparent vector layer. The gyroscope orientation data and the coordinates of the smoke distribution position are spatially calculated. The frontal obstruction event is determined by the cosine angle model. The optimal channel cost is generated by the path cost calculation model, which integrates distance, curvature and smoke concentration gradient. Eight-directional keywords are output and bound to the remaining time to generate a composite semantic structure. S3. Broadcast voice commands with complex semantic structures at the compensated target volume, and simultaneously output a simplified display interface. The voice content integrates the remaining time and location keywords, and the interface highlights and enlarges the remaining time and dynamic direction indicators. S4. Set a touch hotspot in the simplified interface. Respond to a single click operation to freeze the interface refresh and activate the directional indicator pulse flashing. After a specified duration, the data update will be automatically restored.

[0037] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A fire rescue intelligent auxiliary system based on multimodal perception, characterized in that, include: The data acquisition unit (1) is used to acquire the ambient noise intensity value through the helmet microphone, acquire the firefighter displacement rate through the positioning module, acquire on-site images through the helmet camera and identify the distribution location of dense smoke and the location of the safety passage, and acquire the remaining time of the air respirator through the sensor. The multimodal coupling processing unit (2) performs the following operations when a high-priority alarm is triggered during the remaining time of the air respirator: First, the voice output volume is adjusted based on the ambient noise intensity value. This includes comparing the current noise intensity value with a preset noise threshold. If the current noise intensity value exceeds the preset noise threshold, a target volume equal to the current noise intensity value plus a compensation amount is generated. The current movement state is determined based on the firefighter's displacement rate. When in the first movement state, the interface simplification engine is activated. The interface simplification engine hides non-emergency parameter data fields, retaining only the remaining time value and escape direction indicator. The remaining time value is fixed in a magnified form in the top center area of ​​the display interface. At the same time, the escape direction indicator is superimposed on the preset coordinate position of the real-time camera image. Finally, directional escape instructions are generated based on the location of the dense smoke distribution and the direction of the safety passage. This includes analyzing the relative relationship between the location of the dense smoke distribution and the direction the firefighter is facing. If it is determined that there is a frontal obstruction, the nearest safety passage is automatically associated, and the direction keyword is bound to the remaining time value to generate a compound semantic structure. The integrated alarm unit (3) broadcasts voice content containing complex semantic structures at the target volume and simultaneously outputs the display interface simplified by the interface simplification engine.

2. The intelligent auxiliary system for fire rescue based on multimodal perception according to claim 1, characterized in that: The operation of comparing the current noise intensity value with the preset noise threshold specifically includes: The raw sound wave signal collected by the helmet microphone is processed in frames by a real-time audio processing circuit. Each frame of the signal is converted into a spectral energy distribution by a fast Fourier transform, and then the ambient noise intensity value is calculated by weighted integration. The ambient noise intensity value is input into a threshold comparator and differentially calculated with a preset noise threshold. When the difference value is positive, a voice compensation flag signal is triggered. The threshold comparator uses hardware logic gate circuits to realize real-time judgment, and the output terminal is connected to the enable port of the target volume generation module.

3. The intelligent auxiliary system for fire rescue based on multimodal perception according to claim 1, characterized in that, The operation of generating a target volume equal to the current noise intensity value plus a compensation amount specifically includes: When the threshold comparator outputs a voice compensation flag signal, the digital gain controller is activated. The digital gain controller reads the current noise intensity value and calls the compensation lookup table. The compensation lookup table stores fixed compensation values ​​corresponding to different noise intensity ranges. The fixed compensation values ​​are loaded onto the gain control pin of the programmable amplifier through the digital-to-analog converter, so that the amplitude of the alarm voice signal output by the voice synthesis chip is increased to the compensated level. At the same time, the amplitude limiting circuit is used to prevent signal overload distortion, and finally the speaker is driven to output penetrating alarm voice.

4. The intelligent auxiliary system for fire rescue based on multimodal perception according to claim 1, characterized in that: The operation of determining the current movement status based on the firefighter's displacement rate specifically includes: The displacement data stream output by the positioning module is input to the motion state classifier. The motion state classifier calculates the displacement rate change gradient through a sliding time window. When the displacement rate change gradient exceeds the dynamic threshold and the average rate is greater than the state judgment benchmark value within three consecutive time windows, the first motion state flag is activated. The start signal of the interface simplification engine is electrically interlocked with the flag signal to ensure that the simplified interface is forcibly enabled in the rapid movement state.

5. The intelligent auxiliary system for fire rescue based on multimodal perception according to claim 4, characterized in that: The operation of hiding non-urgent parameter data fields in the interface simplification engine specifically includes: A priority mapping database is established, marking the remaining time of the air respirator and the escape direction indicator as the highest priority parameters, and the battery voltage and altitude data as the lowest priority parameters. When the first movement status flag is activated, the graphics rendering interface is called to close the memory address of the display layer corresponding to the lowest priority parameter, while the highest priority parameter is written to the video memory buffer. The escape direction indicator is dynamically generated by a vector graphics engine, and its shape is adaptively switched to a directional arrow or a path light strip according to the orientation of the safety passage.

6. The intelligent auxiliary system for fire rescue based on multimodal perception according to claim 5, characterized in that: The operation of fixing the remaining time value in a magnified form at the top center area of ​​the display interface specifically includes: The font scaling engine is called to parse the character width of the remaining time value, calculate the maximum readable size according to the screen resolution, generate an enlarged font with the rule of occupying 30% of the top area of ​​the interface, and at the same time enable the anti-aliasing rendering pipeline to eliminate pixelated edges. The operation of overlaying the escape direction indicator onto the real-time camera footage is achieved through an augmented reality fusion module, which includes acquiring frame buffer data of the camera video stream, establishing a transparent layer at preset coordinates, overlaying the vector direction indicator onto the video stream in alpha blending mode, and adjusting the overlay angle in real time by matching the camera attitude data.

7. The intelligent auxiliary system for fire rescue based on multimodal perception according to claim 1, characterized in that: The operation of analyzing the relative relationship between the distribution location of dense smoke and the direction the firefighters are facing specifically includes: The orientation data output from the helmet gyroscope is input into the spatial relationship solver. The solver transforms the coordinates of the smoke distribution location obtained from image recognition to the helmet coordinate system, calculates the angle between the center point of the smoke area and the facing direction axis, and establishes an obstruction determination model based on the cosine value of the angle. When the cosine value is greater than the orientation correlation threshold, it is determined that the smoke is located in the frontal field of view and a path obstruction event signal is generated.

8. The intelligent auxiliary system for fire rescue based on multimodal perception according to claim 7, characterized in that: The operation of automatically associating the location of the nearest safe passage specifically includes: After receiving the path obstruction event signal, the path decision-maker retrieves the set of safe passage locations output by image recognition, evaluates the accessibility of each passage through the path cost calculation model, and generates a cost value by combining the straight-line distance of the passage, the curvature of the path, and the gradient of the smoke diffusion concentration. The passage with the lowest cost value is selected as the optimal escape path. The semantic binding engine reads the direction angle data of the optimal path, quantifies it into eight-directional keywords, and embeds the directional placeholders of the voice template to generate a composite semantic structure.

9. The intelligent auxiliary system for fire rescue based on multimodal perception according to claim 8, characterized in that: The operation of the path cost calculation model further includes: A 3D environmental topology map is constructed, mapping the distribution location of dense smoke identified in the image to the orientation of safety passages as grid nodes. The shortest path distance from the firefighter's current location to each passage node is calculated using the Dijkstra algorithm. A dense smoke diffusion simulator is introduced to predict the smoke coverage range within a specified time period in the future. The path accessibility coefficient is dynamically adjusted. The final cost is generated by weighted summation of path length weight, accessibility coefficient weight, and path turning penalty coefficient. The output is connected to the optimal path selector.

10. A method for implementing the intelligent auxiliary system for fire rescue based on multimodal perception as described in any one of claims 1-9, characterized in that: Includes the following steps: S1. Real-time capture of ambient sound wave signals through helmet microphone and conversion into noise intensity value; continuous acquisition of firefighter displacement rate through positioning module; acquisition of on-site images through helmet camera and identification of smoke distribution location and safety passage location by pre-trained model; and monitoring of remaining time through breathing apparatus sensor. S2. When a high-priority alarm is triggered by the remaining time, the noise intensity value is input to the threshold comparator and compared with the preset threshold in real time. If it exceeds the threshold, the compensation mechanism is activated. The compensation amount is matched by a lookup table and the programmable amplifier is driven to increase the amplitude of the voice signal to the target level. At the same time, the first movement status flag is activated based on the displacement rate change gradient and mean, triggering the video memory control module to close the non-emergency parameter layer. The font engine is called to enlarge the remaining time value to the top center of the interface according to the preset ratio. At the same time, the escape direction indicator is superimposed on the preset coordinate position of the camera video stream by the augmented reality fusion module as a transparent vector layer. The gyroscope orientation data and the coordinates of the smoke distribution position are spatially calculated. The frontal obstruction event is determined by the cosine angle model. The optimal channel cost is generated by the path cost calculation model, which integrates distance, curvature and smoke concentration gradient. Eight-directional keywords are output and bound to the remaining time to generate a composite semantic structure. S3. Broadcast voice commands with complex semantic structures at the compensated target volume, and simultaneously output a simplified display interface. The voice content integrates the remaining time and location keywords, and the interface highlights and enlarges the remaining time and dynamic direction indicators. S4. Set a touch hotspot in the simplified interface. Respond to a single click operation to freeze the interface refresh and activate the directional indicator pulse flashing. After a specified duration, the data update will be automatically restored.