A fire accident scene audio multi-dimensional analysis-based investigation and evaluation technical method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TIANJIN FIRE SCI & TECH RES INST OF MEM
- Filing Date
- 2026-05-12
- Publication Date
- 2026-08-04
AI Technical Summary
但是现有的火灾调查音频分析一方面主要依赖分析者的辨听,主观性强;一方面是要通过音频的波形图和频谱图做比对分析,但这种方法对于一些相似的声音无法准确地进行区分和辨识
[0012]As can be seen from the above technical solution, compared with the prior art, this invention provides an investigation and assessment method based on multi-dimensional audio analysis of fire accident scenes. It changes the traditional model that relies on subjective listening by the analyst, simple comparison analysis through audio spectrograms, or direct audio comparison using neural network algorithms. Instead, it uses a multi-dimensional parameter fusion comparison analysis method, effectively improving the accuracy of fire audio recognition. The sounds present at fire scenes are exceptionally complex, and the possibility of similar sounds is also very high. Simply relying on listening, waveform comparison, spectrum comparison, or directly using deep learning algorithms to identify sounds cannot meet the needs of sound recognition at complex fire scenes. In particular, identification errors can mislead the direction of fire investigations and affect the determination of the cause of the case. For example, at arson fire scenes, there are often sounds of ignition devices being used. However, there are many types of common ignition devices, including windproof lighters, roller lighters, blowtorch lighters, piezoelectric ceramic lighters, matches, etc. In particular, the audio similarity between windproof lighters and blowtorch lighters is extremely high, and traditional analysis methods cannot accurately determine which ignition device was used. The audio identification technology method invented in this invention effectively distinguishes the sounds of ignition from several types of ignition devices, accurately determining the type of ignition device. Electrical fires are often accompanied by characteristics of electrical faults, such as the sounds of short circuits in electrical circuits or thermal runaway of lithium batteries. When the angle of the monitoring video is limited, it is impossible to accurately identify the cause of the fire. The audio analysis technology method invented in this invention can effectively distinguish what kind of fault sound caused the fire.
Smart Images

Figure CN122511294A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fire investigation and assessment technology, and more specifically to an investigation and assessment method based on multidimensional audio analysis of fire accident scenes. Background Technology
[0002] Surveillance equipment now covers the vast majority of places where people engage in production activities, and surveillance video has become one of the key pieces of evidence in fire investigations. Traditional surveillance video applications focus on analyzing the location and cause of a fire through visual information. With the widespread availability of surveillance equipment with audio acquisition capabilities, surveillance audio can also capture some key sound information, such as explosions, footsteps, and the sound of a lighter. If these sounds can be distinguished, they can provide important guidance for fire investigations. However, existing fire investigation audio analysis relies heavily on the analyst's listening skills, which is highly subjective, and it also involves comparing and analyzing audio waveforms and spectrograms, but this method cannot accurately distinguish and identify some similar sounds.
[0003] Therefore, how to solve the above-mentioned technical problems still requires further research by those skilled in the art. Summary of the Invention
[0004] In view of this, the present invention provides an investigation and assessment technique based on multidimensional analysis of audio at fire scene. This method compares audio signals at the fire scene through comprehensive analysis of multidimensional characteristic parameters, including time-domain characteristic parameters (mean, variance, kurtosis, short-time energy, zero-crossing rate, crest factor, and waveform factor), frequency-domain characteristic parameters (spectral centroid, spectral roll-off, and spectral flatness), and time-frequency-domain characteristic parameters (energy entropy, time-frequency correlation mean, and energy variance). This enables accurate identification of audio at the fire scene and provides technical support for determining the cause of the fire investigation.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: A survey and assessment technique based on multidimensional audio analysis of fire accident scenes includes the following steps: Extract the target audio signal from the fire scene surveillance video and convert it into a digital audio signal; The digital audio signal is segmented into frames to obtain multiple short-time audio signals. The time-domain feature parameters, frequency-domain feature parameters, and time-frequency-domain feature parameters of each frame of short audio signal are extracted respectively; The weights of each feature parameter are determined based on the CNN convolutional neural network algorithm; A grouped weighted matching strategy is used to calculate the similarity between the audio to be tested and various types of audio in the reference audio library, and the fire audio type is determined based on the highest similarity value.
[0006] Optionally, time-domain feature parameters of each frame of short-time audio signal can be extracted, including mean, variance, kurtosis, short-time energy, zero-crossing rate, crest factor, and waveform factor.
[0007] Optionally, the frequency domain characteristic parameters are determined as follows: The centroid of a spectrum is used to characterize the centroid frequency of the spectral energy distribution; Spectral roll-off is used to select the frequency point where the cumulative energy accounts for 95% of the total energy; Spectral flatness is the ratio of the geometric mean to the arithmetic mean of the spectrum, characterizing the smoothness of the spectral distribution.
[0008] Optionally, the time-frequency domain feature parameters are calculated based on the time-frequency energy matrix of continuous wavelet transform; where energy entropy is used to measure the complexity of signal energy distribution in each sub-time window; energy variance is used to measure the degree of fluctuation of short-time energy in different time windows; and the time-frequency correlation mean is used to calculate the correlation coefficient between the time axis and the frequency axis, characterizing the time-varying correlation of voiceprints.
[0009] Optionally, a group-weighted matching strategy is used to calculate the similarity between the audio to be tested and various types of audio in the reference audio library. Specifically, the multidimensional feature parameters are divided into time-domain feature groups, frequency-domain feature groups, and time-frequency-domain feature groups. At least one feature in each group must be successfully matched for the weight of that group to be included in the total score. The tolerance threshold for each feature is set according to experimental statistics to cover 99.7% of the normal fluctuation range, and the similarity score is calculated in combination with the weights.
[0010] Optionally, for ignition-related fire audio, a multi-level feature combination can be used for differentiation: first, kurtosis and variance are used to differentiate the audio of a spray gun ignition; then, zero-crossing rate and peak factor are used to differentiate the audio of a windproof lighter ignition; then, spectral centroid and spectral flatness are used to differentiate the audio of a match ignition; finally, energy entropy is used to differentiate the audio of a piezoelectric ceramic lighter from that of a roller lighter.
[0011] Optionally, for the audio of electrical short circuit fires, a combination of features is used for judgment: zero-crossing rate is used to distinguish between AC and DC short circuits; short-time energy, variance, and spectral centroid are used to determine the voltage level, wire diameter, and material of AC short circuits; and energy entropy and time-frequency correlation mean are used to determine the voltage level of DC short circuits.
[0012] As can be seen from the above technical solution, compared with the prior art, this invention provides an investigation and assessment method based on multi-dimensional audio analysis of fire accident scenes. It changes the traditional model that relies on subjective listening by the analyst, simple comparison analysis through audio spectrograms, or direct audio comparison using neural network algorithms. Instead, it uses a multi-dimensional parameter fusion comparison analysis method, effectively improving the accuracy of fire audio recognition. The sounds present at fire scenes are exceptionally complex, and the possibility of similar sounds is also very high. Simply relying on listening, waveform comparison, spectrum comparison, or directly using deep learning algorithms to identify sounds cannot meet the needs of sound recognition at complex fire scenes. In particular, identification errors can mislead the direction of fire investigations and affect the determination of the cause of the case. For example, at arson fire scenes, there are often sounds of ignition devices being used. However, there are many types of common ignition devices, including windproof lighters, roller lighters, blowtorch lighters, piezoelectric ceramic lighters, matches, etc. In particular, the audio similarity between windproof lighters and blowtorch lighters is extremely high, and traditional analysis methods cannot accurately determine which ignition device was used. The audio identification technology method invented in this invention effectively distinguishes the sounds of ignition from several types of ignition devices, accurately determining the type of ignition device. Electrical fires are often accompanied by characteristics of electrical faults, such as the sounds of short circuits in electrical circuits or thermal runaway of lithium batteries. When the angle of the monitoring video is limited, it is impossible to accurately identify the cause of the fire. The audio analysis technology method invented in this invention can effectively distinguish what kind of fault sound caused the fire. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0014] Figure 1 This is a schematic diagram of the fire audio comparison process of the present invention; Figure 2 A comparison chart of the average values for different types of ignition devices; Figure 3 A comparison chart of variances for different ignition devices; Figure 4 A comparison chart of zero-crossing rates for different ignition devices; Figure 5 A comparison chart of peak density for different ignition devices; Figure 6 A short-time energy comparison chart of different ignition devices; Figure 7 A comparison chart of peak factors for different ignition devices; Figure 8 A comparison chart of waveform factors for different ignition devices; Figure 9 Comparison of spectral centroids for different ignition devices; Figure 10 A comparison chart of roll-off values for different ignition devices; Figure 11 A comparison chart of spectrum flatness for different ignition devices; Figure 12 A comparison chart of energy entropy for different ignition devices; Figure 13 A comparison chart of the mean time-frequency correlation values for different ignition devices; Figure 14 A graph comparing the energy variance of different ignition devices. Detailed Implementation
[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] This invention discloses a method for identifying fire audio based on multidimensional feature parameters, comprising the following steps: Extract the target audio signal from the fire scene surveillance video and convert it into a digital audio signal; The digital audio signal is segmented into frames to obtain multiple short-time audio signals. The time-domain feature parameters, frequency-domain feature parameters, and time-frequency-domain feature parameters of each frame of short audio signal are extracted respectively; The weights of each feature parameter are determined based on the CNN convolutional neural network algorithm; A grouped weighted matching strategy is used to calculate the similarity between the audio to be tested and various types of audio in the reference audio library, and the fire audio type is determined based on the highest similarity value.
[0017] Audio analysis was performed on surveillance video extracted from the fire scene, such as... Figure 1 As shown, the specific steps are as follows: Step 1: After extracting the video of the fire time period from the surveillance video, use professional audio and video processing tools (such as Format Factory, CapCut, WPS video editing software, etc.) to separate the audio of the target time period. Generally, the separated audio is a digital signal. If it is an analog signal, it can be converted into a digital signal through an analog-to-digital converter. Step 2: Frame the digital signal, that is, divide the digital audio into segments of 20-30ms in length, with each segment representing a frame. The purpose of framing is to divide long audio into multiple short segments (i.e., "frames") so that operations such as spectrum analysis and feature extraction can be performed on each frame.
[0018] Step 3: Analyze the time-domain characteristics of the extracted audio digital signal, mainly including seven parameters: mean, variance, short-time energy, zero-crossing rate, crest factor, waveform factor, and kurtosis. The calculation methods for each parameter are as follows: ① Mean: Used to characterize the average amplitude of a signal within a sampling interval. The feature extraction step calculates the time-domain mean of each frame by first dividing the signal into frames, and then calculating the time-domain mean of each frame using the following formula:
[0019] Where, μ i x represents the time-domain mean of the signal in the i-th frame; i (n) represents the amplitude of the nth sampling point in the i-th frame signal; N represents the number of sampling points in a single frame signal.
[0020] ② Variance: Used to measure the degree of fluctuation in signal amplitude, that is, the deviation between the observed signal value and its mean. In the feature extraction step, when calculating the time-domain variance of each frame of signal, the degree of fluctuation is calculated based on the mean of each frame of signal using the following formula:
[0021] in, x represents the time-domain variance of the signal in the i-th frame; i (n) represents the amplitude of the nth sampling point in the i-th frame of the signal; μ i represents the time-domain mean of the i-th frame signal; N represents the number of sampling points in a single frame signal.
[0022] ③ Short-time energy: Used to quantify the ratio between the signal peak value and its effective value (RMS), characterizing the intensity of sudden spikes in the signal. In the feature extraction step, when calculating the short-time energy of each frame of the signal, it is calculated using the following formula:
[0023] Among them, E i x represents the short-time energy of the signal in the i-th frame; i (n) represents the amplitude of the nth sampling point in the i-th frame signal; N represents the number of sampling points in a single frame signal.
[0024] ④ Zero-crossing rate: This reflects the fluctuation frequency and non-periodic components of the audio signal by counting the number of times the signal waveform crosses zero points per unit time. In the feature extraction step, when calculating the zero-crossing rate of each frame, the number of times the signal crosses the zero level is counted, and then calculated using the following formula:
[0025] Among them, ZCR i x represents the zero-crossing rate of the signal in the i-th frame; i (n) represents the amplitude of the nth sampling point in the i-th frame signal; sgn() is the sign function; N represents the number of sampling points in a single frame signal.
[0026] ⑤ Crest Factor: Used to quantify the ratio between the peak value and the effective value (RMS) of a signal, characterizing the intensity of sudden spikes in the signal. In the feature extraction step, when calculating the crest factor for each frame of the signal, the effective value of the signal is calculated first, and then the following formula is used:
[0027] Among them, CF i x represents the peak value factor of the signal in the i-th frame; i (n) represents the amplitude of the nth sampling point in the i-th frame signal; N represents the number of sampling points in a single frame signal.
[0028] ⑥ Waveform Factor: This feature is defined as the ratio of the effective value of the signal to the average amplitude, used to describe the overall contour features of the signal waveform. In the feature extraction step, when calculating the waveform factor for each frame of the signal, it is calculated using the following formula:
[0029] in, SF i Indicates the first i Waveform factor of frame signal; x i ( n ) indicates the first i The first frame signal n The amplitude of each sampling point; N This indicates the number of sampling points in a single frame of signal.
[0030] ⑦ Kurtosis: Used to describe the steepness of the signal probability distribution, reflecting the sharpness of the data distribution. In the feature extraction step, when calculating the kurtosis of each frame of signal, the steepness of the signal distribution is measured by the following formula:
[0031] in, Kurt i express i Kubularity of the frame signal; xi ( n ) indicates the first i The first frame signal n The amplitude of each sampling point; μ i Indicates the first i The time-domain mean of the frame signal; N This indicates the number of sampling points in a single frame of signal.
[0032] Step four involves analyzing the frequency domain features of the extracted audio. Frequency domain features primarily focus on the frequency components of the signal, and are analyzed by converting the time-domain digital signal to the frequency domain using Fourier Transform (FFT). Frequency domain features reveal the frequency characteristics of the signal and identify the frequency features of the voiceprint. Key indicators in frequency domain features include spectral centroid, spectral roll-off, and spectral flatness.
[0033] ① Spectral centroid: The "center of gravity" frequency describing the distribution of spectral energy, measuring the degree of shift in the signal's spectral center. In the feature extraction step, when calculating the spectral centroid of each frame of the signal, a short-time Fourier transform is first performed on the signal, and then the following formula is used to calculate it:
[0034] in, SC i Indicates the first i The spectral centroid of the frame signal; f k Indicates the first k The frequencies corresponding to the spectral lines; S i ( k ) indicates the first i After the short-time Fourier transform of the frame signal, the first k The energy of each spectral line; K This indicates the total number of spectral lines in the spectrum.
[0035] ② Spectral roll-off: This feature is defined as the ratio of the effective value of the signal to the average amplitude, used to describe the overall contour features of the signal waveform. In the feature extraction step, when calculating the spectral roll-off for each frame of the signal, the frequency point where the cumulative energy accounts for 95% of the total energy is found, and then determined using the following formula:
[0036] in, K roll Spectral line index to satisfy cumulative energy conditions ; S i ( k ) indicates the first i After the short-time Fourier transform of the frame signal, the first k Energy of spectral lines , , , w ( n ) is a window function; K This indicates the total number of spectral lines in the spectrum.
[0037] ③ Spectral flatness: By evaluating the smoothness of the spectral distribution, it is determined whether the signal energy is concentrated at a specific resonance peak or uniformly distributed like white noise. In the feature extraction step, when calculating the spectral flatness of each frame of the signal, the ratio of the geometric mean to the arithmetic mean of the spectrum is obtained using the following formula:
[0038] Among them, SFM i S represents the spectral flatness of the signal in the i-th frame; i (k) represents the energy of the k-th spectral line after the short-time Fourier transform of the i-th frame signal; K represents the total number of spectral lines; ε is to prevent the occurrence of a minimum value of 0 when taking the logarithm.
[0039] Step 5: Analyze the time-frequency domain features of the extracted audio digital signal. Time-frequency domain features combine the advantages of the time and frequency domains, making them suitable for analyzing time-varying signals, especially non-stationary signals such as speech signals, as they can reflect frequency components that change over time. Time-frequency domain features mainly include parameters such as energy entropy and energy variance.
[0040] ① Energy entropy: Used to measure the complexity and uncertainty of signal energy distribution within each sub-time window. In the feature extraction step, when calculating the time-frequency energy entropy of the signal, the signal is first subjected to continuous wavelet transform to obtain the time-frequency energy matrix, and then calculated using the following formula:
[0041] Where H represents the time-frequency energy entropy of the signal; pj,k=C(j,k) / ∑C(j,k) is the energy proportion of the j-th frequency and k-th time point in the time-frequency matrix; C(j,k) is the time-frequency energy matrix, where each element represents the energy of the signal at the i-th frequency scale and k-th time point. x(n) is the original audio signal sequence, and resample() is the resampling function that converts the signal from... f s Convert to , Let ψ(t) be a continuous wavelet transform with Morlet wavelet as the mother wavelet; i is the frequency scale index; k is the time index; M is the number of frequency dimension points; L is the number of time dimension points; ε is used to prevent the occurrence of a minimum value of 0 when taking the logarithm.
[0042] ② Energy Variance: Measures the degree of short-term energy fluctuation in an audio signal between different time windows. In the feature extraction step, when calculating the time-frequency energy variance of the signal, it is calculated based on the time-frequency energy matrix obtained from continuous wavelet transform using the following formula:
[0043] Among them, Var tf Represents the time-frequency energy variance of the signal; C(:) is the vector formed by flattening the time-frequency energy matrix C into a one-dimensional vector; Var( ) is the variance calculation function.
[0044] ③ Time-frequency correlation mean: By calculating the correlation coefficient of the signal on the time axis and frequency axis, the degree of correlation of the voiceprint features in the time-varying process is described. When calculating the time-frequency correlation mean of the signal in the feature extraction step, the time-frequency energy matrix obtained based on the continuous wavelet transform is calculated by the following formula:
[0045] in, Corr tf This represents the mean of the time-frequency correlation of the signal; C Time-frequency energy matrix , x ( n The original audio signal sequence is represented by , and resample() is the resampling function that converts the signal from . f s Convert to , Using Morlet wavelets as ψ ( t ) Continuous wavelet transformation of the mother wave; corr( ) is the correlation coefficient calculation function , The first element in the transpose matrix t line, number p The time-frequency energy value of the column, For the transpose matrix, the first... p The mean of a column; mean( ) is the mean function.
[0046] Step Six: Perform a comprehensive comparative analysis using the thirteen feature parameters mentioned above. Single feature comparison has limitations and is difficult to distinguish fire audio types in complex environments. Combining the thirteen features in different ways can effectively solve the problem of fire audio recognition.
[0047] In arson and fires caused by embers such as cigarette butts, determining the presence of ignition is crucial evidence in identifying the cause of the fire. Common ignition tools include matches, piezoelectric lighters, windproof lighters, blowtorches, and roller lighters. Single-domain feature analysis cannot determine the type of ignition tool used by the suspect, as the means of the five types of ignition tools show a high degree of overlap. Matches, roller lighters, and piezoelectric lighters are very similar in characteristics such as variance, zero-crossing rate, kurtosis, and short-time energy. However, a combination of several features can effectively distinguish the type of ignition tool.
[0048] ① First, the spray gun can be effectively distinguished from other ignition tools by its "kurtosis" and "variance". The "kurtosis" and "variance" of the spray gun are significantly higher than those of other types of ignition tools. ② Furthermore, the "zero crossing rate" and "crest factor" can be used to distinguish windproof lighters from other ignition tools. The "crest factor" and "zero crossing rate" of windproof lighters are significantly higher than those of matches, piezoelectric ceramic lighters, and roller lighters. ③ The “spectral centroid” and “spectral flatness” can be used to distinguish matches from piezoelectric ceramic lighters and roller lighters. When a match is lit, the “spectral centroid” and “spectral flatness” of the audio signal are significantly lower than those of piezoelectric ceramic lighters and roller lighters.
[0049] ④ The energy entropy can be used to distinguish between piezoelectric ceramic lighters and roller lighters. The energy entropy of roller lighters is significantly higher than that of piezoelectric ceramic lighters.
[0050] For example, electrical short circuits are the most common cause of fires. When an electrical short circuit occurs, different materials, wire diameters, voltages, etc., all affect the audio signal produced.
[0051] ① Use "zero-crossing rate" as the main parameter to distinguish the form of fault current. If the zero-crossing rate is significantly >0 (e.g., 10), - If the variance is on the order of 3, it is AC; if the zero-crossing rate is approximately 0, it is DC. Using variance and short-time energy as auxiliary verification methods, the variance and energy of DC signals are lower than those of AC signals and have a very narrow distribution.
[0052] ② When the fault current is determined to be AC, the fault voltage level can be judged by the "short-time energy". The short-time energy value for 220V is 1.66 × 10⁻⁶. -4 It is nearly 3 times higher than the 100V / 150V group (0.52×10). -4 ).
[0053] The wire diameter can then be differentiated using "variance" and "spectral centroid". A 1.5 mm² wire diameter not only has high short-time energy but also significantly higher variance. Other thicker wire diameters (2.5~10 mm²)... 2The energy and variance are not significantly different, so the "spectral centroid" can be used as an auxiliary indicator. The finer the line diameter, the higher the spectral centroid, and the sharper the sound.
[0054] ③ When the fault current is determined to be DC, the voltage level can be judged by the "energy entropy" and "time-frequency correlation mean". The 24V DC entropy value is the highest (13.01), and the energy entropy decreases as the voltage increases. As the voltage increases, the time-frequency correlation mean increases (from 0.15 to 0.27), which can be used as an auxiliary indicator to judge the voltage level.
[0055] ④ Determine the material of the short-time circuit by using "short-time energy" and "spectral centroid". The average energy of aluminum wire is generally slightly higher than that of copper wire of the same specification (e.g., 4 square millimeters aluminum has an energy of 2.51 × 10⁻⁶). -4 Copper 4 square meters 1.62 × 10 -4 The spectral centroid of copper wire is usually higher than that of aluminum wire (e.g., 2203 Hz for 10 square millimeters of copper compared to 1670 Hz for 10 square millimeters of aluminum), indicating that copper wire has a richer high-frequency component.
[0056] Step 7: Using the CNN convolutional neural network algorithm, comprehensively analyze and compare the time-domain features, frequency-domain features, and time-frequency-domain features of 1000 samples in the dataset to determine the weights of different feature parameters. The weights of each feature parameter in four scenarios—lighter, battery, short circuit, and air switch—are as follows (if the application scenario changes, the weights of different feature parameters in the dataset can be recalculated using the neural network algorithm): mean 0.5, variance 3.0, kurtosis 2.0, short-time energy 3.0, zero-crossing rate 2.0, peak factor 2.0, waveform factor 0.5, spectral centroid 3.0, spectral roll-off 3.0, spectral flatness 3.0, energy entropy 3.0, time-frequency correlation mean 1.0, energy variance 3.0.
[0057] Step 8: Determine which audio files the test audio has a higher similarity to using similarity evaluation methods.
[0058] Traditional methods that rely solely on weighted summation of global similarity may suffer from abrupt changes in values for a single dimension (such as temporal energy) due to environmental noise, thus rejecting the correct audio type outright. The grouped forced matching rule ensures that the audio to be tested must be consistent with the reference audio in three physical dimensions: temporal profile, frequency structure, and dynamic evolution, avoiding incorrect exclusions caused by anomalies in a single dimension.
[0059] ① Let the feature vector of the audio to be identified be:
[0060] The feature vector of a certain type of reference audio in the database is:
[0061] Each feature To correspond to a preset matching tolerance (an absolute or relative threshold, derived from experimental statistics), define the feature matching function:
[0062] in or The range of experimental distribution for each feature is determined (e.g., the mean feature is set to 10). -4 The order of magnitude, with a zero-crossing rate of 10. -3 (e.g., magnitude).
[0063] The threshold was determined based on extensive experimental statistical data, specifically as follows: Through statistical analysis of over 1000 samples from four typical scenarios (ignition, lithium battery, short circuit, and air switch), the mean and standard deviation of various characteristic parameters were calculated. Threshold or The mean feature is typically set within the confidence interval of the reference sample characteristics to ensure coverage of 99.7% of the normal fluctuation range; therefore, a value of 10 is used for the mean feature. -4 The order of magnitude, with a zero-crossing rate of 10. -3 Magnitude.
[0064] Differentiated tolerance levels: Different levels of tolerance are set based on the sensitivity of different characteristics to fire sound signatures. For example: the average tolerance level is 10. -4 The fluctuations are extremely small, therefore a very small absolute tolerance is set to strictly control the consistency of the signal baseline. Zero-crossing rate / short-time energy: on the order of 10. -3 Up to 10-4, the propagation path has a slightly greater impact, so the tolerance should be appropriately relaxed to take into account the slight distortion of the same sound source at different recording distances.
[0065] Energy entropy / spectral flatness: It has strong randomness and adopts a relative deviation threshold to adapt to the normal fluctuation of complexity under different signal-to-noise ratio environments.
[0066] ② Divide the 13 features into three groups: Time-domain feature set Gtime={1,2,3,4,5,6,7} Frequency domain feature set Gfreq={8,9,10} Time-frequency domain feature set Gtf={11,12,13} Each group must have at least one feature match; otherwise, all matching scores for that group will be zero.
[0067] ③ For the audio to be identified and a certain reference type, calculate the similarity score S:
[0068] in: For the first The weights of each feature; The indicator function is set to 1 if at least one feature in the group is successfully matched, and 0 otherwise.
[0069] Maximum possible score:
[0070] Normalized similarity: ; ④ Result Judgment: Calculate the similarity Sim between the audio to be identified and each type in the database. Select the type with the highest Sim as the identification result. If Sim=0 for all types (i.e., no match in any of the three groups), then it is judged as "unknown".
[0071] The group-weighted matching strategy exhibits significant technical advantages in complex fire scene environments, including: (1) strong resistance to noise interference. When the audio contains sudden transient noise (such as the sound of an object falling and hitting), the time domain characteristics (such as kurtosis and peak factor) will change abruptly, causing the Gtime matching score to decrease or even become 0. However, if the essential frequency structure of the sound remains unchanged (for example, it is still the high-frequency characteristic of lithium battery depressurization), the Gfreq group can still be successfully matched. According to the rules, as long as all three groups have effective matches, the system can still correctly identify it. This solves the problem of traditional global matching where local noise masks the overall signal characteristics. (2) Reduced misjudgment rate of "similar sounds". For difficult and confusing items in fire investigations (such as blowtorch lighters and windproof lighters, both of which have high time domain energy and similar waveforms), it is easy to misjudge based on a single feature. In this method, the energy variance (time-frequency domain group) and spectral flatness (frequency domain group) distribution range of the blowtorch lighter are significantly different from those of the windproof lighter. By using the time-frequency correlation mean and the differential threshold of the spectral centroid, the system can capture the subtle differences between the Gfreq and Gtf groups in terms of high-frequency decay rate and spectral evolution law, which greatly improves the recognition accuracy of similar sounds that are difficult to distinguish by traditional subjective listening or waveform comparison, and effectively avoids misleading the direction of fire investigation due to incorrect judgment of ignition device type.
[0072] In general, the characteristics of fire audio in different scenarios differ in their physical excitation mechanisms. It is difficult to identify complex fire audio using only a single feature parameter. However, by using a unified feature system for quantitative comparison and analyzing multiple feature parameters, accurate identification of different types of fire audio can be achieved. The systematic comparison of fire audio not only verifies its diversity in energy structure, frequency distribution, and dynamic evolution, but also provides a solid data foundation for the construction of subsequent audio recognition, fault monitoring, and diagnosis algorithms. In the future, combining deep learning and multimodal perception methods holds promise for achieving physical event recognition and state prediction based on audio, expanding the application boundaries of acoustic analysis in areas such as intelligent manufacturing, safety detection, and behavioral modeling.
[0073] By systematically analyzing the acoustic signatures of ignition devices, key acoustic features were extracted, and the methods described in this invention were used to classify and identify the sounds of different types of ignition devices. The experiment included ignition tests with four different types of lighters: windproof, roller, spray gun, and piezoelectric ceramic, as well as a control group of match ignition tests. Ignition sounds were recorded for each group, with at least 50 recordings per group. The recorded sounds were then categorized and a voiceprint comparison database was established. The results are as follows: Figures 2 to 14 As shown: 1. Piezoelectric ceramic lighters. Piezoelectric ceramic lighters produce sound signals with intense energy fluctuations, with the strongest high-frequency pulse signals. The energy is released in a concentrated period of time, making it the most transient ignition method.
[0074] II. Windproof Lighters. Windproof lighters exhibit high energy fluctuations (high energy in short periods) and significant changes in sound intensity. They contain a high proportion of high-frequency components (high spectral centroid), indicating significant airflow noise from the flame. They show obvious non-steady-state characteristics (low energy entropy), with the sound signal varying considerably over time and exhibiting uneven energy distribution. A high zero-crossing rate indicates the presence of more high-frequency pulses in the sound signal. High kurtosis and crest factor suggest the possible presence of numerous sudden high-energy pulses in the signal, such as those from flame ignition or gas impact. The sound signal exhibits significant energy fluctuations, rich high-frequency components, strong pulse characteristics, and significant changes in time-frequency characteristics, demonstrating strong characteristics of sudden signal occurrence.
[0075] III. Flamethrower Lighter. The flamethrower exhibits moderate energy variation (high short-term energy and energy variance, but lower than piezoelectric ceramic lighters). It displays a noticeable high-frequency component (high spectral centroid), but is lower than windproof lighters. The sound signal energy is high, with a noticeable high-frequency component, containing some sudden pulses, but overall more uniform and stable than windproof lighters.
[0076] IV. Flint-and-Steel Wheel Lighters. Flint-and-steel wheel lighters exhibit relatively stable energy (lower short-term energy) and a higher proportion of low-frequency components (lower spectral centroid). They have a lower zero-crossing rate and less pronounced high-frequency components. Low crest factor and waveform factor indicate a smoother sound without sharp, sudden pulses. They also have higher energy entropy and a more uniform energy distribution. Furthermore, they exhibit higher time-frequency correlation and more stable energy.
[0077] V. Matches. Matches exhibit stable energy (lowest short-time energy and lowest energy variance), resulting in minimal changes in the sound signal. Low-frequency components dominate (lowest spectral centroid), indicating that the sound is primarily concentrated in the low-frequency region. The lowest zero-crossing rate indicates fewer high-frequency components. High energy entropy suggests a relatively uniform energy distribution in the sound signal, without drastic fluctuations. A low crest factor indicates a relatively stable signal with no obvious high-energy bursts. High time-frequency correlation indicates a uniform signal energy distribution with minimal variation.
[0078] By comparing the sound signature characteristics of different types of lighters during ignition, it can be found that the energy fluctuations, high-frequency pulse intensity, energy fluctuations, and energy distribution decrease sequentially from piezoelectric ceramic lighters, windproof lighters, blowtorch lighters, flame-sparkling steel wheel lighters, and match ignition. Specific characteristics are summarized in Table 1.
[0079] Table 1. Summary of Ignition Soundprint Characteristics of Different Types of Lighters
[0080] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0081] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A survey and assessment technique based on multidimensional audio analysis of fire accident scenes, characterized in that, Includes the following steps: Extract the target audio signal from the fire scene surveillance video and convert it into a digital audio signal; The digital audio signal is segmented into frames to obtain multiple short-time audio signals. The time-domain feature parameters, frequency-domain feature parameters, and time-frequency-domain feature parameters of each frame of short audio signal are extracted respectively; The weights of each feature parameter are determined based on the CNN convolutional neural network algorithm; A grouped weighted matching strategy is used to calculate the similarity between the audio to be tested and various types of audio in the reference audio library, and the fire audio type is determined based on the highest similarity value.
2. The investigation and assessment technique based on multidimensional audio analysis at a fire accident scene, as described in claim 1, is characterized in that... Extract the temporal characteristic parameters of each frame of short-time audio signal, including mean, variance, kurtosis, short-time energy, zero-crossing rate, crest factor, and waveform factor.
3. The investigation and assessment technique based on multidimensional audio analysis at a fire accident scene according to claim 1, characterized in that, Frequency domain characteristic parameters are determined as follows: The centroid of a spectrum is used to characterize the centroid frequency of the spectral energy distribution; Spectral roll-off is used to select the frequency point where the cumulative energy accounts for 95% of the total energy; Spectral flatness is the ratio of the geometric mean to the arithmetic mean of the spectrum, characterizing the smoothness of the spectral distribution.
4. The investigation and assessment technique based on multidimensional audio analysis at a fire accident scene, as described in claim 1, is characterized in that... The time-frequency domain feature parameters are calculated based on the time-frequency energy matrix of continuous wavelet transform; among them, energy entropy is used to measure the complexity of signal energy distribution in each sub-time window; energy variance is used to measure the degree of fluctuation of short-time energy in different time windows; and the time-frequency correlation mean is used to calculate the correlation coefficient between the time axis and the frequency axis, characterizing the time-varying correlation of voiceprints.
5. The investigation and assessment technique based on multidimensional audio analysis at a fire accident scene according to claim 1, characterized in that, A group-weighted matching strategy is used to calculate the similarity between the audio to be tested and various types of audio in the reference audio library. Specifically, the multidimensional feature parameters are divided into time-domain feature group, frequency-domain feature group, and time-frequency domain feature group. At least one feature in each group must be successfully matched for the weight of that group to be included in the total score. The tolerance threshold of each feature is set according to experimental statistics to cover 99.7% of the normal fluctuation range, and the similarity score is calculated in combination with the weight.
6. The investigation and assessment technique based on multidimensional audio analysis at a fire accident scene according to claim 1, characterized in that, For ignition-related fire audio, a multi-level feature combination is used for differentiation: first, kurtosis and variance are used to differentiate the audio of a spray gun ignition; then, zero-crossing rate and peak factor are used to differentiate the audio of a windproof lighter ignition; then, spectral centroid and spectral flatness are used to differentiate the audio of a match ignition; finally, energy entropy is used to differentiate the audio of a piezoelectric ceramic lighter from that of a roller lighter.
7. The investigation and assessment technique based on multidimensional audio analysis at a fire accident scene according to claim 1, characterized in that, For audio signals of electrical short circuit fires, a characteristic combination judgment method is adopted: zero-crossing rate is used to distinguish between AC and DC short circuits; short-time energy, variance, and spectral centroid are used to determine the voltage level, wire diameter, and material of AC short circuits; and energy entropy and time-frequency correlation mean are used to determine the voltage level of DC short circuits.