A method and device for identifying audio features of panicked pedestrians screaming
By analyzing audio features in the 0-750Hz frequency band and establishing a recognition matrix model, the robustness problem of audio feature recognition under panic was solved, and accurate identification of panic behavior and gender was achieved, improving the accuracy of crowd flow safety analysis in public places.
Patent Information
- Application Number
- CN202411366296.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-29
- Publication Date
- 2026-03-06
- Estimated Expiration
- 2044-09-29
AI Technical Summary
Existing technologies lack robustness and environmental adaptability in recognizing pedestrian audio features during panic, and audio-based panic behavior analysis mainly relies on the characteristics of crowd limb movements, lacking in-depth analysis of audio content.
By analyzing the audio characteristics of the 0-750Hz frequency band, calculating the audio energy distribution, and establishing a panic behavior audio feature recognition matrix model, combined with gender feature recognition, the audio energy range is divided using Fourier transform and normal distribution fitting plot to identify panic behavior and gender.
It effectively identifies screaming behavior under panic, possesses robustness and environmental adaptability, can accurately determine panic behavior and gender, and improves the accuracy of panic behavior analysis.
Smart Images

Figure CN119380746B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of crowd flow stability analysis, and in particular to a method and apparatus for identifying audio features of panicked pedestrians screaming. Background Technology
[0002] With the increasing coverage of network surveillance cameras in public places and the abundance of audio data, human pedestrian behavior recognition technology based on audio features has been developed and applied in the field of public place crowd stability research. Domestic and international scholars began researching crowd evacuation and emergencies at the beginning of the last century. In various public place emergencies, pedestrians are easily influenced by the environment and fall into panic, which in turn triggers various panic behaviors. Therefore, this paper, based on existing audio feature recognition technology, considers the audio content that may appear in different emergencies, analyzes the audio features of pedestrians screaming during accidents, and applies scream audio feature recognition to the analysis of panicked crowd behavior and the study of disturbance propagation, establishing a crowd panic behavior audio feature recognition model. This lays the foundation for further exploration of the dynamic impact of panic behavior disturbances on crowd stability and provides new ideas for research on the safety of crowd flow in public places.
[0003] Current research still has several shortcomings:
[0004] 1. Currently, research on pedestrian audio feature recognition rarely considers the situation where audio content is distorted or incomplete due to panic, and there is a lack of robust and environmentally adaptable audio feature criteria for panic behavior.
[0005] 2. Existing research on panic behavior analysis based on deep learning methods is mostly based on the movement characteristics of people's limbs, and there are few studies that use audio features such as sound frequency, pitch or semantic content extracted from audio data to analyze panic behavior. Summary of the Invention
[0006] The purpose of this invention is to provide a method and apparatus for identifying audio features of panicked pedestrians screaming. By analyzing audio features, a panic behavior audio feature identification matrix model is established to identify the screaming behavior of panicked pedestrians, which has robustness and environmental adaptability.
[0007] The objective of this invention can be achieved through the following technical solutions:
[0008] A method for identifying audio features of panicked pedestrian screaming behavior, the method comprising:
[0009] Extracting audio from video;
[0010] Extract the audio features of the first frequency band from the audio and calculate the audio energy distribution;
[0011] The audio energy is divided into intervals, and a matrix model for recognizing audio features of panic behavior is established.
[0012] The audio feature recognition matrix model is used to determine the type of behavior and gender in the audio.
[0013] Furthermore, the first frequency band is the 0-750Hz band.
[0014] Furthermore, the process of extracting the audio features of the first frequency band in the audio and calculating the audio energy distribution includes:
[0015] Perform Fourier transform on the audio data to obtain the waveform and spectrum of the audio signal;
[0016] Extract the audio characteristics of the first frequency band from the spectrum of the audio signal;
[0017] The audio energy of the first frequency band is normalized, and the audio energy distribution is calculated.
[0018] Furthermore, the process of dividing the audio energy into intervals includes:
[0019] Obtain an audio dataset of panicked pedestrians screaming.
[0020] Randomly select audio recordings of male and female panic behaviors from the dataset and analyze the energy distribution of the first frequency band.
[0021] Plot the normal distribution fit of the energy of audio recordings of male and female panic behaviors, and calculate the mean and standard deviation of the energy.
[0022] By combining the normal distribution fitting plots of audio energy for different genders, the mean energy, and the standard deviation of energy, the audio energy is divided into intervals.
[0023] Furthermore, the audio energy division intervals include: [0, 0.009], [0.009, 0.053], [0.053, 0.076], [0.076, 0.172], [0.172, 0.178], [0.178, 0.206], [0.206, 0.634] and [0.634, 1].
[0024] Furthermore, the process of establishing the panic behavior audio feature recognition matrix model includes:
[0025] Extract a frame from the audio and calculate the energy value of this frame in the first frequency band;
[0026] Normalize audio energy values;
[0027] The normalized audio energy values are assigned to the corresponding partitioned intervals, and the corresponding partitioned intervals are set to 1.
[0028] Repeat the above steps to process the complete audio frame by frame, and finally establish a panic behavior audio feature recognition matrix model.
[0029] Furthermore, the expression for the panic behavior audio feature recognition matrix model is as follows:
[0030]
[0031] Among them, C1, C2, C3, C4, C5, C6, C7, and C8 are the division intervals of audio energy.
[0032] Furthermore, C1, C3, and C5 are the division intervals for audio recordings of panic behavior with ambiguous gender distinction; C2 is the division interval for audio recordings of female panic behavior; C4 is the division interval for audio recordings of male panic behavior; C6 and C8 are the division intervals for audio recordings of behaviors that cannot be distinguished; and C7 is the division interval for audio recordings of normal communication behavior.
[0033] Furthermore, based on the type of interval into which the audio energy mainly falls in the audio feature recognition matrix model, the behavioral type and gender in the audio can be determined.
[0034] An electronic device comprising:
[0035] One or more processors;
[0036] Memory;
[0037] One or more programs stored in memory, the one or more programs including instructions for performing any of the above-described methods for identifying audio features of panicked pedestrian screaming behavior.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] 1. This invention considers the possibility that the content of pedestrians' shouts may be unclear under the influence of panic. It focuses on screaming behavior and analyzes the audio features of the first frequency band that can effectively distinguish panic behavior from normal communication behavior. It obtains the audio energy distribution and divides the audio energy into intervals to establish a panic behavior audio feature recognition matrix model to determine the type of behavior and gender in the audio. This method makes the determination of panicked pedestrian screaming behavior robust and environmentally adaptable, and can more effectively determine whether pedestrians are panicked.
[0040] 2. This invention analyzes the audio energy of panicked pedestrians screaming in different genders, plots a normal distribution fit graph, calculates the mean and standard deviation of audio energy, divides the energy values of video audio into intervals, and then establishes a panic behavior audio feature recognition matrix model, which can effectively identify panic behavior while distinguishing gender in the audio. Attached Figure Description
[0041] Figure 1 This is a process flow diagram of the present invention;
[0042] Figure 2 Audio spectrum analysis of normal male communication behavior;
[0043] Figure 3 Audio spectrum analysis of normal female communication behavior;
[0044] Figure 4 This is a fitting graph of the normal distribution of normal AC audio energy.
[0045] Figure 5 Audio spectrum energy normalization and error bar statistics for normal communication behavior;
[0046] Figure 6 Audio spectrum analysis of women's panic screaming behavior;
[0047] Figure 7 Audio spectrum analysis of male panic screaming behavior;
[0048] Figure 8 Audio spectral energy normalization and error bar statistics for panic screaming behavior;
[0049] Figure 9 A normal distribution fitting plot of the energy of audio features of male and female panic screaming behaviors. Detailed Implementation
[0050] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.
[0051] Example
[0052] This embodiment discloses a method for recognizing audio features of panicked pedestrians screaming, such as... Figure 1 As shown, the method includes:
[0053] S1, extract the audio portion from the video;
[0054] S2, extract the audio features of the first frequency band in the audio and calculate the audio energy distribution;
[0055] S3, divide the audio energy into intervals and establish a matrix model for recognizing audio features of panic behavior;
[0056] S4, determine the behavior type and gender in the audio based on the audio feature recognition matrix model.
[0057] The specific implementation process of the method is as follows:
[0058] First, based on audio data of normal communication behavior, we analyze the characteristics of the audio data.
[0059] The dataset for normal communication behavior in this embodiment uses a 200-hour Mandarin Chinese speech dataset selected from the complete dataset. The dataset includes 600 participants, a sampling frequency of 16kHz 16bit, and 50 sets of data randomly selected, including 30 males and 20 females.
[0060] Each frame of audio from normal communication behavior in the dataset is extracted and subjected to Fourier transform to obtain the waveform and spectrum of the audio signal. The frequencies of the audio signals from normal communication behavior are mainly distributed below 6000Hz. A segmented analysis of the audio characteristics of normal communication behavior is used to calculate the average energy of each frequency range, extracting effective information about the spectral characteristics of the audio signal. Here, energy represents the average length of the sound signal after Fourier transform within each frequency range, derived from the collected data, and is a dimensionless value. The segmented calculation formula is shown below:
[0061]
[0062] In the formula, f internal f is the length of each frequency segment. min To capture the minimum frequency of the audio, f max To capture the maximum frequency of the audio, n real To divide the frequency bands, f σ This is the lower limit of the data collection frequency. Based on the statistical results of the collected information and the rationality of the data analysis, f is taken as... max =6000Hz, f min =0Hz, n real =8.
[0063] The energy values are summed across the eight frequency bands of interest and then normalized, as shown in the following formula:
[0064]
[0065] Among them, e′ i e represents the average energy percentage across different frequency bands. i e represents the average energy across all frequency bands. σ This represents the lower limit of the average energy percentage, where e is the lower limit. σ =0.
[0066] The final statistical results show the average energy proportion of each frequency band for normal male and female communication behaviors. The spectrum analysis diagram of normal communication behaviors is shown below. Figure 2 and Figure 3 As shown, Figure 2Audio spectrum analysis of normal male communication behavior. Figure 3 Audio spectrum analysis of normal female communication behavior.
[0067] Furthermore, such as Figure 4 Based on statistical results, a histogram is plotted using the first segment of frequency distribution as an example, and a normal density function is fitted.
[0068] Figure 4 In this context, σ represents the standard deviation, which is derived from the confidence interval.
[0069] A confidence interval is an estimated interval for a population parameter constructed from sample statistics. In statistics, a confidence interval for a probability sample is an interval estimate of a population parameter for that sample. A confidence interval shows the degree to which the true value of the parameter has a certain probability of falling within the range of the measurement result; it gives the confidence level of the measured value of the measured parameter, i.e., "a probability". The probability value of one standard deviation (1σ) is 68.3%, the probability value of two standard deviations (2σ) is 95.5%, and the probability value of three standard deviations (3σ) is 99.7%.
[0070] Figure 4 In the above, a confidence interval of two standard deviations, i.e., a probability value of 95.5%, was selected. The mean and standard deviation obtained from the fitting plot of the normal distribution of audio energy in normal communication behavior are shown in Table 1.
[0071] Table 1 Mean and Standard Deviation of Audio Recordings of Normal Communication Behavior
[0072] Serial Number frequency band Range (Hz) Energy mean σ 1. <![CDATA[f s1 ]]> 0-750 0.405734 0.114096 2. <![CDATA[f s2 ]]> 750-1500 0.200814 0.070349 3. <![CDATA[f s3 ]]> 1500-2250 0.096148 0.047655 4. <![CDATA[f s4 ]]> 2250-3000 0.068588 0.030714 5. <![CDATA[f s5 ]]> 3000-3750 0.054678 0.025836 6. <![CDATA[f s6 ]]> 3750-4500 0.061152 0.037849 7. <![CDATA[f s7 ]]> 4500-5250 0.060240 0.046054 8. <![CDATA[f s8 ]]> 5250-6000 0.052646 0.039515
[0073] Using the method described above, energy normalization and error bar statistics for eight frequency bands of the audio spectrum of normal communication behavior were plotted, as shown below. Figure 5 As shown.
[0074] f in the figure s6 (3750-4500Hz), f s7 (4500-5250Hz) and f s8 The lower limit of the error bar (5250-6000Hz) is all below 0. Based on the characteristic that the average energy percentage value is not less than 0, the lower limit of the error bar below 0 is changed to 0.
[0075] The audio characteristics of panic behavior are mainly manifested in the screaming of pedestrians; therefore, we will begin our discussion with screaming behavior. The audio dataset for panic behavior is selected from the Dataset-AOB: Urban Sound Event Classification dataset. The Dataset-AOB dataset is an audio dataset for urban sound event classification collected and manually edited for a master's thesis using convolutional neural networks. The dataset contains 10 audio events: sirens, children playing, dogs barking, engines, footsteps, glass breaking, gunshots, subway trains, rain, and screams. Panic behavior in crowds is often accompanied by screams. Screams were randomly selected from the Dataset-AOB dataset for analysis, yielding statistical results on the average energy proportion of each frequency band in the audio of male and female panic behavior. The audio spectrum analysis of panic behavior is shown in the figure below. Figure 6 and Figure 7 As shown, Figure 6 Audio spectrum analysis of women's panic screaming behavior. Figure 7 Audio spectrum analysis of male panic screaming behavior.
[0076] The mean and standard deviation obtained from the normal distribution fitting plot of audio energy of panic behavior are shown in Table 2.
[0077] Table 2 Mean and Standard Deviation of Panic Behavior Audio Recordings
[0078] Serial Number frequency band Range (Hz) Energy mean σ 1. <![CDATA[f s1 ]]> 0-750 0.085062 0.060544 2. <![CDATA[f s2 ]]> 750-1500 0.404742 0.149134 3. <![CDATA[f s3 ]]> 1500-2250 0.149954 0.083362 4. <![CDATA[f s4 ]]> 2250-3000 0.114022 0.0895983 5. <![CDATA[f s5 ]]> 3000-3750 0.06373 0.0350749 6. <![CDATA[f s6 ]]> 3750-4500 0.07059 0.0387502 7. <![CDATA[f s7 ]]> 4500-5250 0.055618 0.0411158 8. <![CDATA[f s8 ]]> 5250-6000 0.050986 0.0303498
[0079] Based on the audio spectrum energy normalization and error bar statistical analysis method used for normal communication behavior, the audio spectrum energy normalization and error bar statistical chart of panic behavior are obtained as follows: Figure 8 As shown.
[0080] Figure 8 Chinese f s1 (0-750Hz), f s3 (1500-2250Hz), f s4 (2250-3000Hz), f s5 (3000-3750Hz), f s6 (3750-4500Hz), f s7 (4500-5250Hz) and f s8 The lower limit of the error bar (5250-6000Hz) is all below 0. Based on the characteristic that the average energy percentage value is not less than 0, the lower limit of the error bar below 0 is changed to 0.
[0081] Based on the above data, the conclusion is that the audio feature analysis of normal communication behavior and the audio feature analysis of panic behavior are similar in f s1 There is a clear difference between (0-750Hz).
[0082] Therefore, in the audio feature recognition method for panicked pedestrian screaming behavior in this embodiment, the first frequency band is the 0-750Hz frequency band.
[0083] In practical applications, besides effectively identifying panic behavior audio, effectively identifying the gender within panic behavior audio is also of great significance. Therefore, we randomly selected panic behavior audio from male and female participants in the dataset and analyzed f s1 (0-750Hz), the mean and standard deviation of the normal distribution fitting plots of the audio energy of male panic behavior and female panic behavior are shown in Table 3. The normal distribution fitting plots of the audio energy features of male panic behavior and female panic behavior are shown in Table 3. Figure 9 As shown.
[0084] Table 3. Audio Mean and Standard Deviation of Panic Behavior in Men and Women
[0085] Serial Number gender frequency band Range (Hz) Energy mean σ 1. male <![CDATA[f s1 ]]> 0-750 0.124179 0.047869 2. female <![CDATA[f s1 ]]> 0-750 0.031043 0.022134
[0086] according to Figure 9 As shown in Table 3, the audio energy distribution of panic behavior differs between men and women.
[0087] Furthermore, the energy percentage is divided into eight intervals: 0-0.009, 0.009-0.053, 0.053-0.076, 0.076-0.172, 0.172-0.178, 0.178-0.206, 0.206-0.634, and 0.634-1.
[0088] Furthermore, a matrix model for recognizing audio features of panic behavior is established, the process of which includes:
[0089] Extract a frame from the audio and calculate the energy value of this frame in the first frequency band;
[0090] Normalize audio energy values;
[0091] The normalized audio energy values are assigned to the corresponding partitioned intervals, and the corresponding partitioned intervals are set to 1.
[0092] Repeat the above steps to process the complete audio frame by frame, and finally establish a panic behavior audio feature recognition matrix model.
[0093] Furthermore, the expression for the audio feature recognition matrix model is:
[0094]
[0095] Among them, C1, C2, C3, C4, C5, C6, C7, and C8 are the division intervals of audio energy.
[0096] Furthermore, C1, C3, and C5 are the division intervals for audio recordings of panic behavior with ambiguous gender distinction; C2 is the division interval for audio recordings of female panic behavior; C4 is the division interval for audio recordings of male panic behavior; C6 and C8 are the division intervals for audio recordings of behaviors that cannot be distinguished; and C7 is the division interval for audio recordings of normal communication behavior.
[0097] Furthermore, based on the type of interval into which the audio energy mainly falls in the audio feature recognition matrix model, the behavioral type and gender in the audio can be determined.
[0098] Taking a certain location as an example, the surveillance video of that location was used to verify the audio feature recognition model of panicked pedestrians screaming through real panic event videos.
[0099] Extract the audio portion of a video from a specific location, and then extract f from it. s1 Audio characteristics in the (0-750Hz) frequency band. f is calculated by normalizing the audio spectral energy. s1 The audio energy distribution in the frequency band, after normalization, should have an energy value within the (0,1) interval. Based on the eight energy distribution intervals, the normalized f values are calculated... s1 The energy result in the frequency band is placed into the corresponding energy range. For example, at time t=1s in the video, if f s1 The audio energy of the frequency band, after normalization, is 0.033. Therefore, the C2 interval at that moment is set to 1, indicating that a woman's panicked screaming behavior was detected in the audio at this time. Following the above method, the complete audio is normalized and data is inserted into the corresponding intervals to obtain the calculation result of the panic behavior audio feature recognition matrix model, represented as [0,1,0,0,0,0,0,0]. That is, C2 panic behavior was identified in the audio, and it is believed that a woman's panicked screaming behavior occurred in the video.
[0100] If the above methods are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0101] In another embodiment, an electronic device is provided, characterized in that it includes one or more processors, a memory, and one or more programs stored in the memory, said one or more programs including instructions for performing the method for identifying audio features of panicked pedestrian screaming behavior as described above.
[0102] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for identifying audio features of panicked pedestrian screaming behavior, characterized in that, The method comprises: extracting an audio part from a video; cutting the audio features of a first frequency band in the audio and calculating the audio energy distribution; the first frequency band is 0-750 Hz; dividing the audio energy by interval and establishing a panic behavior audio feature recognition matrix model; determining the behavior type and gender in the audio according to the audio feature recognition matrix model; The process of dividing the audio energy by interval comprises: obtaining a panic pedestrian screaming behavior audio data set; randomly selecting male panic behavior audio and female panic behavior audio in the data set, and analyzing the energy distribution of the first frequency band audio; drawing a normal distribution fitting graph of male panic behavior audio and female panic behavior audio energy, and calculating the energy mean and energy standard deviation; combining the normal distribution fitting graph of audio energy of different genders, the energy mean and the energy standard deviation to divide the audio energy by interval; The division interval of the audio energy includes: , , , , , , , ; The expression of the panic behavior audio feature recognition matrix model is: , wherein, a division interval of audio energy; The and is a division interval for panic behavior but gender-ambiguous audio, is a division interval for female panic behavior audio, is a division interval for male panic behavior audio, is a division interval for behavior audio that cannot be distinguished, is a division interval for normal communication behavior audio.
2. The method of claim 1, wherein the method further comprises: The process of cutting the audio features of the first frequency band in the audio and calculating the audio energy distribution comprises: performing Fourier transform on the audio data to obtain the waveform and spectrum of the audio signal; cutting the audio features of the first frequency band from the spectrum of the audio signal; normalizing the audio energy of the first frequency band and calculating the audio energy distribution.
3. The method of claim 1, wherein the method further comprises: The process of establishing the panic behavior audio feature recognition matrix model comprises: cutting a frame of audio and calculating the energy value of the frame of audio in the first frequency band; normalizing the audio energy value; putting the normalized audio energy value into the corresponding divided interval and setting the corresponding divided interval to 1; repeating the above steps to process the complete audio frame by frame, and finally establishing the panic behavior audio feature recognition matrix model.
4. The method of claim 1, wherein the method further comprises: Determine the behavior type and gender in the audio according to the type of the divided interval where the audio energy in the audio feature recognition matrix model is mainly concentrated.
5. An electronic device, comprising: It includes: one or more processors; memory; one or more programs stored in the memory, the one or more programs including instructions for performing the panic pedestrian screaming behavior audio feature recognition method as claimed in any one of claims 1-4.
Citation Information
Patent Citations
Voice gender recognition method and device, storage medium and computer equipment
CN114049881A
Scream detecting device for surveillance systems based on audio data and, the method thereof
KR101578108B1