Voice signal edge processing system and method for audio monitoring

By collecting sound pressure polarization vector flow and vibration signals in an audio monitoring system, extracting features, and using a lightweight decision model and generative adversarial network for event determination, the problem of insufficient recognition of privacy violations and verbal violence in existing technologies is solved, achieving accurate identification and rapid response to violent events.

CN120932682AInactive Publication Date: 2025-11-11WUHAN DASHENGJI TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511395438.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2025-11-11
Estimated Expiration
Not applicable · inactive patent

Smart Images

  • Figure CN120932682A_ABST
    Figure CN120932682A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of mechanical perception, and discloses a voice signal edge processing system and method for audio monitoring. The method comprises the following steps: acquiring sound pressure polarization vector flow, a vibration signal and a sound pressure signal; performing feature extraction and fusion on the sound pressure polarization vector flow, the vibration signal and the sound pressure signal to obtain a sound pressure polarization feature and a sound force coupling feature; taking the sound pressure polarization characteristics and the sound force coupling characteristics as input of a lightweight decision model to obtain an event judgment result; if the event judgment result is a violent event, recording records W seconds before and after the judgment are encrypted and stored, synchronously giving an alarm to a teacher end and a security system, and pushing three-dimensional positioning information; if the event judgment result is normal, no intervention is performed; according to the method, dependence on voice content is avoided so as to effectively protect privacy of monitored objects such as students, and violent events are accurately recognized through multi-dimensional anomaly detection integrating sound pressure polarization characteristics and sound force coupling characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of mechanical sensing technology, and more specifically, to a voice signal edge processing system and method for audio monitoring. Background Technology

[0002] School bullying, a prominent issue endangering students' physical and mental health, presents significant challenges to traditional prevention methods due to its hidden, sudden, and non-verbal nature. Current school security systems suffer from limitations such as blind spots in video surveillance, restricted deployment in privacy-sensitive areas, and the potential to provoke resistance. Manual patrols, relying on manpower efficiency, struggle to achieve 24 / 7, comprehensive monitoring. Audio signals, including sounds of physical impact, crying, and arguments, serve as crucial clues in bullying incidents and play an irreplaceable supplementary role in surveillance.

[0003] Chinese patent application CN118761867A discloses a campus violence early warning system: It acquires environmental audio signals in real time through audio acquisition devices installed on campus, transmits the acquired audio data to a cloud server for speech recognition processing, and extracts semantic information from the speech content; it constructs a lexicon containing violence-related sensitive words, and determines the existence of violence risk by comparing and analyzing the matching degree between the identified speech semantics and the sensitive word lexicon; when a matching sensitive word is detected, the system triggers an early warning mechanism, sending alarm information to teachers and the security system, and providing the location information of the incident using positioning technology. This invention uses a privacy-preserving audio recognition model for privacy areas and processes the acquired audio data; it also stores privacy-preserving audio data only when violent behavior is identified, achieving privacy protection of audio data without affecting the investigation of violent incidents.

[0004] While the above methods can meet the needs of most scenarios, research and practical application of these methods and existing technologies have revealed at least the following shortcomings:

[0005] The aforementioned system relies on parsing speech and semantic data for violence identification, which poses a risk of infringing on student privacy and has a weak ability to identify violent behavior without verbal communication.

[0006] In view of this, the present invention proposes a voice signal edge processing system and method for audio monitoring to solve the above problems. Summary of the Invention

[0007] To overcome the aforementioned deficiencies of the prior art and to achieve the above objectives, the present invention provides the following technical solution: a voice signal edge processing method for audio monitoring, comprising the following steps:

[0008] Acquire sound pressure polarization vector flow, vibration signal, and sound pressure signal;

[0009] Feature extraction and fusion are performed on the sound pressure polarization vector flow, vibration signal, and sound pressure signal to obtain sound pressure polarization features and acoustic-mechanical coupling features. The sound pressure polarization features include polarization chaos, polarization angle chaos, and curl-divergence ratio. The acoustic-mechanical coupling features include vibration spectral entropy, sound pressure energy abrupt gradient, and acoustic-mechanical coupling anomaly.

[0010] The acoustic pressure polarization characteristics and acoustic-mechanical coupling characteristics are used as inputs to a lightweight decision model to obtain event determination results;

[0011] If the event is determined to be a violent event, the audio recording of W seconds before and after the determination is encrypted and stored, and an alarm is simultaneously sent to the teacher's terminal and the security system, and 3D location information is pushed; if the event is determined to be normal, no intervention is performed.

[0012] Furthermore, methods for obtaining the acoustic pressure polarization vector flow include:

[0013] The sound pressure force components in the x, y, and z axes of a preset characteristic scene are collected by an edge device.

[0014] The polarization angle is calculated using trigonometric functions.

[0015] Time series data describing the system's pressure measurements over time are acquired, preprocessed, and then arranged into a matrix to obtain a data matrix. and ,in, and yes matrix, This refers to the dimension of state variables, i.e., the number of force components. The number of samples is represented by each column; each column represents the state at a given moment, and each row corresponds to a different displacement.

[0016] For data matrix Perform singular value decomposition, decomposing it into a left singular vector matrix, a singular value diagonal matrix, and a right singular vector matrix; based on the left singular vector matrix, the singular value diagonal matrix, and the data matrix... The linear operator is obtained through computation;

[0017] Eigenvalue decomposition is performed on a linear operator to obtain eigenvalues ​​and corresponding eigenvectors;

[0018] The DMD mode is obtained by calculating the left singular vector matrix and the corresponding eigenvector. The system state is reconstructed based on the DMD mode and eigenvalues ​​to predict the system state at future time.

[0019] Obtain the imaginary part of the eigenvalues, and calculate the phase angle based on the DMD mode and the real part of the eigenvalues;

[0020] The sound pressure amplitude is obtained by vector synthesis of the force components along the x, y, and z axes.

[0021] By splicing together the polarization angle, phase angle, and sound pressure amplitude, a sound pressure polarization vector flow is obtained.

[0022] Furthermore, methods for obtaining sound pressure polarization characteristics include:

[0023] Calculate the sound pressure amplitude of the sound pressure polarization vector flow, calculate the standard deviation and mean of the sound pressure amplitude, calculate the ratio of the standard deviation to the mean, and obtain the initial polarization chaos degree.

[0024] Calculate the first-order difference of the phase angle, count the number of times the first-order difference exceeds the difference threshold per unit time, and calculate the polarization chaos degree based on the number of times and the initial polarization chaos degree.

[0025] Calculate the time standard deviation and the mean of the absolute value of the polarization angle, and calculate the ratio of the time standard deviation to the mean of the absolute value of the polarization angle to obtain the polarization angle chaos degree;

[0026] The phase angle and sound pressure amplitude are divided into K intervals. All K sets of data are iterated through, and the polarization angle is statistically analyzed on the [Kth]th interval. Each phase angle divides the interval And the sound pressure modulus falls on the first Each sound pressure amplitude is divided into intervals The number of samples is determined, the ratio of the number of samples to K is calculated, the joint probability distribution is obtained, and the phase amplitude coupling entropy is calculated based on the joint probability distribution.

[0027] Calculate the curl and divergence of the sound pressure polarization vector flow, calculate the magnitudes of the curl and divergence respectively, and then calculate the ratio of the magnitude of the curl to the magnitude of the divergence to obtain the initial curl-divergence ratio; calculate the magnitude of the spatial gradient of the polarization angle, and then calculate the curl-divergence ratio based on the initial curl-divergence ratio and the magnitude of the spatial gradient of the polarization angle.

[0028] Furthermore, methods for obtaining acoustic-mechanical coupling characteristics include:

[0029] The vibration signal is filtered by using a Hanning window with a window length of 1 / L of the signal length, moving in steps of L / 2. A fast Fourier transform is performed on the signal within each window to obtain the power spectrum of each segment. The power spectrum of segment H is averaged to obtain the power spectral density. Based on the power spectral density, the vibration spectral entropy is calculated using information entropy.

[0030] Calculate the first derivative of the sound pressure signal, obtain the maximum value of the first derivative, and obtain the abrupt gradient of the sound pressure energy.

[0031] Calculate the angle of the vibration direction, and obtain the directional consistency based on the angle and polarization angle;

[0032] Calculate the phase angle and the phase difference of the vibration signal, and obtain the synchronization based on the phase difference;

[0033] The power spectral density of the sound pressure signal is calculated. The conjugate of the sound pressure signal is calculated based on its frequency domain representation. The cross power spectrum of the sound pressure signal and the vibration signal is calculated based on their conjugate and frequency domain representations. The cross power spectral density is then calculated based on the cross power spectrum, the power spectral density of the sound pressure signal, and the power spectral density of the vibration signal.

[0034] Calculate the time difference between the moment of abrupt change in the sound pressure phase angle and the moment of peak vibration spectrum, and obtain the time synchronization based on the time difference;

[0035] The acoustic pressure polarization features and acoustic-mechanical coupling features are used as inputs to the anomaly analysis model to obtain anomaly scores. For the acoustic pressure polarization features and acoustic-mechanical coupling features to be detected, the reconstruction features are reconstructed based on the output of the generative adversarial network, and the reconstruction error is calculated. The mean and standard deviation of the reconstruction error are calculated through the training set. The anomaly scores and reconstruction errors are weighted and fused to obtain the acoustic-mechanical coupling anomaly values.

[0036] Furthermore, training methods for generative adversarial networks include:

[0037] This will contain Z sets of real data pairs. The training set data is divided according to a preset batch size, and one batch of fused multimodal data, namely the joint features of sound pressure polarization characteristics and acoustic coupling characteristics, is extracted for training each time; among them, It is the set of all acoustic pressure polarization features extracted from the acoustic pressure polarization vector flow; It is a collection of all acoustic-mechanical coupling features extracted and fused from vibration signals and sound pressure signals;

[0038] The combined features of acoustic pressure polarization and acoustic-mechanical coupling are used as input to the generator G to obtain reconstructed features; the discriminator determines the probability that the input features belong to the real data pairs.

[0039] A batch of real data is randomly selected from the training set. The selected real data and the generated data generated by the generator are used as inputs to the discriminator, and the discriminator outputs the corresponding discrimination probability. The loss function corresponding to the discriminator is designed. The network parameters of the discriminator are updated through the backpropagation algorithm.

[0040] The process continues until a preset optimal performance condition is reached, at which point the generative adversarial network (GAN) corresponding to the optimal performance condition is taken as the final output GAN.

[0041] Furthermore, methods for obtaining event determination results include:

[0042] If the sound pressure polarization feature or the acoustic coupling feature is triggered, it is determined whether the preset combination conditions are met: at least any 2 items of the sound pressure polarization feature and any 3 items of the acoustic coupling feature are triggered, and it is not an interference scene, then it is determined to be a violent scene;

[0043] Methods for triggering acoustic pressure polarization characteristics or acoustic-mechanical coupling characteristics include:

[0044] The sound pressure polarization feature is triggered if any of the following conditions are met:

[0045] The polarization chaos degree is greater than the polarization dynamic threshold;

[0046] The polarization angle chaos degree is greater than the polarization angle threshold;

[0047] The curl-divergence ratio is greater than the ratio threshold;

[0048] The acoustic-mechanical coupling feature is triggered if any of the following conditions are met:

[0049] The vibration spectral entropy is greater than the dynamic threshold of spectral entropy.

[0050] The abrupt gradient of sound pressure energy exceeds the dynamic threshold of the gradient.

[0051] Directional consistency is below the dynamic consistency threshold;

[0052] Synchronization is below the dynamic threshold of synchronization;

[0053] Time synchronization is below the time dynamic threshold;

[0054] The abnormal value of acoustic-mechanical coupling exceeds the abnormal threshold.

[0055] Furthermore, methods for obtaining the dynamic consistency threshold include:

[0056] The mean and standard deviation of directional consistency are calculated based on historical directional consistency data, and the dynamic threshold of consistency is obtained based on the mean and standard deviation.

[0057] Methods for obtaining the dynamic threshold of synchronization include:

[0058] The synchronization mean and synchronization standard deviation are calculated based on historical synchronization data, and the synchronization dynamic threshold is obtained based on the mean and standard deviation.

[0059] Methods for obtaining time dynamic thresholds include:

[0060] An initial dynamic time threshold is defined based on historical time synchronization data, and then corrected according to real-time noise level and real-time time delay to obtain the dynamic time threshold.

[0061] Furthermore, methods for obtaining the polarization dynamic threshold include:

[0062] Based on historical normal data, a probability density function is constructed using kernel density estimation, and an initial threshold is set as the preset quantile of the probability density function. The ratio of real-time number of people to maximum capacity and the ratio of environmental noise to the noise limit are calculated. Combined with the initial threshold, the polarization dynamic threshold is calculated.

[0063] Furthermore, methods for obtaining the gradient dynamic threshold include:

[0064] The sound pressure energy is calculated based on the sound pressure signal and air density. The mean and standard deviation of the current background sound pressure energy are calculated in real time using a sliding window of length Twin.

[0065] The mean value of the current background sound pressure energy is smoothed by using an exponentially weighted moving average to obtain a smoothed mean.

[0066] The base threshold is calculated based on the smoothed mean and the standard deviation of the current background sound pressure energy. The base threshold is then adjusted based on the real-time number of people and time to obtain the gradient dynamic threshold.

[0067] Furthermore, methods for obtaining the dynamic threshold of spectral entropy include:

[0068] Perform a short-time Fourier transform on the vibration signal to obtain the spectrum matrix, and extract the frequency components.

[0069] Calculate the energy percentage of each frequency component, and obtain the Shannon entropy based on the energy percentage.

[0070] Beforehand, K-means clustering analysis is used to analyze the vibration spectrum entropy distribution of different locations, generating a mapping relationship between location type and baseline entropy. The baseline entropy corresponding to the current location is obtained, and the updated baseline entropy is obtained based on the baseline entropy corresponding to the current location.

[0071] The dynamic threshold of spectral entropy is calculated based on the baseline entropy, the updated baseline entropy, and the standard deviation.

[0072] Furthermore, methods for generating the mapping relationship between location type and baseline entropy include:

[0073] The number of clusters is set to be equal to the number of site types, and Euclidean distance is used to measure the dispersion of the vibration spectrum entropy;

[0074] The dataset is grouped by location, and K-means clustering is performed on the vibration spectrum entropy subset for each location;

[0075] Output the vibration spectrum entropy corresponding to the cluster center of each location as the baseline entropy; at the same time, record the standard deviation of the cluster corresponding to the location.

[0076] A voice signal edge processing system for audio monitoring, comprising implementing the aforementioned voice signal edge processing method for audio monitoring, including:

[0077] Data acquisition module: Acquires sound pressure polarization vector flow, vibration signal, and sound pressure signal through edge devices;

[0078] Feature extraction module: Extracts and fuses features from sound pressure polarization vector flow, vibration signal and sound pressure signal to obtain sound pressure polarization features and acoustic-mechanical coupling features; sound pressure polarization features include polarization chaos degree, polarization angle chaos degree and curl-divergence ratio; acoustic-mechanical coupling features include vibration spectrum entropy, sound pressure energy abrupt gradient and acoustic-mechanical coupling anomaly value;

[0079] Event determination module: It takes the sound pressure polarization characteristics and acoustic coupling characteristics as inputs to the lightweight decision model to obtain the event determination results;

[0080] Event handling module: If the event is determined to be a violent event, an alarm will be triggered, and the audio recordings before and after the alarm (W seconds) will be encrypted and stored. Simultaneously, the 3D location information will be pushed to the teacher's end and the security system. If the event is determined to be normal, no intervention will be performed.

[0081] The technical effects and advantages of the speech signal edge processing system and method for audio monitoring of the present invention are as follows:

[0082] This invention extracts non-semantic acoustic features from speech signals at the edge and combines them with dynamic threshold judgment and a lightweight decision model for event analysis. This avoids dependence on speech content, effectively protecting the privacy of monitored subjects such as students. Furthermore, through multi-dimensional anomaly detection that integrates sound pressure polarization features and acoustic coupling features, it accurately identifies violent events, especially significantly improving the detection rate of violent behavior without verbal communication. At the same time, by leveraging the real-time nature of edge processing and the targeted response mechanism after judgment, it achieves rapid response and precise handling of violent events. Ultimately, it improves the security and effectiveness of audio monitoring while protecting privacy. Attached Figure Description

[0083] Figure 1 This is a schematic flowchart of a voice signal edge processing method for audio monitoring according to the present invention;

[0084] Figure 2 This is a schematic diagram of the data flow in this invention;

[0085] Figure 3 This is the intended feature extraction process of the present invention;

[0086] Figure 4 This is a schematic diagram of a voice signal edge processing system for audio monitoring according to the present invention. Detailed Implementation

[0087] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0088] Example 1

[0089] Please see Figure 1 , Figure 2 As shown, this embodiment provides a voice signal edge processing method for audio monitoring, including the following steps:

[0090] The sound pressure polarization vector flow of the sound pressure field, the vibration signal generated by the impact of the limb, and the sound pressure signal accompanied by the impact are collected by the edge device; the vibration signal and sound pressure signal can be directly acquired by the sensor; the data collected by the edge device has been preprocessed, which is a routine operation and will not be described in detail here.

[0091] Methods for obtaining acoustic pressure polarization vector flow include:

[0092] The sound pressure force components along the x, y, and z axes of preset characteristic scenes, such as classrooms, restrooms, and dormitories, are collected by edge devices.

[0093] The polarization angle can be calculated using trigonometric functions; for example, the polarization angle... ,in, and They are respectively Force components along the x-axis and y-axis at any given time;

[0094] Time series data describing the system's pressure measurements over time are acquired, preprocessed, and then arranged into a matrix to obtain a data matrix. and ,in, and yes matrix, The dimension of state variables, i.e., the number of force components. , The number of samples is represented by each column; each column represents the state at a given moment, and each row corresponds to a different displacement.

[0095] For data matrix Perform singular value decomposition. ,in, for Left singular vector matrix; for Singular value diagonal matrix, For matrix rank; for Right singular vector matrix; linear operators are calculated based on the decomposition results; such as linear operators ;

[0096] Eigenvalue decomposition is performed on the linear operator to obtain the eigenvalues. and the corresponding feature vector ;satisfy ,in, ;

[0097] From the left singular vector matrix and the corresponding feature vector The DMD modes are calculated, such as The system state is reconstructed based on DMD modes and eigenvalues ​​to predict the system's state in [the following context is missing]. The state at any given moment; such as ,in, The coefficients are set by the initial conditions; for The first moment One eigenvalue;

[0098] Obtain the imaginary part of the eigenvalues, based on the DMD mode. The phase angle is obtained by calculating the real part of the eigenvalues; for example, the phase angle... ;in, Eigenvalues The imaginary part; for The initial phase.

[0099] Vector synthesis of the force components along the x, y, and z axes is performed to obtain the sound pressure amplitude, such as... ,in, for The force component along the z-axis at time;

[0100] By splicing together the polarization angle, phase angle, and sound pressure amplitude, a sound pressure polarization vector flow is obtained.

[0101] The above steps involve collecting the sound pressure polarization vector flow of the sound pressure field, the vibration of limb impact, and the accompanying sound pressure signal. Edge devices are used to acquire the force components of sound pressure along the x, y, and z axes in scenarios such as classrooms and restrooms. The polarization angle is calculated using trigonometric functions, and the phase angle is obtained by processing time-series data through singular value decomposition and DMD modal analysis. The sound pressure amplitude is then obtained through vector synthesis and spliced ​​to form the sound pressure polarization vector flow. This process consistently focuses on physical characteristics unrelated to speech content in violent incidents—whether it's the spatial polarization pattern and dynamic changes of sound pressure generated by limb impact, or the vector synthesis results of sound pressure amplitude—all of which accurately capture the unique acoustic and vibrational characteristics of both ordinary violent behavior and non-verbal violent behavior. This avoids dependence on speech content to protect student privacy, and by extracting abnormal changes in these physical characteristics, it effectively improves the detection rate of non-verbal violent behavior, thus achieving accurate identification of violent incidents without infringing on privacy.

[0102] Please see Figure 3 As shown, feature extraction and fusion are performed on the sound pressure polarization vector flow, vibration signal, and sound pressure signal to obtain sound pressure polarization features and acoustic-mechanical coupling features. The sound pressure polarization features include polarization chaos, polarization angle chaos, and curl-divergence ratio. The acoustic-mechanical coupling features include vibration spectrum entropy, sound pressure energy abrupt gradient, and acoustic-mechanical coupling anomaly.

[0103] Methods for obtaining sound pressure polarization characteristics include:

[0104] Calculate the sound pressure amplitude of the sound pressure polarization vector flow, calculate the standard deviation and mean of the sound pressure amplitude, calculate the ratio of the standard deviation to the mean, and obtain the initial polarization chaos degree.

[0105] Calculate the first-order difference of the phase angle, such as First-order difference in statistical unit of time Number of times greater than the difference threshold According to the number and initial polarization chaos The polarization chaos degree is calculated; for example, the polarization chaos degree. ,in, This is a proportionality coefficient, typically set to 0.1, used to control the contribution of phase abrupt changes;

[0106] Calculate the time standard deviation of the polarization angle and the mean of the absolute values ​​of the polarization angles The polarization angle chaos degree is obtained by calculating the ratio of the time standard deviation to the mean of the absolute values ​​of the polarization angles; for example, the polarization angle chaos degree. ;

[0107] phase angle The magnitude of the sound pressure amplitude is divided into K intervals. All K sets of data are iterated through, and the polarization angle is counted at the [missing value]. Each phase angle divides the interval And the sound pressure modulus falls on the first Each sound pressure amplitude is divided into intervals Given the sample size, calculate the ratio of the sample size to K to obtain the joint probability distribution. According to the joint probability distribution Calculate the phase amplitude coupling entropy; such as ,in, The phase amplitude coupling entropy;

[0108] Calculate the curl and divergence of the sound pressure polarization vector flow, calculate the magnitudes of the curl and divergence respectively, and then calculate the ratio of the magnitude of the curl to the magnitude of the divergence to obtain the initial curl-divergence ratio. ;like ,in, Sound pressure polarization vector flow curl; Sound pressure polarization vector flow The divergence; Let be the magnitude of the vector; This is a local minimum value used to avoid a denominator of 0; it is used to calculate the spatial gradient of the polarization angle. modulus Such as the spatial gradient of the polarization angle The curl-divergence ratio is then calculated based on the initial curl-divergence ratio and the magnitude of the spatial gradient of the polarization angle; for example, the curl-divergence ratio... .

[0109] The method for obtaining sound pressure polarization characteristics obtains the initial polarization chaos by calculating the ratio of the standard deviation to the mean of the sound pressure amplitude. It then obtains the polarization chaos by combining the number of abrupt changes in the first-order difference of the phase angle with the scaling factor. The polarization angle chaos is obtained by the ratio of the time standard deviation of the polarization angle to the mean of its absolute value. The phase amplitude coupling entropy is calculated using the joint probability distribution of the phase angle and the sound pressure amplitude interval. Finally, the curl-divergence ratio is obtained by the ratio of the curl to the divergence magnitude of the sound pressure polarization vector flow and the magnitude of the spatial gradient of the polarization angle. These features are based on the physical properties of sound pressure rather than speech content. This avoids reliance on student speech information to protect privacy, while accurately capturing the disorder of sound pressure polarization state, the disorder of phase and amplitude coupling, and the abnormal changes in the curl-divergence relationship of the sound pressure field in both ordinary and non-verbal acts of violence. These unique physical differences improve the detection rate of non-verbal violent acts, enabling accurate identification of violent events independent of speech content.

[0110] Methods for obtaining acoustic-mechanical coupling characteristics include:

[0111] The vibration signal is filtered using a Hanning window with a window length of 1 / L (signal length), moving in steps of L / 2. A Fast Fourier Transform is performed on the signal within each window to obtain the power spectrum of each segment. The power spectrum of segment H is averaged to obtain the power spectral density. Based on the power spectral density, the vibration spectral entropy is calculated using information entropy. ,in, The power spectral density is obtained after the vibration signal undergoes FFT transformation.

[0112] Calculate the first derivative of the sound pressure signal, obtain its maximum value, and thus obtain the abrupt gradient of the sound pressure energy; for example... ,in, The first derivative of the sound pressure signal;

[0113] Calculate the angle of the vibration direction, and obtain directional consistency based on the angle and polarization angle; such as the angle of the vibration direction. Consistency in direction ,in, The vibration vector is located in the x-axis direction. The vibration vector is located in the y-axis direction.

[0114] Calculate the phase angle and the phase difference of the vibration signal, and obtain the synchronization based on the phase difference; such as the phase difference... synchronicity ,in, for The sound pressure phase angle at that moment; for The phase of the vibration signal at any given moment; The attenuation coefficient is used to adjust the intensity of the effect of the phase difference on synchronization. It is a preset constant that controls the rate at which synchronization changes with the phase difference.

[0115] The calculation method for the power spectral density of a sound pressure signal is similar to that for a vibration signal, and will not be repeated here.

[0116] The conjugate of the sound pressure signal is calculated based on its frequency domain representation. The cross-power spectrum of the sound pressure signal and the vibration signal is then calculated based on their respective frequency domain representations. ,in, for Conjugate; This is an element-wise multiplication method; the cross-power spectral density is calculated based on the cross-power spectrum, the power spectral density of the sound pressure signal, and the power spectral density of the vibration signal; for example... ,in, and These are the power spectral density of the sound pressure signal and the power spectral density of the vibration signal, respectively.

[0117] Calculate the time difference between the moment of abrupt change in the sound pressure phase angle and the moment of peak vibration spectrum, and obtain the time synchronization based on the time difference; such as time synchronization. ,in, It is the time difference between the moment of abrupt change in the sound pressure phase angle and the moment of peak vibration spectrum. This is an adjustment factor, typically set to a value of 10;

[0118] The acoustic pressure polarization features and acoustic-mechanical coupling features are used as inputs to the anomaly analysis model to obtain anomaly scores. For the acoustic pressure polarization features and acoustic-mechanical coupling features to be detected, the reconstruction features are reconstructed based on the output of the generative adversarial network, and the reconstruction error is calculated. The mean and standard deviation of the reconstruction error are calculated through the training set. The anomaly scores and reconstruction errors are weighted and fused to obtain the acoustic-mechanical coupling anomaly values.

[0119] Training methods for generative adversarial networks include:

[0120] This will contain Z sets of real data pairs. The training set data is divided according to a preset batch size, and one batch of fused multimodal data, namely the joint features of sound pressure polarization characteristics and acoustic coupling characteristics, is extracted for training each time; among them, It is the set of all acoustic pressure polarization features extracted from the acoustic pressure polarization vector flow; It is a collection of all acoustic-mechanical coupling features extracted and fused from vibration signals and sound pressure signals;

[0121] The combined features of acoustic pressure polarization and acoustic-mechanical coupling are used as input to the generator G to obtain reconstructed features; the discriminator determines the probability that the input features belong to the real data pairs.

[0122] A batch of real data is randomly selected from the training set. The selected real data and the generated data generated by the generator are used as inputs to the discriminator, and the discriminator outputs the corresponding discrimination probability. The loss function corresponding to the discriminator is designed. The network parameters of the discriminator are updated through the backpropagation algorithm.

[0123] The process continues until a preset optimal performance condition is reached, at which point the generative adversarial network (GAN) corresponding to the optimal performance condition is taken as the final output GAN.

[0124] The above steps involve filtering and Fourier transforming the vibration signal to obtain the power spectral density and calculating the vibration spectral entropy. The maximum value of the first derivative of the sound pressure signal is calculated to obtain the sound pressure energy mutation gradient. Directional consistency is obtained by combining the vibration direction angle and polarization angle. Synchronization is calculated through phase difference. The frequency domain correlation between sound pressure and vibration is analyzed based on the cross-power spectral density. Time synchronization is obtained based on the time difference between the sound pressure phase angle mutation and the vibration spectral peak. The sound pressure polarization characteristics and acoustic-mechanical coupling characteristics are then input into an anomaly analysis model to obtain an anomaly score. Anomaly values ​​for acoustic-mechanical coupling are obtained by weighted fusion of reconstruction errors from a generative adversarial network. All these processes revolve around the physical coupling characteristics of sound pressure and vibration, without involving the parsing of speech content. This avoids infringing on student privacy and accurately captures the abnormal coupling patterns of sound pressure and vibration in ordinary violent behavior and non-verbal violent behavior, such as energy mutations during impact, decreased directional consistency, and synchronization imbalance. Through comprehensive analysis of multi-dimensional coupling characteristics and anomaly value judgment, the detection rate of non-verbal violent behavior is effectively improved, achieving accurate identification of violent events without relying on speech content.

[0125] The acoustic pressure polarization characteristics and acoustic-mechanical coupling characteristics are used as inputs to a lightweight decision model to obtain event determination results;

[0126] Methods for obtaining event determination results include:

[0127] If the sound pressure polarization feature or the acoustic coupling feature is triggered, it is determined whether the preset combination conditions are met: at least any 2 items of the sound pressure polarization feature and any 3 items of the acoustic coupling feature are triggered, and it is not an interference scene, then it is determined to be a violent scene.

[0128] Methods for triggering acoustic pressure polarization characteristics or acoustic-mechanical coupling characteristics include:

[0129] The sound pressure polarization feature is triggered if any of the following conditions are met:

[0130] The polarization chaos degree is greater than the polarization dynamic threshold;

[0131] The polarization angle chaos degree is greater than the polarization angle threshold;

[0132] The curl-divergence ratio is greater than the ratio threshold;

[0133] The acoustic-mechanical coupling feature is triggered if any of the following conditions are met:

[0134] The vibration spectral entropy is greater than the dynamic threshold of spectral entropy.

[0135] The abrupt gradient of sound pressure energy exceeds the dynamic threshold of the gradient.

[0136] Directional consistency is below the dynamic consistency threshold;

[0137] Synchronization is below the dynamic threshold of synchronization;

[0138] Time synchronization is below the time dynamic threshold;

[0139] The abnormal value of acoustic-mechanical coupling exceeds the abnormal threshold.

[0140] Methods for obtaining a dynamic consistency threshold include:

[0141] The mean and standard deviation of directional consistency are calculated based on historical directional consistency data. A dynamic consistency threshold is then calculated based on the mean and standard deviation. (The dynamic consistency threshold is then used to calculate the mean and standard deviation.) ,in, The mean of directional consistency; Standard deviation of directional consistency; This is the consistency coefficient, which is typically set to 1.5.

[0142] Methods for obtaining the dynamic threshold of synchronization include:

[0143] The synchronization mean and standard deviation are calculated based on historical synchronization data. A dynamic synchronization threshold is then calculated based on the mean and standard deviation. (e.g., the dynamic synchronization threshold.) ,in, The mean of synchronicity; The standard deviation of synchronicity; This is the synchronization coefficient, which is typically set to 1.5.

[0144] Methods for obtaining time dynamic thresholds include:

[0145] An initial dynamic time threshold is defined based on historical time synchronization data, such as a confidence interval within the normal range of the scene (e.g., a 95% confidence interval). This threshold is then adjusted based on real-time noise levels and real-time time delay to obtain the final dynamic time threshold. ,in, This is a smoothing coefficient, typically set to 0.8. The initial time dynamic threshold; This represents the actual measured time delay between the sound pressure signal and the vibration signal at the current moment.

[0146] Methods for obtaining the polarization dynamic threshold include:

[0147] Real-time number of people and environmental noise are collected. Based on historical normal data, a probability density function is constructed using kernel density estimation. An initial threshold is set as the preset quantile of the probability density function, such as the 95th quantile. The ratio of real-time number of people to maximum capacity and the ratio of environmental noise to the noise limit are calculated. Combined with the initial threshold, a polarization dynamic threshold is calculated.

[0148] Methods for obtaining gradient dynamic thresholds include:

[0149] The sound pressure energy is calculated based on the sound pressure signal and air density. The mean and standard deviation of the current background sound pressure energy are calculated in real time using a sliding window of length Twin.

[0150] The mean of the current background sound pressure level is smoothed using an exponentially weighted moving average method to obtain a smoothed mean; for example, the smoothed mean... ,in, This represents the average value of the current background sound pressure energy. This represents the sound pressure energy calculated within the current sliding window.

[0151] The base threshold is calculated based on the smoothed mean and the standard deviation of the current background sound pressure energy, such as... ,in, The proportional coefficient of the base threshold can be obtained through natural heuristic optimization algorithms; The standard deviation of the current background sound pressure energy is used; the base threshold is adjusted based on the real-time number of people and time to obtain a gradient dynamic threshold; such as... ,in, The adjustment coefficient for the dynamic threshold can be obtained through natural heuristic optimization algorithms; Real-time number of users; The current time; For rectangle functions, It is set to 1 when the situation is noisy and 0 otherwise, and is used to increase the gradient dynamic threshold for noisy scenarios during breaks.

[0152] Methods for obtaining the dynamic threshold of spectral entropy include:

[0153] Perform a short-time Fourier transform on the vibration signal to obtain the spectrum matrix, and extract the frequency components.

[0154] Calculate the energy percentage of each frequency component, such as the first... Energy percentage at each frequency point ,in, For the first The amplitude at each frequency point; For the first The amplitude at each frequency point; The frequency point number; Shannon entropy is calculated based on the energy percentage; e.g., Shannon entropy. ;

[0155] Beforehand, K-means clustering analysis is used to analyze the vibration spectrum entropy distribution of different locations, generating a mapping relationship between location type and baseline entropy. The baseline entropy corresponding to the current location is obtained, and the updated baseline entropy is obtained based on the baseline entropy corresponding to the current location. ,in, The baseline value of the vibration spectrum entropy before the update; The coefficients for baseline updates can be obtained through natural heuristic optimization algorithms. The real-time vibration spectrum entropy is calculated based on the vibration signal at the current location at the current moment.

[0156] The dynamic threshold of spectral entropy is calculated based on the baseline entropy, the updated baseline entropy, and the standard deviation. For example... ,in, Baseline entropy; The adjustment coefficients for calculating the dynamic threshold of spectral entropy can be obtained through optimization using a natural heuristic algorithm. This represents the standard deviation of the cluster corresponding to the location.

[0157] Methods for generating the mapping relationship between location type and baseline entropy include:

[0158] The number of clusters is set to be equal to the number of site types, and Euclidean distance is used to measure the dispersion of the vibration spectrum entropy;

[0159] The dataset is grouped by location, and K-means clustering is performed on the vibration spectrum entropy subset for each location;

[0160] Output the vibration spectrum entropy corresponding to the cluster center of each location as the baseline entropy; at the same time, record the standard deviation of the cluster corresponding to the location.

[0161] In the process of inputting sound pressure polarization features and acoustic coupling features into a lightweight decision model to obtain event judgment results, the triggering conditions of sound pressure polarization features and acoustic coupling features are clearly defined. Judgment is made by combining the preset combination condition of "at least 2 sound pressure polarization features + 3 acoustic coupling features triggered in a non-interference scenario". At the same time, the dynamic threshold of each feature can be adapted to the normal baseline of different scenarios. The whole process is always based on the physical characteristics of sound pressure and vibration rather than speech content. This avoids the invasion of students' privacy and effectively captures the unique abnormal combination of physical features in ordinary violent behavior and non-verbal violent behavior through multi-feature collaborative triggering and accurate adaptation of dynamic thresholds. This reduces misjudgment caused by environmental differences, thereby improving the detection rate of non-verbal violent behavior and achieving accurate identification of violent events that do not depend on speech content.

[0162] If the event is determined to be a violent event, the audio recordings of W seconds before and after the determination are encrypted and stored, and an alarm is simultaneously sent to the teacher's terminal and the security system, along with the 3D location information. If the event is determined to be normal, no intervention is required, and all collected and processed data can be discarded.

[0163] When the event is determined to be a violent event, the sound pressure polarization characteristics and sound force coupling characteristics are stored. A dynamic key can be generated through the PUF physical non-cloning function. The feature vector is encrypted using the national cryptographic SM4 algorithm to obtain encrypted data. The encrypted data is stored in fragments on the edge device. The key is managed by the school and the education bureau in a two-factor manner.

[0164] The above steps, without involving the analysis of voice content, provide a basis for subsequent handling by responding to the results of violent incident determination through encrypted storage of recordings and accurate push of location information. This avoids unnecessary intervention in normal scenarios. At the same time, since the entire determination is based on the physical characteristics of sound pressure and vibration, it can effectively cover a variety of violent behaviors. Thus, while protecting student privacy, timely response further ensures the effectiveness of violent incident identification, helps improve the efficiency of handling non-verbal violent behaviors after detection, and ultimately serves the technical goal of accurately identifying violent incidents without relying on voice content.

[0165] Example 2

[0166] Please see Figure 4 As shown, this embodiment provides a voice signal edge processing system for audio monitoring, including:

[0167] Data acquisition module: Acquires sound pressure polarization vector flow, vibration signal, and sound pressure signal through edge devices;

[0168] Feature extraction module: Extracts and fuses features from sound pressure polarization vector flow, vibration signal and sound pressure signal to obtain sound pressure polarization features and acoustic-mechanical coupling features; sound pressure polarization features include polarization chaos degree, polarization angle chaos degree and curl-divergence ratio; acoustic-mechanical coupling features include vibration spectrum entropy, sound pressure energy abrupt gradient and acoustic-mechanical coupling anomaly value;

[0169] Event determination module: It takes the sound pressure polarization characteristics and acoustic coupling characteristics as inputs to the lightweight decision model to obtain the event determination results;

[0170] Event handling module: If the event is determined to be a violent event, an alarm will be triggered, and the audio recordings before and after the alarm (W seconds) will be encrypted and stored. Simultaneously, the 3D location information will be pushed to the teacher's end and the security system. If the event is determined to be normal, no intervention will be performed.

[0171] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

[0172] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for edge processing of voice signals for audio monitoring, characterized in that, Includes the following steps: Acquire sound pressure polarization vector flow, vibration signal, and sound pressure signal; Feature extraction and fusion are performed on the sound pressure polarization vector flow, vibration signal, and sound pressure signal to obtain sound pressure polarization features and acoustic-mechanical coupling features. The sound pressure polarization features include polarization chaos, polarization angle chaos, and curl-divergence ratio. The acoustic-mechanical coupling features include vibration spectral entropy, sound pressure energy abrupt gradient, and acoustic-mechanical coupling anomaly. The acoustic pressure polarization characteristics and acoustic-mechanical coupling characteristics are used as inputs to a lightweight decision model to obtain event determination results; If the event is determined to be a violent event, the audio recording of W seconds before and after the determination is encrypted and stored, and an alarm is simultaneously sent to the teacher's terminal and the security system, and 3D location information is pushed; if the event is determined to be normal, no intervention is performed.

2. The voice signal edge processing method for audio monitoring according to claim 1, characterized in that, Methods for obtaining acoustic pressure polarization vector flow include: The sound pressure force components in the x, y, and z axes of a preset characteristic scene are collected by an edge device. The polarization angle is calculated using trigonometric functions. Time series data describing the system's pressure measurements over time are acquired, preprocessed, and then arranged into a matrix to obtain a data matrix. and ,in, and yes matrix, This refers to the dimension of state variables, i.e., the number of force components. The number of samples is represented by each column; each column represents the state at a given moment, and each row corresponds to a different displacement. For data matrix Perform singular value decomposition, decomposing it into a left singular vector matrix, a singular value diagonal matrix, and a right singular vector matrix; based on the left singular vector matrix, the singular value diagonal matrix, and the data matrix... The linear operator is obtained through computation; Eigenvalue decomposition is performed on a linear operator to obtain eigenvalues ​​and corresponding eigenvectors; The DMD mode is obtained by calculating the left singular vector matrix and the corresponding eigenvector. The system state is reconstructed based on the DMD mode and eigenvalues ​​to predict the system state at future time. Obtain the imaginary part of the eigenvalues, and calculate the phase angle based on the DMD mode and the real part of the eigenvalues; The sound pressure amplitude is obtained by vector synthesis of the force components along the x, y, and z axes. By splicing together the polarization angle, phase angle, and sound pressure amplitude, a sound pressure polarization vector flow is obtained.

3. The method for edge processing of voice signals for audio monitoring according to claim 1, characterized in that, Methods for obtaining sound pressure polarization characteristics include: Calculate the sound pressure amplitude of the sound pressure polarization vector flow, calculate the standard deviation and mean of the sound pressure amplitude, calculate the ratio of the standard deviation to the mean, and obtain the initial polarization chaos degree. Calculate the first-order difference of the phase angle, count the number of times the first-order difference exceeds the difference threshold per unit time, and calculate the polarization chaos degree based on the number of times and the initial polarization chaos degree. Calculate the time standard deviation and the mean of the absolute value of the polarization angle, and calculate the ratio of the time standard deviation to the mean of the absolute value of the polarization angle to obtain the polarization angle chaos degree; The phase angle and sound pressure amplitude are divided into K intervals. All K sets of data are iterated through, and the polarization angle is statistically analyzed on the [Kth]th interval. Each phase angle divides the interval And the sound pressure modulus falls on the first Each sound pressure amplitude is divided into intervals The number of samples is determined, the ratio of the number of samples to K is calculated, the joint probability distribution is obtained, and the phase amplitude coupling entropy is calculated based on the joint probability distribution. Calculate the curl and divergence of the sound pressure polarization vector flow, calculate the magnitudes of the curl and divergence respectively, and then calculate the ratio of the magnitude of the curl to the magnitude of the divergence to obtain the initial curl-divergence ratio; calculate the magnitude of the spatial gradient of the polarization angle, and then calculate the curl-divergence ratio based on the initial curl-divergence ratio and the magnitude of the spatial gradient of the polarization angle.

4. The method for edge processing of voice signals for audio monitoring according to claim 1, characterized in that, Methods for obtaining acoustic-mechanical coupling characteristics include: The vibration signal is filtered by using a Hanning window with a window length of 1 / L of the signal length, moving in steps of L / 2. A fast Fourier transform is performed on the signal within each window to obtain the power spectrum of each segment. The power spectrum of segment H is averaged to obtain the power spectral density. Based on the power spectral density, the vibration spectral entropy is calculated using information entropy. Calculate the first derivative of the sound pressure signal, obtain the maximum value of the first derivative, and obtain the abrupt gradient of the sound pressure energy. Calculate the angle of the vibration direction, and obtain the directional consistency based on the angle and polarization angle; Calculate the phase angle and the phase difference of the vibration signal, and obtain the synchronization based on the phase difference; The power spectral density of the sound pressure signal is calculated. The conjugate of the sound pressure signal is calculated based on its frequency domain representation. The cross power spectrum of the sound pressure signal and the vibration signal is calculated based on their conjugate and frequency domain representations. The cross power spectral density is then calculated based on the cross power spectrum, the power spectral density of the sound pressure signal, and the power spectral density of the vibration signal. Calculate the time difference between the moment of abrupt change in the sound pressure phase angle and the moment of peak vibration spectrum, and obtain the time synchronization based on the time difference; The acoustic pressure polarization features and acoustic-mechanical coupling features are used as inputs to the anomaly analysis model to obtain anomaly scores. For the acoustic pressure polarization features and acoustic-mechanical coupling features to be detected, the reconstruction features are reconstructed based on the output of the generative adversarial network, and the reconstruction error is calculated. The mean and standard deviation of the reconstruction error are calculated through the training set. The anomaly scores and reconstruction errors are weighted and fused to obtain the acoustic-mechanical coupling anomaly values.

5. The voice signal edge processing method for audio monitoring according to claim 4, characterized in that, The training method for the generative adversarial network includes: This will contain Z sets of real data pairs. The training set data is divided according to a preset batch size, and one batch of fused multimodal data, namely the joint features of sound pressure polarization characteristics and acoustic coupling characteristics, is extracted for training each time; among them, It is the set of all acoustic pressure polarization features extracted from the acoustic pressure polarization vector flow; It is a collection of all acoustic-mechanical coupling features extracted and fused from vibration signals and sound pressure signals; The combined features of acoustic pressure polarization and acoustic-mechanical coupling are used as input to the generator G to obtain reconstructed features; the discriminator determines the probability that the input features belong to the real data pairs. A batch of real data is randomly selected from the training set. The selected real data and the generated data generated by the generator are used as inputs to the discriminator, and the discriminator outputs the corresponding discrimination probability. The loss function corresponding to the discriminator is designed. The network parameters of the discriminator are updated through the backpropagation algorithm. The process continues until a preset optimal performance condition is reached, at which point the generative adversarial network (GAN) at which the preset optimal performance condition is reached is taken as the final output GAN.

6. The method for edge processing of voice signals for audio monitoring according to claim 1, characterized in that, Methods for obtaining event determination results include: If the sound pressure polarization feature or the acoustic coupling feature is triggered, it is determined whether the preset combination conditions are met: at least any 2 items of the sound pressure polarization feature and any 3 items of the acoustic coupling feature are triggered, and it is not an interference scene, then it is determined to be a violent scene; Methods for triggering acoustic pressure polarization characteristics or acoustic-mechanical coupling characteristics include: The sound pressure polarization feature is triggered if any of the following conditions are met: The polarization chaos degree is greater than the polarization dynamic threshold; The polarization angle chaos degree is greater than the polarization angle threshold; The curl-divergence ratio is greater than the ratio threshold; The acoustic-mechanical coupling feature is triggered if any of the following conditions are met: The vibration spectral entropy is greater than the dynamic threshold of spectral entropy. The abrupt gradient of sound pressure energy exceeds the dynamic threshold of the gradient. Directional consistency is below the dynamic consistency threshold; Synchronization is below the dynamic threshold of synchronization; Time synchronization is below the time dynamic threshold; The abnormal value of acoustic-mechanical coupling exceeds the abnormal threshold.

7. The method for edge processing of voice signals for audio monitoring according to claim 6, characterized in that, Methods for obtaining a dynamic consistency threshold include: The mean and standard deviation of directional consistency are calculated based on historical directional consistency data, and the dynamic threshold of consistency is obtained based on the mean and standard deviation. Methods for obtaining the dynamic threshold of synchronization include: The synchronization mean and synchronization standard deviation are calculated based on historical synchronization data, and the synchronization dynamic threshold is obtained based on the mean and standard deviation. Methods for obtaining time dynamic thresholds include: An initial dynamic time threshold is defined based on historical time synchronization data, and then corrected according to real-time noise level and real-time time delay to obtain the dynamic time threshold.

8. A method for edge processing of voice signals for audio monitoring according to claim 6, characterized in that, Methods for obtaining the polarization dynamic threshold include: Based on historical normal data, a probability density function is constructed using kernel density estimation, and an initial threshold is set as the preset quantile of the probability density function. The ratio of real-time number of people to maximum capacity and the ratio of environmental noise to the noise limit are calculated. Combined with the initial threshold, the polarization dynamic threshold is calculated.

9. A method for edge processing of voice signals for audio monitoring according to claim 6, characterized in that, Methods for obtaining gradient dynamic thresholds include: The sound pressure energy is calculated based on the sound pressure signal and air density. The mean and standard deviation of the current background sound pressure energy are calculated in real time using a sliding window of length Twin. The mean value of the current background sound pressure energy is smoothed by using an exponentially weighted moving average to obtain a smoothed mean. The base threshold is calculated based on the smoothed mean and the standard deviation of the current background sound pressure energy. The base threshold is then adjusted based on the real-time number of people and time to obtain the gradient dynamic threshold.

10. A method for edge processing of voice signals for audio monitoring according to claim 6, characterized in that, Methods for obtaining the dynamic threshold of spectral entropy include: Perform a short-time Fourier transform on the vibration signal to obtain the spectrum matrix, and extract the frequency components. Calculate the energy percentage of each frequency component, and obtain the Shannon entropy based on the energy percentage. Beforehand, K-means clustering analysis is used to analyze the vibration spectrum entropy distribution of different locations, generating a mapping relationship between location type and baseline entropy. The baseline entropy corresponding to the current location is obtained, and the updated baseline entropy is obtained based on the baseline entropy corresponding to the current location. The dynamic threshold of spectral entropy is calculated based on the baseline entropy, the updated baseline entropy, and the standard deviation.

11. The method for edge processing of voice signals for audio monitoring according to claim 1, characterized in that, Methods for generating the mapping relationship between location type and baseline entropy include: The number of clusters is set to be equal to the number of site types, and Euclidean distance is used to measure the dispersion of the vibration spectrum entropy; The dataset is grouped by location, and K-means clustering is performed on the vibration spectrum entropy subset for each location; Output the vibration spectrum entropy corresponding to the cluster center of each location as the baseline entropy; at the same time, record the standard deviation of the cluster corresponding to the location.

12. A voice signal edge processing system for audio monitoring, implementing the voice signal edge processing method for audio monitoring as described in any one of claims 1-11, characterized in that, include: Data acquisition module: Acquires sound pressure polarization vector flow, vibration signal, and sound pressure signal through edge devices; Feature extraction module: Extracts and fuses features from sound pressure polarization vector flow, vibration signal and sound pressure signal to obtain sound pressure polarization features and acoustic-mechanical coupling features; sound pressure polarization features include polarization chaos degree, polarization angle chaos degree and curl-divergence ratio; acoustic-mechanical coupling features include vibration spectrum entropy, sound pressure energy abrupt gradient and acoustic-mechanical coupling anomaly value; Event determination module: It takes the sound pressure polarization characteristics and acoustic coupling characteristics as inputs to the lightweight decision model to obtain the event determination results; Event handling module: If the event is determined to be a violent event, an alarm will be triggered, and the audio recordings before and after the alarm (W seconds) will be encrypted and stored. Simultaneously, the 3D location information will be pushed to the teacher's end and the security system. If the event is determined to be normal, no intervention will be performed.

Citation Information

Patent Citations

  • Campus violence early warning system

    CN118761867A

  • Indoor emergent abnormal event alarm system

    CN103198605A

  • Sound-vibration fusion signal identification method and system, computer equipment and medium

    CN117292494A