Campus hidden area abnormal voice sensing and classification recognition system

By using a distributed acoustic sensor array and an integrated processing pipeline, the contradiction between high sensitivity and low false alarm rate in the voice monitoring system for concealed areas of the campus is resolved, achieving efficient abnormal voice recognition and privacy protection, thus meeting the needs of campus security and privacy.

CN121811918APending Publication Date: 2026-04-07DINGDIAN TECHNOLOGY (SUQIAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-12
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies in voice monitoring systems for concealed areas on campus present a contradiction between high-sensitivity voice sensing and low false alarm rate classification and recognition, and are difficult to meet the needs of privacy protection and sustainable system deployment.

Method used

It employs a distributed acoustic sensor array, an environmental context awareness module, an adaptive signal conditioning module, a sound source credibility assessment module, and a privacy-preserving data processing module. Through a unified time synchronization bus and a secure communication link, it forms an integrated processing pipeline to achieve dynamic collaboration and privacy protection.

Benefits of technology

It effectively reduces the false alarm rate, ensures that alarms are generated only when acoustic features, semantic content, and emotional state are consistent, protects privacy, reduces communication and storage overhead, and meets the security and privacy requirements of the campus environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121811918A_ABST
    Figure CN121811918A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence and intelligent security and protection, and particularly relates to a campus hidden area abnormal voice sensing and classification recognition system. Comprising a distributed acoustic sensing array, an environment context sensing module, an adaptive signal conditioning module, a sound source credibility evaluation module, a context enhanced abnormal voice classification module and a privacy protection type data processing module. Through a multi-mode environment perception and dynamic feedback coupling mechanism, sound source space verification, adaptive gain control and multi-level classification decision are realized, and through an adaptive pickup mechanism driven by an environment context, excessive acquisition of environment noise in an unmanned state is avoided; therefore, false triggering caused by non-human-body sound sources can be eliminated by utilizing consistency verification of sound source spatial orientation and human-body existence information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence and intelligent security technology, specifically a system for sensing and classifying abnormal voices in concealed areas of a campus. Background Technology

[0002] In the process of building smart campuses, campus security systems are evolving from traditional human and physical defenses to intelligent and perceptive ones. Early detection and intelligent warning of abnormal behavior in hidden areas such as stairwells of teaching buildings and secluded passages in dormitories have become key to improving the comprehensive governance capabilities of campuses. As a natural accompanying signal of human behavior, speech, with its emotional characteristics, semantic content, and acoustic patterns, can serve as a highly sensitive basis for identifying abnormal events such as bullying and calls for help. Therefore, constructing a highly sensitive sensing and low false alarm rate classification and recognition system for speech signals in hidden areas has significant practical significance and application value.

[0003] Current mainstream campus voice monitoring solutions adopt a technical approach that combines general speech recognition (ASR) with keyword triggering. They collect environmental audio streams through fixed microphone arrays or microphones and rely on a preset sensitive word database for matching and judgment. Some solutions introduce deep neural network acoustic event classification models to filter non-verbal acoustic events such as crying and glass breaking. Such technologies have certain practical applications in open or semi-open scenarios. They are derived from the traditional security post-event retrospective and keyword alarm paradigms, effectively alleviating the problems of large blind spots and slow response of manual inspections, and have engineering feasibility.

[0004] When existing technologies are migrated to concealed areas on campus, a fundamental contradiction arises between high-sensitivity speech sensing and low-false-alarm-rate classification and recognition. Concealed areas have strong reverberation, low signal-to-noise ratio, and complex background noise, leading to speech signal distortion. To avoid missing weak abnormal speech, increasing sensor gain or expanding the pickup range introduces a large amount of interference. Static thresholds or offline-trained general recognition models are prone to misclassifying normal sounds as abnormal. This contradiction stems from the lack of a dynamic coordination mechanism between sensing and recognition. The full capture at the sensing end is disconnected from the static criteria at the recognition end. Even with higher-precision speech recognition models (such as those based on temporal convolutional networks or Transformer architectures), the input data itself is already mixed with a large amount of invalid or even misleading information due to front-end oversampling, causing a sharp decline in the model's generalization ability. Furthermore, due to the extremely high privacy protection requirements of the campus environment, the system usually cannot store the original audio for long-term verification. Once a false alarm occurs, it is difficult to trace and verify, and it may cause unnecessary panic or waste of management resources, thereby weakening the system's credibility and sustainable deployment capability.

[0005] Therefore, the present invention provides an abnormal voice sensing and classification recognition system for hidden areas on campus. Summary of the Invention

[0006] In order to overcome the shortcomings of the prior art, at least one technical problem raised in the background art is solved.

[0007] The technical solution adopted by the present invention to solve its technical problem is as follows: The present invention provides an abnormal voice sensing and classification system for concealed areas on campus, which includes a distributed acoustic sensing array module, an environmental context perception module, an adaptive signal conditioning module, a sound source credibility assessment module, a context-enhanced abnormal voice classification module, and a privacy-preserving data processing module. Each module is interconnected with a unified time synchronization bus and a secure communication link to form an integrated processing pipeline with status feedback and parameter coordination capabilities.

[0008] Preferably, the distributed acoustic sensor array module is deployed at key locations in the concealed area of ​​the target, and consists of multiple directional microphone units. Each microphone unit integrates an independent preamplifier circuit and an analog-to-digital converter, and is connected to the local edge computing node via power line carrier. The array adopts an asymmetric spatial arrangement strategy, so that the main axis of adjacent microphone units is offset by a preset angle, thereby forming a spatial suppression capability against sound sources in non-target directions at the physical level. All microphone units synchronously sample audio signals and encapsulate the raw digital audio stream into data frames with timestamp tags, which are then uploaded to the adaptive signal conditioning module via a secure communication protocol conforming to the IEC61850 extended specification.

[0009] Preferably, the environmental context awareness module is configured within the same edge computing node and includes an inertial measurement unit, a temperature and humidity sensor, an infrared human presence detector, and a background noise spectrum analysis unit. This module collects environmental physical state parameters at fixed intervals and models the frequency domain energy distribution of the continuously input environmental background noise based on short-time Fourier transform, generating a multi-dimensional context feature vector of the current scene. This feature vector is output in real time to the adaptive signal conditioning module and the sound source credibility assessment module as the basis for adjusting the sound pickup strategy and determining the effectiveness of the sound source.

[0010] Preferably, the adaptive signal conditioning module receives the raw audio frames from the distributed acoustic sensor array module and the context feature vector from the environmental context awareness module, and performs dynamic gain control and frequency band selective filtering. Specifically, the module has a built-in state machine controller that automatically switches the working mode according to the current ambient noise energy level and the presence of a human: when the infrared human presence detector is not triggered and the background noise energy is lower than a preset safety threshold, the system enters a low-power monitoring mode, only enabling the narrowband filter in the center frequency band and maintaining the basic gain; when the presence of a human is detected or the noise energy suddenly increases, the state machine jumps to a high-sensitivity capture mode, activates the full-band filter and increases the gain coefficient, and simultaneously starts the beamforming algorithm of the microphone array to focus on the direction of the most recent human infrared signal; all adjustments to the gain and filtering parameters are linked with the sound source credibility assessment module through a feedback loop to ensure that the sound pickup intensity always matches the probability of anomalies in the current scene.

[0011] Preferably, the sound source credibility assessment module receives the audio stream processed by the adaptive signal conditioning module and the corresponding context feature vector, and performs sound source authenticity judgment. This module first calculates the time difference features between each microphone channel through cross-correlation analysis, and calculates the spatial azimuth of the sound source by combining the geometric layout of the microphone array. Then, it performs consistency verification between the azimuth and the positioning information of the infrared human presence detector. If the deviation between the two exceeds the allowable range, the sound source is determined to be non-human. Further, this module extracts three acoustic features of the audio signal: fundamental frequency stability, harmonic structure integrity, and instantaneous energy mutation rate, and performs dynamic matching with a pre-stored typical abnormal speech template library. If the matching score is lower than the convergence judgment condition, and the sound source azimuth is inconsistent with the human body position, the audio segment is marked as low credibility data and prohibited from entering the subsequent classification process. Otherwise, a high-confidence sound source identifier is generated and attached to the header of the audio data frame for use by the context-enhanced abnormal speech classification module.

[0012] Preferably, the context-enhanced abnormal speech classification module is deployed in a trusted execution environment on an edge computing node. Its input consists of high-confidence audio frames filtered by the sound source confidence assessment module and their associated contextual feature vectors. This module employs a hierarchical decision architecture. The first layer is a coarse acoustic event classifier, which uses a convolutional neural network to perform binary discrimination between non-linguistic acoustic events (such as screams, cries, and broken glass) and linguistic speech in the input audio. The second layer is a semantic-emotion joint analyzer, which simultaneously performs keyword matching and emotion state recognition for linguistic speech. The keyword matching uses an encrypted dictionary structure, completing the process without decrypting the original speech. The first layer verifies the existence of sensitive words, while the second layer identifies the emotion state by fusing speech prosody features with time, location, and population density information in the context through an attention mechanism, and outputs the probability of abnormal emotion. The third layer is a context fusion decision unit, which integrates the output results of the first two layers, the current time period (e.g., night / break time), the area type (e.g., stairwell / experiment corridor), and the historical alarm frequency, and generates the final abnormal event judgment signal through a weighted logic gate circuit. This judgment signal is activated only when multiple conditions are met, including but not limited to: the acoustic event category belongs to a preset abnormal set, the probability of abnormal emotion is higher than the dynamic threshold, and there are no frequent false alarm records in the current area recently.

[0013] Preferably, the privacy-preserving data processing module is responsible for end-to-end security control of the data flow throughout the system's entire lifecycle; the raw audio data is immediately destroyed after feature extraction is completed by the adaptive signal conditioning module, retaining only irreversible acoustic feature vectors and intermediate classification results; all data transmitted to the central management platform is homomorphically encrypted to ensure that the voice content cannot be restored even if intercepted on the network; in addition, the module has a built-in access control policy engine that dynamically allocates data viewing permissions based on user roles, and ordinary security personnel can only receive abnormal event alarm summaries and cannot obtain any raw or intermediate audio data; the system log is stored using a blockchain-style hash chain structure to ensure the immutability and auditability of operation records.

[0014] Preferably, each microphone unit in the distributed acoustic sensor array module interacts with the local edge computing node through a shared buffer in dual-port RAM. The DMA controller triggers an interrupt when it detects a sudden change in audio energy, thereby achieving event-driven low-latency data upload and avoiding resource waste caused by continuous polling.

[0015] Preferably, the background noise spectrum analysis unit in the environmental context awareness module uses a sliding window mechanism to perform frequency domain energy integration on continuous audio segments. When the energy of a certain frequency band is stably higher than the baseline within multiple consecutive windows, it is marked as a steady-state interference source, and the adaptive signal conditioning module is notified to perform notch filtering on that frequency band in subsequent processing.

[0016] Preferably, when performing sound source azimuth calculation, the sound source reliability assessment module introduces equipment attitude angle compensation provided by the inertial measurement unit to eliminate geometric model inaccuracies caused by the tilt of the mounting surface, thereby ensuring spatial positioning accuracy.

[0017] Preferably, the encrypted dictionary tree structure in the context-enhanced abnormal speech classification module uses a combination of Bloom filters and homomorphic encryption to verify the existence of keywords, ensuring both query efficiency and preventing the leakage of dictionary content.

[0018] Preferably, the privacy-preserving data processing module maintains a rolling feature cache queue locally at the edge node. Only when an abnormal event judgment signal is activated will the feature vectors for the corresponding time period be packaged and uploaded. Data for other time periods will be automatically cleared before being overwritten by the cache, thus preventing unnecessary data retention from the source.

[0019] The beneficial effects of this invention are as follows: This invention discloses an abnormal voice sensing and classification system for concealed areas on campus. Through an environment-context-driven adaptive sound pickup mechanism, it avoids excessive collection of environmental noise in unoccupied environments. It utilizes consistency verification between the spatial location of the sound source and the presence of a human to eliminate false triggers caused by non-human sound sources. It achieves deep coupling between high-confidence sound source screening and multi-level classification decisions, ensuring that alarms are generated only when acoustic features, semantic content, emotional state, and scene context all point to an anomaly. It ensures that raw voice data is destroyed immediately after feature extraction, preventing sensitive information from remaining within the system for extended periods. Furthermore, through an event-driven data upload strategy and rolling cache management, it effectively reduces communication load and storage overhead. Simultaneously, by leveraging a trusted execution environment and homomorphic encryption technology, it ensures the privacy and security of the classification process and data transmission, meeting the stringent requirements for personal information protection in campus settings. Attached Figure Description

[0020] The invention will now be further described with reference to the accompanying drawings.

[0021] Figure 1 This is a schematic diagram of the overall architecture of an abnormal voice sensing and classification system for concealed areas on campus according to the present invention. Figure 2 This is a schematic diagram of the distributed acoustic sensor array module structure in this invention. Detailed Implementation

[0022] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention will be further described below in conjunction with specific embodiments.

[0023] like Figure 1As shown in the embodiment of the present invention, an abnormal voice sensing and classification system for a hidden area on campus comprises a distributed acoustic sensor array module, an environmental context awareness module, an adaptive signal conditioning module, a sound source credibility assessment module, a context-enhanced abnormal voice classification module, and a privacy-preserving data processing module. Each module is interconnected with a secure communication link conforming to the IEC 61850 extended specification via a unified time synchronization bus, forming an integrated processing pipeline with status feedback and parameter coordination capabilities.

[0024] The following will describe in detail the various components of the present invention and their collaborative working mechanism, in conjunction with the accompanying drawings and specific engineering implementation details.

[0025] In some implementations, the distributed acoustic sensor array module is deployed in key locations in concealed areas such as stairwells, laboratory corridors, and equipment storage rooms on campus. As mentioned earlier, the module consists of no fewer than three directional microphone units, each of which uses a supercardioid condenser pickup head. Each microphone unit integrates an independent preamplifier circuit and is equipped with a 24-bit Σ-Δ analog-to-digital converter. All microphone units are connected to the local edge computing node via power line carrier (PLC) technology, with the carrier frequency set at 150kHz, the modulation method being OFDM, and the effective data transmission rate being 2.5Mbps. As mentioned above, the array adopts an asymmetric spatial arrangement strategy, and the main axis pointing of adjacent microphone units is offset by a preset angle, the offset angle θ satisfying Where N is the total number of microphones, k is the integer index (k=0,1,…,N-1), and δ is a random perturbation term with a value range of ±5° to disrupt the grating lobe effect caused by the periodic structure. All microphone units achieve sub-microsecond synchronous sampling through the IEEE1588 precision time protocol. The raw digital audio stream is encapsulated into data frames with timestamp tags, the timestamp accuracy is better than ±100ns, and uploaded to the adaptive signal conditioning module via a secure communication protocol.

[0026] Furthermore, each microphone unit interacts with the local edge computing node through a shared buffer of dual-port RAM. This dual-port RAM has a capacity of 64kB and is divided into two 32kB circular buffers, which are used to store the audio data of the current sampling window and the previous window, respectively. The DMA controller continuously monitors changes in audio energy. When it detects a short-term energy surge (with a window length of 20ms) in any channel that exceeds the static baseline by more than 3dB and lasts for more than 50ms, it triggers an interrupt signal and starts an event-driven data upload mechanism. Understandably, this mechanism avoids the waste of CPU resources caused by continuous polling, and actual tests show that it can reduce the average power consumption of edge nodes by 42% in a typical campus environment.

[0027] In some implementations, the environmental context awareness module is configured within the same edge computing node and includes a six-axis inertial measurement unit (IMU), a digital temperature and humidity sensor, a passive infrared human presence detector, and a background noise spectrum analysis unit. It should be noted that the IMU model is BMI270, which provides triaxial acceleration (range ±16g) and triaxial angular velocity (range ±2000° / s, resolution 0.06° / s) outputs; It should also be noted that the temperature and humidity sensor is model SHT45, with a temperature measurement range of -40°C to 125°C and an accuracy of ±0.1°C; the relative humidity measurement range is 0% to 100%RH and an accuracy of ±1.8%RH. The infrared human presence detector uses a dual-element pyroelectric sensor with a Fresnel lens, with a detection distance of 0.5m to 8m, a field of view of 110°, and a response time of less than 1s. The background noise spectrum analysis unit uses a 100ms sliding window to perform a short-time Fourier transform (STFT) on continuous audio segments from any reference channel in the distributed acoustic sensor array module. The window function used is a Hanning window, with 2048 FFT points and a frequency resolution of 23.4Hz. This unit calculates the average energy of each frequency band (divided into 1 / 3 octave bands, with center frequencies from 63Hz to 8kHz) and compares it with the baseline energy distribution of the same time period over the past 7 days, generating a multi-dimensional context feature vector. Where m=24 (corresponding to 8 1 / 3 octave band × 3 statistical measures: mean, variance, and mutation rate), each component is normalized and then input to the adaptive signal conditioning module and the sound source credibility assessment module.

[0028] In some implementations, the background noise spectrum analysis unit uses a sliding window mechanism to perform frequency domain energy integration on continuous audio segments; Specifically, the energy of the i-th frequency band in the t-th window is defined as... ,in For STFT coefficients, This is the set of frequency points corresponding to the i-th 1 / 3 octave band. As mentioned earlier, if there exists a frequency band i such that It holds true within a consecutive L=5 windows. , If the mean and standard deviation of the frequency band over the past 24 hours are respectively, it is marked as a steady-state interference source, and the adaptive signal conditioning module is notified through the control bus to perform notch filtering on the frequency band in subsequent processing, with a notch depth of not less than 20dB and a bandwidth of ±100Hz.

[0029] In some implementations, the adaptive signal conditioning module receives raw audio frames from the distributed acoustic sensor array module and context feature vectors from the environmental context awareness module, performs dynamic gain control and band-selective filtering, and incorporates a finite state machine controller whose state transition logic is determined by two key inputs: the binary output of the infrared human presence detector. Total energy of background noise Preset safety threshold dBFS (relative to full scale); It should be noted that when and When this happens, the system enters a low-power monitoring mode (State_LowPower). Following the above, when or When this happens, the state machine transitions to the high-sensitivity capture mode (State_HighSensitivity). Under State_LowPower, only the narrowband Butterworth filter (order 4, passband ripple 0.1dB) in the center frequency band (500Hz-4kHz) is enabled, while maintaining the base gain. dB; Under State_HighSensitivity, activate the full-band (100Hz-8kHz) filter and increase the gain to dB, and simultaneously activate the delay-summation beamforming algorithm of the microphone array to focus on the spatial orientation recorded at the most recent infrared signal trigger. Beamforming weight vector Calculated by the following formula:

[0030] Where N is the number of microphones. Here is the x-coordinate (in meters) of the nth microphone relative to the reference point, f is the signal frequency (Hz), and c is the speed of sound (taken as 343m / s). All adjustments to the gain and filtering parameters are linked to the sound source reliability assessment module through a feedback loop: if the latter determines the sound source to be of low reliability three times in a row, the gain coefficient of the next cycle is automatically reduced by 1dB until the minimum limit of 8dB is reached to prevent system overload caused by continuous false triggering.

[0031] In some implementations, the sound source credibility assessment module receives the multi-channel audio stream processed by the adaptive signal conditioning module and the corresponding context feature vector, and performs sound source authenticity determination. This module first performs cross-correlation analysis on the signals of each channel and calculates the time delay corresponding to the maximum cross-correlation peak. (m≠n), combined with the known geometric layout of the microphone array (coordinate matrix) The spatial azimuth angle of the sound source is calculated using the least squares method. The objective function is optimized as follows:

[0032] in Assuming the location of the sound source (r is taken as 5m). Let m be the coordinates of the m-th microphone; Then, Location information provided by infrared human presence detector (Consistency verification is performed by mapping the installation location to the detection area); If the angle deviation If so, it can be preliminarily determined that the sound source is not from a human body.

[0033] In some implementations, the sound source orientation calculation incorporates device attitude angle compensation provided by the inertial measurement unit (IMU), assuming the pitch angle measured by the IMU is... Horizontal roll angle is Then, the original microphone coordinates are rotated and corrected.

[0034] in , These are the rotation matrices about the x-axis and y-axis, respectively. This compensation can eliminate geometric model inaccuracies caused by uneven wall surfaces or tilted mounting brackets. The measured positioning error was reduced from an average of 18.7° without compensation to 5.3°.

[0035] Furthermore, this module extracts three acoustic features of the audio signal: fundamental frequency stability. Harmonic structure integrity and instantaneous energy mutation rate Fundamental frequency stability is measured by the peak consistency of the autocorrelation function within the fundamental frequency candidate interval; harmonic structure integrity is calculated using the harmonic-to-noise ratio (HNR), defined as the ratio of harmonic component energy to residual noise energy; instantaneous energy mutation rate is defined as the mean absolute value of the energy derivative within a 20ms window. These three features are matched with a pre-stored library of typical abnormal speech templates (containing 12 categories such as screams, cries, calls for help, and fighting sounds) using dynamic time warping (DTW) to obtain a matching score. ; like If the distance to the sound source is normalized and the location of the sound source matches the location of the human body, then a high-confidence sound source identifier (confidence label) is generated. Otherwise, mark it as low credibility. ( ), prohibiting it from entering the subsequent classification process.

[0036] In some implementations, the context-enhanced anomalous speech classification module is deployed in the ARMTrustZone Trusted Execution Environment (TEE) of the edge computing node, ensuring that the classification process is isolated from the operating system. Its input is high-confidence audio frames filtered by the sound source confidence assessment module. ) and its associated context feature vector This module adopts a three-tiered decision-making architecture: The first layer is a coarse classifier for acoustic events. It uses a 1D convolutional neural network (CNN) to perform binary discrimination on the Mel-spectrogram of the input audio. The network structure contains four convolutional layers (with 32, 64, 128, and 256 filters respectively, a kernel size of 5, and a stride of 2), followed by global average pooling and fully connected layers. It outputs the probabilities of non-verbal acoustic events (such as screams and broken glass) and verbal speech. , The training dataset contains 15,000 labeled audio clips collected from campus scenes, with an accuracy of 96.2%.

[0037] Following on from the above, the second layer is a semantic-sentiment joint analyzer, which only analyzes... The sample activation and keyword matching adopt an encrypted dictionary tree structure. The sensitive word library (such as 87 words such as "help", "robbery", "fire" etc.) is stored in the TEE in a homomorphic encrypted form. During a query, the phoneme sequence output by the speech recognition engine (using an end-to-end CTC model) is converted into a ciphertext hash path, enabling existence verification without decrypting the original speech. Emotional state recognition, through an attention mechanism, fuses prosodic features (fundamental frequency profile, intensity envelope, and speech rate) with contextual information such as time (whether it is nighttime, 22:00-6:00), location (whether it is a high-risk area), and population density (estimated from historical infrared trigger frequencies), outputting the probability of abnormal emotion. ; The encrypted dictionary tree structure uses a combination of Bloom filters and Paillier homomorphic encryption to verify the existence of keywords. The Bloom filter bit array is 10,000 bits long, the number of hash functions is 5, and the false positive rate is controlled below 0.1%. The Paillier public key is uniformly distributed by the central management platform, and the private key exists only within the TEE to ensure that the dictionary content is irreversible. Next, the third layer is the context fusion decision unit, which integrates the outputs of the first two layers and the current time period flag. (1 indicates nighttime), area type code (0 = ordinary corridor, 1 = stairwell, 2 = laboratory) and historical alarm frequencies (Number of alarms per square meter in the past 7 days) is used to generate the final abnormal event determination signal through weighted logic gate circuits. ; The decision logic is as follows:

[0038] Among them, dynamic emotion threshold Determined by the following formula:

[0039] The design ensures that an alert is generated only when acoustic features, semantic content, emotional state, and scene context all point to an anomaly.

[0040] In some implementations, the privacy-preserving data processing module is responsible for end-to-end security control of the data stream throughout the system's lifecycle. The raw audio data is immediately destroyed after the adaptive signal conditioning module completes feature extraction (such as Mel spectrum and fundamental frequency trajectory), retaining only irreversible acoustic feature vectors (dimension ≤ 128) and intermediate classification results (such as...). , ); As mentioned earlier, all data transmitted to the central management platform is homomorphically encrypted using Paillier with a public key length of 2048 bits. The encrypted data cannot be used to recover the audio content. This module has a built-in RBAC (Role-Based Access Control) policy engine. Ordinary security personnel can only receive abnormal event alarm summaries (including time, location, and event type) and cannot obtain any original or intermediate audio data. School administrators can apply for temporary decryption permissions, but this requires approval from two people and recording of the operation log.

[0041] In some implementations, the privacy-preserving data processing module maintains a rolling feature cache queue locally at the edge node with a capacity of 10 minutes (600 frames at a 1Hz update rate). Only when the abnormal event determination signal When activated, the feature vectors within 30 seconds before and after the corresponding event are packaged and uploaded; data from other periods are automatically cleared before being overwritten in the cache, eliminating unnecessary data retention from the source. Actual tests show that this strategy reduces the average daily data upload volume of edge nodes from 1.2GB to 85MB.

[0042] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A system for sensing and classifying abnormal voice signals in concealed areas of a campus, characterized in that, include: The system includes a distributed acoustic sensor array module, an environmental context awareness module, an adaptive signal conditioning module, a sound source credibility assessment module, a context-enhanced abnormal speech classification module, and a privacy-preserving data processing module. The modules are interconnected through a unified time synchronization bus and a secure communication link, forming an integrated processing pipeline with status feedback and parameter coordination capabilities. The distributed acoustic sensor array module is deployed in the target concealed area and consists of multiple microphone units. Each microphone unit integrates a preamplifier circuit and an analog-to-digital converter, and adopts an asymmetric spatial arrangement strategy so that the main axis of adjacent microphone units is offset by a preset angle, so as to form a spatial suppression capability against sound sources in non-target directions at the physical level. The microphone unit synchronously samples the audio signal and uploads the raw digital audio frames with timestamp tags to the adaptive signal conditioning module; The environmental context awareness module is configured within the local edge computing node and includes an inertial measurement unit, a temperature and humidity sensor, an infrared human presence detector, and a background noise spectrum analysis unit. It is used to periodically collect environmental physical state parameters and generate multi-dimensional context feature vectors. The adaptive signal conditioning module receives the original audio frame and the multi-dimensional context feature vector, dynamically switches the working mode according to the infrared human presence state and background noise energy level, and links the beamforming algorithm to focus on the direction of the most recent human infrared signal. The sound source confidence assessment module calculates the spatial azimuth angle of the sound source based on the multi-channel audio stream and performs consistency verification with the positioning information of the infrared human presence detector. At the same time, it combines three acoustic characteristics, namely fundamental frequency stability, harmonic structure integrity and instantaneous energy mutation rate, to determine whether the sound source is a high-confidence human sound source. Only when the sound source is marked as high-confidence is it allowed to enter the subsequent classification process. The context-enhanced abnormal speech classification module is deployed in a trusted execution environment and adopts a hierarchical decision architecture; The privacy-preserving data processing module destroys the original audio data immediately after feature extraction, retaining only the irreversible acoustic feature vectors, and performs homomorphic encryption on the data uploaded to the central management platform. At the same time, it dynamically assigns data viewing permissions based on user roles.

2. The abnormal voice sensing and classification system for concealed areas on campus according to claim 1, characterized in that, Each microphone unit in the distributed acoustic sensor array module interacts with the local edge computing node through a shared buffer in dual-port RAM, and an interrupt is triggered by the DMA controller to achieve event-driven low-latency data upload.

3. The abnormal voice sensing and classification system for concealed areas on campus according to claim 1, characterized in that, The background noise spectrum analysis unit in the environmental context awareness module uses a sliding window mechanism to perform short-time Fourier transform on continuous audio segments and divides frequency bands by frequency range. When a steady-state interference source exists in the frequency band, the adaptive signal conditioning module is notified to perform notch filtering on that frequency band.

4. The abnormal voice sensing and classification system for concealed areas on campus according to claim 1, characterized in that, The adaptive signal conditioning module has a built-in finite state machine controller, which is used to determine whether to enter the low-power listening mode. When the presence of a human body is detected or a sudden increase in noise energy is detected, the system switches to a high-sensitivity capture mode and initiates a delay-summation beamforming algorithm to focus on the spatial orientation of the most recently recorded infrared signal.

5. The abnormal voice sensing and classification system for concealed areas on campus according to claim 1, characterized in that, When calculating the spatial azimuth angle of the sound source, the sound source reliability assessment module introduces the device pitch angle and roll angle provided by the inertial measurement unit to perform rotation correction on the geometric coordinates of the microphone array, thereby eliminating spatial positioning errors caused by the tilt of the mounting surface.

6. The abnormal voice sensing and classification system for concealed areas on campus according to claim 1, characterized in that, The hierarchical decision-making architecture includes: The first layer performs binary discrimination between non-verbal acoustic events and verbal speech; the second layer performs keyword existence verification and calculates the probability of abnormal emotions for verbal speech simultaneously; the third layer integrates the outputs of the first two layers, the current time period, the region type, and the historical alarm frequency, and generates the final abnormal event judgment signal through weighted logic gate circuits. The first layer of the context-enhanced abnormal speech classification module uses a one-dimensional convolutional neural network to process the Mel spectrogram of the input audio and output the probabilities of non-linguistic acoustic events and linguistic speech. In the second layer, keyword existence verification is achieved through an encrypted trie structure, the sensitive word library is stored in a trusted execution environment in a homomorphic encryption form, and the phoneme sequence output by speech recognition is converted into a ciphertext hash path for matching and querying.

7. The abnormal voice sensing and classification system for concealed areas on campus according to claim 6, characterized in that, The encrypted dictionary tree structure uses a combination of Bloom filters and Paillier homomorphic encryption to verify the existence of keywords.

8. The abnormal voice sensing and classification system for concealed areas on campus according to claim 1, characterized in that, In the third layer of the context-enhanced abnormal speech classification module, the generation logic for the abnormal event determination signal is activated if any of the following conditions are met: The probability of a non-verbal acoustic event is greater than 0.85 and the current area type is a stairwell or laboratory corridor; The probability of language-related speech is greater than 0.7, the probability of abnormal emotions is higher than the dynamic threshold, and the alarm frequency per square meter in this area is less than 0.05 times in the past 7 days; The dynamic emotion threshold mentioned above is dynamically adjusted based on whether it is nighttime and the regional risk level, and the calculation formula is as follows: in This is a nighttime marker. Encodes the region type.

9. The abnormal voice sensing and classification system for concealed areas on campus according to claim 1, characterized in that, The privacy-preserving data processing module maintains a rolling feature cache queue locally at the edge node; The irreversible acoustic feature vector of an event is packaged and uploaded only when the abnormal event determination signal is activated; data from other time periods is automatically cleared before being overwritten by the cache.

10. A method for sensing and classifying abnormal voice in concealed areas of a campus, applicable to the abnormal voice sensing and classification system for concealed areas of a campus as described in any one of claims 1-9, characterized in that, Includes the following steps: Acquire multi-channel audio frames synchronously acquired by a distributed acoustic sensor array and multi-dimensional context feature vectors generated by an environmental context awareness module; The pickup gain and filtering strategy are dynamically adjusted based on the human body's presence and background noise energy, and beamforming focusing is performed. The azimuth angle of the sound source is calculated based on multi-channel cross-correlation analysis, and consistency verification is performed by combining infrared positioning information. Three features are extracted: fundamental frequency stability, harmonic structure integrity, and energy mutation rate, to screen high-confidence sound sources. A three-level classification is performed on high-confidence audio frames in a trusted execution environment: first, it is determined whether it is a non-linguistic abnormal acoustic event; if it is linguistic speech, the existence of sensitive keywords is verified and the probability of emotional abnormality is calculated simultaneously; finally, the acoustic category, semantic results, emotional probability, time period, region type and historical false alarm frequency are fused to generate the final abnormal event judgment signal. The original audio is destroyed immediately after feature extraction is completed. Only the irreversible feature vector is uploaded in encrypted form, and data access permissions are controlled according to user roles.