Microphone array-based campus acoustic monitoring fusion privacy protection bullying early warning system and method

By using microphone arrays and multimodal fusion decision technology, the privacy protection and adaptability issues of campus indoor security monitoring have been solved, enabling accurate early warning of bullying risks, reducing false alarm rates, and making it suitable for security monitoring in complex acoustic environments on campus.

CN121963383APending Publication Date: 2026-05-01CHONGQING ZHONGKAI TESTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING ZHONGKAI TESTING CO LTD
Filing Date
2026-01-28
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing campus indoor security monitoring technologies have shortcomings in terms of privacy protection, adaptability to complex acoustic environments, multi-dimensional information fusion, and early warning accuracy. In particular, they suffer from severe sound field reverberation, low far-field recognition accuracy, and high false alarm rate in narrow rectangular dormitory spaces, and cannot effectively distinguish between daily noise and high-risk sounds.

Method used

A microphone array is used to acquire multi-channel audio signals. Combined with acoustic front-end processing, voiceprint recognition, emotion recognition and abnormal sound detection, the warning level is output through a multimodal fusion decision module, including DC removal, automatic gain control, beamforming, deep VAD, x-vector network, CRNN network and multi-label sound event detection, to achieve signal-to-noise ratio improvement and multi-evidence triggered logical decision-making.

Benefits of technology

While ensuring privacy, it achieves accurate and real-time early warning of indoor bullying risks on campus, reduces false alarm rate, improves voice signal quality and recognition accuracy, adapts to complex acoustic environments, and meets compliance requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963383A_ABST
    Figure CN121963383A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of audio signal processing and artificial intelligence, in particular to a microphone-array-based campus acoustic monitoring fusion privacy protection bullying early warning system and method.Audio signals are collected through a microphone array deployed in a campus indoor space, after the signal-to-noise ratio is increased through acoustic front-end processing, acoustic features are extracted, and the acoustic front-end processing is completed; and finally, carrying out fusion decision making according to a multi-evidence triggering logic through a multi-modal fusion decision making module, outputting a risk level and triggering early warning. The voice signal quality and the recognition accuracy are effectively improved, the false alarm rate is remarkably reduced through a three-in-one fusion decision-making mechanism of voiceprint, emotion and abnormal sound, meanwhile, the privacy of students is strictly guaranteed through the design of end-side processing, data minimization and the like, and the accurate, real-time and compliant bully risk early warning effect is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of audio signal processing and artificial intelligence technology, and in particular to a campus acoustic monitoring and privacy protection bullying early warning system and method based on microphone array. Background Technology

[0002] Currently, in densely populated and somewhat private indoor environments such as school campuses, safety risk monitoring mainly employs the following technical means, but these methods are clearly insufficient in terms of effectiveness, accuracy, and privacy protection.

[0003] Manual patrol and management systems are common, but they rely entirely on human experience and judgment, resulting in discontinuous monitoring, insufficient coverage at night or during off-peak hours, and difficulty in timely detection and intervention during emergencies. Some video surveillance systems detect abnormal behavior through image recognition or manual inspection. However, the application of video surveillance in indoor spaces such as dormitories is significantly limited: these are high-privacy areas, and video collection can easily lead to privacy and compliance disputes; furthermore, video systems are highly dependent on lighting conditions and obstruction, with recognition effectiveness decreasing at night or when obstructed; and their ability to identify risk events primarily expressed through sound, such as crying, arguing, or emotional outbursts, is limited.

[0004] Existing technologies also include traditional sound monitoring and alarm devices, which typically consist of a single or small number of microphones, a sound acquisition module, and a threshold judgment unit. An alarm is triggered when the detected sound decibel level exceeds a preset threshold. These devices are simple in structure and low in cost, but they have significant shortcomings: they usually judge based solely on sound energy, failing to distinguish between everyday noise and high-risk abnormal sounds; they cannot identify the individual source of the sound and their emotional state; and they are prone to false alarms from normal activity sounds in indoor environments, limiting their practicality. Furthermore, while there are general emotion recognition models based on deep learning in existing technologies, most are trained on recording studio or near-field interference-free data. However, campus indoor environments (such as dormitories) are typically long and narrow rectangular spaces with hard materials for floors and walls, resulting in a sound field characterized by strong reverberation (long RT60 reverberation time) and multipath reflections. In such scenarios, existing general models have serious flaws: strong reverberation blurs the speech spectrum, causing the loss of fine frequency domain features used to distinguish anger or fear, leading to model misjudgment; simultaneously, in far-field, low signal-to-noise ratio conditions, directly applying general models yields extremely low accuracy.

[0005] In summary, existing campus indoor security monitoring technologies generally suffer from the following shortcomings: a lack of effective preprocessing methods for strong reverberation and far-field acoustic environments, resulting in poor front-end voice quality; limited monitoring dimensions, relying solely on decibel thresholds or single emotion recognition, failing to handle the acoustic confusion between "normal activity sounds" and "real risk events," leading to a high false alarm rate; and a failure to establish an effective joint judgment and multimodal fusion mechanism between voiceprint identity, emotional state, and abnormal physical sounds. Therefore, there is an urgent need for a technical solution that, while ensuring privacy, is suitable for complex acoustic environments within campuses and can achieve accurate and reliable abnormal behavior identification and early warning. Summary of the Invention

[0006] The purpose of this invention is to provide a campus acoustic monitoring and privacy protection bullying early warning system and method based on microphone array, which solves the shortcomings of existing campus indoor security monitoring technologies in terms of privacy protection, adaptability to complex acoustic environments, multi-dimensional information fusion, and early warning accuracy.

[0007] To achieve the above objectives, this invention provides a method for campus acoustic monitoring and privacy protection-based bullying early warning based on microphone arrays, comprising the following steps: Multi-channel audio signals are collected using microphone arrays deployed in indoor spaces on campus. The multi-channel audio signal is subjected to acoustic front-end processing to output a high signal-to-noise ratio audio stream; Extract acoustic features from the high signal-to-noise ratio audio stream; The acoustic features are input in parallel to the voiceprint recognition module, the emotion recognition module, and the abnormal sound detection module to obtain voiceprint recognition results containing speaker identity information, emotion recognition results containing emotion category and confidence level, and abnormal sound detection results containing abnormal sound event category and confidence level, respectively. The voiceprint recognition result, the emotion recognition result, and the abnormal sound detection result are input into the multimodal fusion decision module, which makes a decision based on the preset fusion rules and outputs a warning level associated with the bullying risk. Warning information is generated and pushed out based on the warning level.

[0008] Specifically, the multi-channel audio signal undergoes acoustic front-end processing to output a high signal-to-noise ratio audio stream, including... It performs DC rejection, automatic gain control, wind noise and power supply noise suppression, and echo cancellation; and estimates the direction of arrival of the sound source based on the GCC-PHAT and SRP-PHAT algorithms, and guides beamforming based on the estimation results.

[0009] Specifically, extracting acoustic features from the high signal-to-noise ratio audio stream includes: Deep VAD is used for speech activity detection to determine the speech segment to be processed; the speech segment is framed and windowed; 80-dimensional log-Mel filter bank features are calculated based on the windowed speech frames, and the cepstral mean-variance is normalized on the features.

[0010] The voiceprint recognition module uses an x-vector network architecture to extract the speaker embedding vector and combines the sound source arrival direction angle with the position matching relationship of the corresponding voiceprint area to improve the recognition confidence.

[0011] The emotion recognition module uses a CRNN network for classification and outputs a stable emotion state based on a sliding window hysteresis control strategy. The strategy includes: when the proportion of frames belonging to a preset high-risk emotion in a decision window of length N frames exceeds a first threshold, the system state flips from normal to abnormal and locks the abnormal state for at least T seconds.

[0012] The abnormal sound detection module adopts a multi-label sound event detection framework, uses a lightweight CRNN network for streaming inference, identifies event types including crying, screaming, falling and impact, and integrates an open set anomaly detection mechanism.

[0013] The multimodal fusion decision module has a preset fusion rule that is a multi-evidence triggering logic. The logic includes: when a high-risk emotion with a confidence level higher than the first threshold is continuously detected, a physical abnormal sound event is detected at the same time, and the voiceprint identity is confirmed, it is determined to be a high-risk warning level.

[0014] A campus acoustic monitoring and privacy protection bullying early warning system based on a microphone array includes a microphone array, an acoustic front-end processing module, an acoustic feature extraction module, a voiceprint recognition module, an emotion recognition module, an abnormal sound detection module, and a multimodal fusion decision module. The acoustic front-end processing module is connected to the microphone array, the acoustic feature extraction module is connected to the acoustic front-end processing module, the voiceprint recognition module, the emotion recognition module, and the abnormal sound detection module are respectively connected to the acoustic feature extraction module, and the multimodal fusion decision module is respectively connected to the voiceprint recognition module, the emotion recognition module, and the abnormal sound detection module. The microphone array is used to collect multi-channel audio signals from indoor spaces on campus. The acoustic front-end processing module is used to perform de-reverberation and beamforming processing on the audio signal; The acoustic feature extraction module is used to extract features from the processed audio stream and distribute them. The voiceprint recognition module, the emotion recognition module, and the abnormal sound detection module are used to generate recognition results for voiceprint, emotion, and abnormal sound, respectively. The multimodal fusion decision module is used to fuse multimodal recognition results according to preset rules, output a warning level decision, and generate and send warning information based on the decision.

[0015] This invention discloses a campus acoustic monitoring and privacy-protecting bullying early warning system and method based on a microphone array. Audio signals are collected by a microphone array deployed in an indoor campus space. After acoustic front-end processing (including dreverberation and beamforming based on the MVDR algorithm) to improve the signal-to-noise ratio, acoustic features (such as 80-dimensional log-Mel filter bank features) are extracted and fed in parallel into a voiceprint recognition module (using an x-vector network architecture), an emotion recognition module (using a CRNN network combined with a sliding window hysteresis control strategy), and an abnormal sound detection module (using a multi-label sound event detection framework). Finally, a multimodal fusion decision module performs a fusion decision based on multi-evidence triggering logic, outputs a risk level, and triggers an early warning. This system effectively improves the quality and accuracy of speech signals in the complex acoustic environment of a campus with strong reverberation, far-field conditions, and multiple participants, without requiring speech transcription or semantic understanding. The three-in-one fusion decision mechanism of voiceprint, emotion, and abnormal sound significantly reduces the false alarm rate. Simultaneously, edge processing and data minimization designs strictly protect student privacy, achieving accurate, real-time, and compliant bullying risk early warning. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0017] Figure 1 This is an overall structural diagram of the campus acoustic monitoring and privacy protection bullying early warning system based on a microphone array according to the first embodiment of the present invention.

[0018] Figure 2 This is a flowchart of voiceprint recognition according to the second embodiment of the present invention.

[0019] Figure 3 This is a flowchart of emotion recognition according to the second embodiment of the present invention.

[0020] Figure 4 This is a flowchart of the second embodiment of the campus acoustic monitoring and privacy protection bullying early warning method based on microphone array in this invention.

[0021] Figure 5 This is a flowchart of the steps of the campus acoustic monitoring and privacy protection bullying early warning method based on microphone array according to the second embodiment of the present invention.

[0022] In the diagram: 101-Microphone array, 102-Acoustic front-end processing module, 103-Acoustic feature extraction module, 104-Voiceprint recognition module, 105-Emotion recognition module, 106-Abnormal sound detection module, 107-Multimodal fusion decision module. Detailed Implementation

[0023] The embodiments of the present invention are described in detail below. Examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, but should not be construed as limiting the present invention.

[0024] The first embodiment of this application is as follows: Please see Figure 1 The present invention provides a campus acoustic monitoring and privacy protection bullying early warning system based on a microphone array, including a microphone array 101, an acoustic front-end processing module 102, an acoustic feature extraction module 103, a voiceprint recognition module 104, an emotion recognition module 105, an abnormal sound detection module 106, and a multimodal fusion decision module 107.

[0025] The acoustic front-end processing module 102 is connected to the microphone array 101, the acoustic feature extraction module 103 is connected to the acoustic front-end processing module 102, the voiceprint recognition module 104, the emotion recognition module 105 and the abnormal sound detection module 106 are respectively connected to the acoustic feature extraction module 103, and the multimodal fusion decision module 107 is respectively connected to the voiceprint recognition module 104, the emotion recognition module 105 and the abnormal sound detection module 106; The microphone array 101 is used to collect multi-channel audio signals from indoor spaces on campus. The acoustic front-end processing module 102 is used to perform de-reverberation and beamforming processing on the audio signal; The acoustic feature extraction module 103 is used to extract features from the processed audio stream and distribute them. The voiceprint recognition module 104, the emotion recognition module 105, and the abnormal sound detection module 106 are used to generate recognition results for voiceprint, emotion, and abnormal sound, respectively. The multimodal fusion decision module 107 is used to fuse multimodal recognition results according to preset rules, output a warning level decision, and generate and send warning information based on the decision.

[0026] Specifically, each room, such as a dormitory, deploys a 4-8 channel digital microphone array 101, supporting a sampling rate of 16kHz or 48kHz, with clock synchronization and gain consistency calibration between channels.

[0027] The acoustic front-end processing module 102 further includes: Preprocessing unit: used to perform DC removal, automatic gain control (AGC), wind noise / power supply noise suppression, and echo cancellation (eliminating interference from corridor broadcasts / rings, etc.). DOA estimation unit: Based on GCC-PHAT (Generalized Cross-Correlation-Phase Transform) and SRP-PHAT (Controllable Power Response-Phase Transform) algorithms, the direction of arrival of the sound source is estimated; De-reverberation and beamforming unit: For the typical "shoe-box" spatial model of dormitories, the MVDR (Minimum Variance Distortionless Response) adaptive beamforming algorithm is adopted instead of simple omnidirectional reception.

[0028] In addition, the system also adjusts the scene-specific parameters: the main sound pickup beam is preset to point to the bed area on both sides of the dormitory (ROI area), and "Null-steering" is set in the beamforming algorithm to actively suppress external broadcast or bell noise interference from the door and corridor, significantly improving the signal-to-noise ratio (SNR) of the target area.

[0029] The acoustic feature extraction module 103 performs the following processing flow: Speech segmentation: Deep VAD (Voice Activity Detection) is used for speech activity detection. Overlapping speech is marked with overlapping speaker detection tags, and speaker separation is performed when necessary. Framing and windowing: The speech signal is divided into short time frames (20-40ms), and Hamming windows are used to reduce spectral leakage; log-Mel features: Calculate the 80-dimensional log-Mel filter bank features and perform cepstral mean-variance normalization (CMVN) to eliminate the influence of channel and environmental changes.

[0030] The voiceprint recognition module 104 uses an x-vector network architecture to extract speaker embedding vectors. The improvement of this invention lies in the registration strategy: in the initial stage of system deployment, a small amount of voice recordings from dormitory members in both "quiet" and "generally noisy" environments are collected to construct an adaptive base library, and an array spatial consistency check is introduced—that is, confidence is increased only when the source localization angle (DOA) roughly matches the bed area corresponding to the voiceprint, thus preventing recording playback attacks.

[0031] The emotion recognition module 105 (with added hysteresis control logic) uses a CRNN (Convolutional Recurrent Neural Network) for classification. To address false alarms caused by instantaneous noise, this invention designs a sliding window-based hysteresis control strategy: Sliding window mechanism: Set a decision window with a length of N frames (e.g., N=50, approximately 1.5 seconds). State flip logic: The system state flips from "normal" to "abnormal" only when the proportion of frames in the window that belong to the high-risk emotion of "anger / fear" exceeds a threshold (e.g., 80%). State Locking: Once an abnormal state is triggered, the system will lock that state for at least T seconds (e.g., 3 seconds) and will not easily revert to that state unless a clear "calm" signal or silence is detected. This effectively avoids frequent jumps in emotional state caused by pauses in speech or breathing sounds.

[0032] The abnormal sound detection module 106 is used to identify specific risky sound events, and its design is as follows: Event collection: Supports detection of typical campus risk sounds such as crying, screaming, throwing, impact, continuous arguing, and suspected fighting; supports expansion to high-risk events such as broken glass; Detection framework: The multi-label sound event detection (SED) framework is adopted to output the probability of various events over time; the network adopts a lightweight CRNN and supports streaming inference; Low false alarm strategy: Design a two-level cascade (rapid screening + verification) and "background noise suppression" joint modeling to effectively distinguish daily dormitory noises such as dragging chairs, closing doors, and washing up, and significantly reduce false alarms; Open set anomalies: By combining contrastive learning methods to calculate the Out-of-Distribution (OOD) score or reconstruction error, we provide alerts for unseen anomalies.

[0033] The multimodal fusion decision module 107 is responsible for integrating information from three dimensions: voiceprint, emotion, and abnormal sound. Its workflow includes aligning the voiceprint ID, emotion probability, and abnormal sound probability on the time axis, and outputting the risk level using a rule engine and a lightweight model. Specifically, it includes: Fusion Strategy: The rule engine defines multi-evidence triggering logic (A (emotion) > threshold + B (voiceprint) confirmation + C (abnormal sound) = D (risk level)), and the lightweight model learns the weighted fusion of multimodal features; Triggering mechanism: Set hysteresis threshold, cooldown time and multi-evidence trigger to avoid frequent alarms; support backtracking the most recent N seconds of cache according to the authorized process for review; Risk Level: Output three risk levels: alert, general, and emergency, and push the dormitory number, time, event type, risk level, and speaker ID.

[0034] The core decision logic of its rule engine is illustrated below: Scene Emotion recognition results (lasting >1.5 seconds) Abnormal sound detection results Voiceprint identity features Risk level determination Response Action Scenario A: Physical Conflict Fear / Crying (Confidence > 0.8) Impact / slamming sound detected Identified as a dormitory member Level 1 (Emergency) Immediate push notifications + dormitory management pop-up messages Scene B: Heated Argument Anger (confidence level > 0.7) Detected sustained high volume Recognize that different members can speak alternately. Level 2 (Warning) Mark events to notify patrols Scene C: Routine playful fighting Neutral / Excited (with laughter) A short impact sound was detected. - Level 3 (Ignore / Hint) Log only, no alarms. Scene D: Solo venting Anxiety / Sobbing No other abnormal sounds Single member continued Level 2 (Attention) Send psychological counseling suggestions This invention incorporates privacy protection principles from the outset, and specific measures include: Without ASR and semantic understanding: The system does not output specific speech content, but only outputs emotion, event category and confidence level based on acoustic features; Data minimization: By default, the original audio is not saved; only the necessary event metadata and the encrypted voiceprint vector are saved. A minimum retention period can be set. Device-side priority: Feature extraction and inference should be completed on the device side as much as possible, and the transmission should use an encrypted channel. Platform-side access permissions should be hierarchical and operation logs should be recorded.

[0035] Through the collaborative work of the above modules, high-precision, low-false-alarm monitoring of bullying risks is achieved in complex indoor acoustic environments, while strictly protecting users' personal privacy at the technical level.

[0036] The second embodiment of this application is as follows: Based on the first embodiment, please refer to Figures 2 to 5 The method for campus acoustic monitoring and privacy protection-based bullying early warning based on microphone array in this embodiment includes the following steps: S201: Acquire multi-channel audio signals through a microphone array 101 deployed in the indoor space of the campus; S202: Perform acoustic front-end processing on the multi-channel audio signal to output a high signal-to-noise ratio audio stream; S203: Extract acoustic features from the high signal-to-noise ratio audio stream; S204: The acoustic features are input in parallel to the voiceprint recognition module 104, the emotion recognition module 105 and the abnormal sound detection module 106 to obtain voiceprint recognition results containing speaker identity information, emotion recognition results containing emotion category and confidence level, and abnormal sound detection results containing abnormal sound event category and confidence level, respectively. S205: Input the voiceprint recognition result, the emotion recognition result, and the abnormal sound detection result into the multimodal fusion decision module 107, make a decision according to the preset fusion rules, and output the warning level associated with the bullying risk; S206: Generate and push warning information according to the warning level.

[0037] Specifically, firstly, multi-channel audio signals are acquired using microphone arrays 101 deployed in indoor spaces on campus (such as dormitories and activity rooms). Specifically, 4-8 channel digital microphone arrays 101 can be used, supporting sampling rates of 16kHz or 48kHz, and ensuring inter-channel clock synchronization and gain consistency.

[0038] Next, the multi-channel audio signal undergoes acoustic front-end processing to output a high signal-to-noise ratio audio stream. Specifically, the acoustic front-end processing includes: performing DC removal, automatic gain control (AGC), wind noise and power supply noise suppression, and echo cancellation (e.g., eliminating interference from corridor broadcasts or ringtones) through a preprocessing unit; estimating the direction of arrival of the sound source based on GCC-PHAT (Generalized Cross-Correlation-Phase Transform) and SRP-PHAT (Controllable Power Response-Phase Transform) algorithms; and then employing the MVDR (Minimum Variance Distortionless Response) adaptive beamforming algorithm for déreverberation and beamforming, pre-setting the main pickup beam towards the target monitoring area (e.g., the bed area) for typical indoor layouts, while simultaneously setting nulls in the interference direction (e.g., the corridor direction) in the algorithm to actively suppress external noise and significantly improve the signal-to-noise ratio of the target speech.

[0039] Then, acoustic features are extracted from the high signal-to-noise ratio audio stream. This step further includes: using deep VAD (voice activity detection) to determine valid speech segments; performing frame segmentation (e.g., 20-40ms) and windowing (e.g., Hamming window) on the speech segments; calculating 80-dimensional log-Mel filter bank features based on the windowed speech frames, and performing cepstral mean-variance normalization on them to eliminate the influence of channel and environmental differences.

[0040] Subsequently, the normalized acoustic features are input in parallel to the three recognition modules. In the voiceprint recognition module 104, an x-vector network architecture is used to extract the speaker embedding vector, and spatial consistency is checked by combining the sound source arrival direction angle estimated in step S202 with the preset voiceprint region (such as bed position), thereby improving the recognition confidence and preventing recording attacks. In the emotion recognition module 105, a CRNN network is used to classify speech, and a sliding window-based hysteresis control strategy is introduced to stabilize the output: that is, only when the proportion of frames judged as high-risk emotions such as "anger" or "fear" exceeds a preset threshold (e.g., 80%) within a decision window of length N frames (e.g., 50 frames, about 1.5 seconds), the system emotion state is flipped to "abnormal", and once an abnormality is triggered, the state will be locked for at least T seconds (e.g., 3 seconds) to avoid misjudgment and state jump caused by short pauses or noise. In the abnormal sound detection module 106, a multi-label sound event detection framework is adopted, and a lightweight CRNN network is used for streaming inference to identify preset risk events such as crying, screaming, falling, and impact. An open set anomaly detection mechanism (such as calculating OOD score) is also integrated to provide prompts for unseen abnormal sounds.

[0041] Subsequently, the voiceprint recognition results (including speaker ID and confidence level), emotion recognition results (including emotion category and confidence level), and abnormal sound detection results (including event category and confidence level) are input into the multimodal fusion decision module 107. Based on preset fusion rules, a decision is made, and a warning level associated with bullying risk is output. The fusion rules employ multi-evidence triggering logic, implemented collaboratively by a rule engine and a lightweight model. For example, the rule engine defines: when a high-risk emotion with a confidence level exceeding a first threshold is continuously detected, a physical abnormal sound event (such as a collision sound) is detected simultaneously, and the voiceprint identity is confirmed as that of an indoor resident, a high-risk warning level is determined. This module also effectively avoids frequent false alarms by setting hysteresis thresholds, cooling-off times, and requiring multi-evidence collaborative triggering.

[0042] Finally, based on the generated warning level, a warning message is generated and pushed out. The warning message includes at least the room number where the event occurred, a timestamp, event type, risk level, and the identified speaker ID (if applicable), and is sent to the management platform according to a preset response strategy (such as pop-up alerts, notifications for patrols, and log recording). Throughout the process, the system adheres to privacy protection design: it does not perform speech-to-text transcription or semantic understanding, does not save the original audio on the server by default, and feature extraction and inference are primarily completed on the edge device, uploading only the necessary encrypted event metadata.

[0043] Through the above steps, the method provided in this embodiment achieves real-time, accurate, and low-false-alarm monitoring of bullying risk events in complex indoor acoustic environments on campuses, while strictly protecting personal privacy at the technical level and meeting compliance requirements.

[0044] This invention addresses the "strong reverberation + far field + multiple people" scenario in school dormitories by constructing an integrated link of "array enhancement + voiceprint recognition + emotion recognition + abnormal sound detection" to achieve security warnings without collecting semantic data, effectively protecting student privacy. This invention introduces the MVDR beamforming algorithm, which effectively improves the signal-to-noise ratio of far-field speech. Combined with the x-vector voiceprint recognition network, it enables accurate speaker identification in complex dormitory environments. This invention uses a CRNN network for emotion recognition, combined with time smoothing and hysteresis control strategies, to effectively distinguish between everyday playful fighting and bullying scenarios, thereby reducing the false alarm rate of emotion recognition. This invention introduces a cascaded verification and open set anomaly detection mechanism, combined with background noise suppression class joint modeling, which effectively reduces false alarms caused by the complex dormitory environment and can also provide effective prompts for unseen abnormal sound events; This invention enables lightweight deployment on the client side, supports real-time processing and low-bandwidth upload, while taking into account privacy protection and compliance requirements, and has good engineering feasibility. This invention significantly improves the stability and accuracy of far-field emotion recognition and abnormal event determination by using front-end acoustic processing designed for strong reverberation environments in dormitories, multimodal evidence fusion strategies, and time hysteresis control mechanisms. It avoids the problems of high false alarm rates and poor robustness of existing technologies in actual dormitory deployments.

[0045] The above-disclosed embodiments are merely one or more preferred embodiments of this application and should not be construed as limiting the scope of this application. Those skilled in the art can understand that all or part of the processes for implementing the above embodiments and equivalent changes made in accordance with the claims of this application still fall within the scope of this application.

Claims

1. A method for campus acoustic monitoring and privacy-protecting bullying early warning based on microphone array, characterized in that, Includes the following steps: Multi-channel audio signals are collected using microphone arrays deployed in indoor spaces on campus. The multi-channel audio signal is subjected to acoustic front-end processing to output a high signal-to-noise ratio audio stream; Extract acoustic features from the high signal-to-noise ratio audio stream; The acoustic features are input in parallel to the voiceprint recognition module, the emotion recognition module, and the abnormal sound detection module to obtain voiceprint recognition results containing speaker identity information, emotion recognition results containing emotion category and confidence level, and abnormal sound detection results containing abnormal sound event category and confidence level, respectively. The voiceprint recognition result, the emotion recognition result, and the abnormal sound detection result are input into the multimodal fusion decision module, which makes a decision based on the preset fusion rules and outputs a warning level associated with the bullying risk. Warning information is generated and pushed out based on the warning level.

2. The method for campus acoustic monitoring and privacy protection-based bullying early warning based on microphone array as described in claim 1, characterized in that, The multi-channel audio signal is subjected to acoustic front-end processing to output a high signal-to-noise ratio audio stream, specifically including... Performs DC rejection, automatic gain control, wind noise and power supply noise suppression, and echo cancellation; Furthermore, the sound source arrival direction is estimated based on the GCC-PHAT and SRP-PHAT algorithms, and beamforming is guided based on the estimation results.

3. The method for campus acoustic monitoring and privacy protection-based bullying early warning based on microphone array as described in claim 1, characterized in that, Extracting acoustic features from the high signal-to-noise ratio audio stream specifically includes: Deep VAD is used for speech activity detection to determine the speech segment to be processed; the speech segment is framed and windowed; 80-dimensional log-Mel filter bank features are calculated based on the windowed speech frames, and the cepstral mean-variance is normalized on the features.

4. The method for campus acoustic monitoring and privacy protection-based bullying early warning based on microphone array as described in claim 1, characterized in that, The voiceprint recognition module uses an x-vector network architecture to extract the speaker embedding vector and combines the sound source arrival direction angle with the position matching relationship of the corresponding voiceprint area to improve the recognition confidence.

5. The method for campus acoustic monitoring and privacy protection-based bullying early warning based on microphone array as described in claim 1, characterized in that, The emotion recognition module uses a CRNN network for classification and outputs a stable emotion state based on a sliding window hysteresis control strategy. The strategy includes: when the proportion of frames belonging to a preset high-risk emotion in a decision window of length N frames exceeds a first threshold, the system state flips from normal to abnormal and locks the abnormal state for at least T seconds.

6. The method for campus acoustic monitoring and privacy protection-based bullying early warning based on microphone array as described in claim 1, characterized in that, The abnormal sound detection module adopts a multi-label sound event detection framework, uses a lightweight CRNN network for streaming inference, identifies event types including crying, screaming, falling and impact, and integrates an open set anomaly detection mechanism.

7. The method for campus acoustic monitoring and privacy protection-based bullying early warning based on microphone array as described in claim 1, characterized in that, The preset fusion rule in the multimodal fusion decision module is a multi-evidence triggering logic, which includes: when a high-risk emotion with a confidence level higher than the first threshold is continuously detected, a physical abnormal sound event is detected at the same time, and the voiceprint identity is confirmed, it is determined to be a high-risk warning level.

8. A campus acoustic monitoring and privacy protection bullying early warning system based on a microphone array, used to implement the campus acoustic monitoring and privacy protection bullying early warning method based on a microphone array as described in claim 1, characterized in that, The system includes a microphone array, an acoustic front-end processing module, an acoustic feature extraction module, a voiceprint recognition module, an emotion recognition module, an abnormal sound detection module, and a multimodal fusion decision module. The acoustic front-end processing module is connected to the microphone array, the acoustic feature extraction module is connected to the acoustic front-end processing module, the voiceprint recognition module, the emotion recognition module, and the abnormal sound detection module are each connected to the acoustic feature extraction module, and the multimodal fusion decision module is connected to the voiceprint recognition module, the emotion recognition module, and the abnormal sound detection module. The microphone array is used to collect multi-channel audio signals from indoor spaces on campus. The acoustic front-end processing module is used to perform de-reverberation and beamforming processing on the audio signal; The acoustic feature extraction module is used to extract features from the processed audio stream and distribute them. The voiceprint recognition module, the emotion recognition module, and the abnormal sound detection module are used to generate recognition results for voiceprint, emotion, and abnormal sound, respectively. The multimodal fusion decision module is used to fuse multimodal recognition results according to preset rules, output a warning level decision, and generate and send warning information based on the decision.