Campus spoofing intelligent prevention and control Internet of Things system with grid management
By using high-precision microphone array and sliding time window division technology in the campus bullying monitoring system, combining beam automatic pointing and width adaptive adjustment, the difficulty of identifying cry sounds in the prior art in the context of group noise is solved, efficient capture and accurate identification of potential instantaneous cry signals is achieved, and real-time discovery and early warning capabilities of bullying events are improved.
Patent Information
- Application Number
- CN202510550223.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-29
AI Technical Summary
The existing campus bullying monitoring system is difficult to accurately separate and identify victims' abnormal crying sounds in the context of group noise, resulting in missed detection and misjudgment, affecting real-time discovery and early warning capabilities.
A multi-point distribution high-precision microphone array combines sliding time window division and beam automatic pointing and width adaptive adjustment, and spatial focus and noise suppression of potential instantaneous crying signals are achieved through audio feature extraction and machine learning analysis.
It significantly improves the real-time discovery, accurate warning and rapid response capabilities of campus bullying incidents, ensures clear collection and reliable identification of victims' abnormal signals, and effectively suppresses noise and interference around them.
Smart Images

Figure CN120388582A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of campus bullying prevention and control, and particularly relates to an intelligent prevention and control Internet of Things system for campus bullying with grid management. Background Art
[0002] The intelligent prevention and control Internet of Things system for campus bullying with grid management is a campus security guarantee system that integrates the concepts of the Internet of Things, artificial intelligence, and grid management. The system divides several "grid" units within the campus (such as areas like teaching buildings, playgrounds, dormitories, and canteens), and deploys Internet of Things devices such as video monitoring, audio collection, behavior recognition, environmental perception, and intelligent terminals within each grid to collect and sense students' behavior and environmental information in real time. Based on technologies such as AI image recognition, abnormal sound detection, and group behavior analysis, the system can automatically identify suspected bullying incidents, and combined with students' trajectories, social networks, and historical data, accurately warn of abnormal situations. At the same time, according to the grid division, the management responsibilities are implemented to specific management personnel or security patrol posts, forming a collaborative mechanism of "manpower prevention + technical prevention" to achieve real-time discovery, rapid intervention, and subsequent tracking of bullying behaviors, effectively improving the prevention and control ability and disposal efficiency of campus bullying, and ensuring the safety and physical and mental health of students.
[0003] The existing technologies have the following deficiencies: In the campus bullying scenario, bullies and onlookers often deliberately or inadvertently form a large amount of background noise through collective noise-making, heckling, screaming, laughing, etc., thus masking the abnormal crying sounds emitted by victims in the sound field. Since crying sounds usually have weak energy, short duration, and spectral characteristics that partially overlap with environmental noise, the existing monitoring systems based on abnormal sound detection are difficult to accurately separate and identify such abnormal signals in the context of group noise, often misjudging the audio containing the victim's crying as ordinary collective activities or normal student noise-making behaviors, resulting in missed detections and misjudgments of real bullying incidents, seriously affecting the system's real-time discovery and early warning capabilities for campus bullying incidents, and potentially delaying subsequent intervention and rescue opportunities, or even causing serious safety consequences.
[0004] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure, and thus it may include information that does not constitute prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] The object of the present invention is to provide an intelligent prevention and control Internet of Things system for campus bullying with grid management. By introducing a high-precision microphone array with multi-point distribution, it realizes the all-round and multi-channel perception of audio information in a large scene; through the division of sliding time windows and the analysis of micro time segments, it significantly improves the capture ability of short-time and weak-energy abnormal signals; secondly, combined with the automatic beam pointing based on credibility and the adaptive adjustment of beam width, it further enhances the spatial focusing and noise suppression ability of potential instantaneous crying signals, ensures that the sound source can be quickly locked and the signal-to-noise ratio can be enhanced when the credibility of the target signal increases, effectively suppresses the surrounding noise interference, guarantees the clear acquisition and reliable identification of abnormal signals of victims, and improves the real-time detection, accurate early warning and rapid response capabilities of the system for campus bullying events, so as to solve the problems in the above-mentioned background technology.
[0006] To achieve the above object, the present invention provides the following technical solution: An intelligent prevention and control Internet of Things system for campus bullying with grid management, including an audio acquisition module, an audio data windowing module, a feature extraction and determination module, a beam pointing control module, a beam width adjustment module, and a deep analysis and alarm module; The audio acquisition module synchronously acquires multi-channel audio signals in the environment through a high-precision microphone array, and obtains sound information from different directions and different distances; The audio data windowing module divides the acquired audio data in the original large scene into multiple consecutive time windows according to the sliding window mode, so that the audio detection is decomposed into multiple local time-domain analysis tasks; The feature extraction and determination module extracts features and performs machine learning analysis on the audio signal in each time window, and calculates the credibility within the micro time segment, and judges whether there is a potential instantaneous crying signal in the audio signal through the credibility data; When the beam pointing control module recognizes that there is a potential instantaneous crying signal in a certain time window, it controls the beamforming unit of the microphone array to automatically point the main lobe of the beam to the direction of the potential instantaneous crying signal, so that the subsequent audio acquisition can spatially focus on the sound source direction of the potential instantaneous crying signal; When the main lobe of the beam automatically points to the direction of the potential instantaneous crying signal, the beam width adjustment module adaptively adjusts the beam width formed by the microphone array according to the credibility of the potential instantaneous crying signal; The deep analysis and alarm module acquires the audio signal from the target direction, and performs real-time deep analysis on the audio data in this direction, further confirms whether there is a campus bullying event, and outputs an alarm signal when the trigger condition is met.
[0007] Preferably, the specific steps of dividing the acquired audio data in the original large scene into multiple consecutive time windows are as follows: Set the fixed length and sliding step of the time window; Starting from the starting moment of the audio data, intercept the first time window with the set window length; Advance forward with the sliding step to intercept the next time window, that is, there is partial overlap between the new window and the previous window; And so on, repeat the sliding and intercepting operations until the entire audio stream is covered.
[0008] Preferably, feature extraction is performed on the audio signal. Among them, the extracted features include the amplitude of the rapid change of the spectral centroid over time and the proportion of non-linear acoustic components in the audio signal. In each time window, after feature engineering processing on the extracted features, a spectral flutter reference value and a non-linear component reference value are generated respectively, and the potential instantaneous crying signal in the audio signal is preliminarily quantified through the spectral flutter reference value and the non-linear component reference value.
[0009] Preferably, the processed spectral flutter reference value and non-linear component reference value are used as feature vectors and input into a pre-trained support vector machine model. Based on the output of the support vector machine model, a credibility coefficient is obtained, and whether there is a potential instantaneous crying signal in the micro time segment is judged through the credibility coefficient.
[0010] Preferably, the credibility coefficient generated when predicting the potential instantaneous crying signal in the micro time segment by the pre-trained support vector machine model is compared and analyzed with the pre-set credibility coefficient reference threshold to judge whether there is a potential instantaneous crying signal in the audio signal. The specific judgment steps are as follows: If the credibility coefficient is greater than the credibility coefficient reference threshold, an abnormal signal is generated, indicating that there is a potential instantaneous crying signal in the audio signal; If the credibility coefficient is less than or equal to the credibility coefficient reference threshold, a normal signal is generated, indicating that there is no potential instantaneous crying signal in the audio signal.
[0011] Preferably, when it is recognized that there is a potential instantaneous crying signal in a certain time window, that is, when the credibility coefficient generated in the time window is greater than the credibility coefficient reference threshold, the spatial position of the potential signal source is located through the generalized cross-correlation technique, and the beamforming unit of the microphone array is controlled to automatically direct the main lobe of the beam to the direction of the potential instantaneous crying signal. The specific steps are as follows: Based on the features extracted in the time window and the prediction result, when it is detected that there is a potential instantaneous crying signal, the sound source localization process is immediately triggered; Adopt the generalized cross-correlation technique, combine the audio signals received by each channel in the microphone array, and determine the specific azimuth angle or three-dimensional coordinates of the potential instantaneous crying signal in space by calculating the time delay difference between each channel; Control the beamforming unit in the microphone array. According to the located sound source position, adjust the beamforming weights, and direct the main lobe of the beam to the potential signal source in real time. In the newly formed beam direction, continue to perform high-precision acquisition of the audio signals in the target area.
[0012] Preferably, adaptively adjust the beam width formed by the microphone array according to the credibility of the potential instantaneous crying signal. The specific steps are as follows: Obtain the credibility coefficient of the potential instantaneous crying signal within the current time window, compare it with the preset credibility coefficient reference threshold, and according to the comparison result, adaptively adjust the beam width using the following adaptive beam width control formula: , where: represents the current beam width of the beam, is the maximum beam width preset by the system; is the beam narrowing rate control coefficient, which controls the narrowing speed of the beam width as the credibility increases when the credibility coefficient exceeds the credibility coefficient reference threshold; is the credibility coefficient of the instantaneous crying signal; is the credibility coefficient reference threshold; Introduce a smoothing regulation mechanism to dynamically update the beam width. Specifically: , where: is the actually used beam width of the current frame, is the beam width of the previous frame, is the beam width smoothing coefficient, .
[0013] Preferably, within each time window, the specific steps for generating the spectral tremor reference value after performing feature engineering on the amplitude of the spectral centroid that changes rapidly with time are as follows: Within the current time window, extract the spectral centroid of each frame of audio, obtain the spectral centroid volatility sequence through the first-order difference absolute change rate of the spectral centroid sequence, and the spectral centroid volatility sequence is used to represent the relative change amplitude of the spectral centroid; After obtaining the spectral centroid volatility sequence, introduce a logarithmic weighted integral strategy to generate the spectral tremor reference value.
[0014] Preferably, within each time window, the specific steps for generating the non-linear component reference value after performing feature engineering on the proportion of non-linear acoustic components in the audio signal are as follows: Within each time window, frame the audio signal, perform spectral decomposition on each frame of the signal, extract the frequency-domain energy distribution of the non-linear components, and measure the disorder degree of the non-linear components in the signal by calculating the spectral structure entropy of this energy distribution to obtain an intermediate variable for subsequent quantization; The obtained characteristic entropy value is compared with the overall spectral energy ratio of the signal to form a normalized non-linear component reference value, which is used to identify the relative proportion and significance degree of the non-linear component within the current time window.
[0015] In the above technical solution, the technical effects and advantages provided by the present invention are as follows: The present invention realizes the all-round and multi-channel perception of audio information in a large scene by introducing a high-precision microphone array with multi-point distribution; through the sliding time window division and micro-time segment analysis, the capture ability of short-time and weak-energy abnormal signals is significantly improved; secondly, combined with the beam automatic pointing based on credibility and the beam width adaptive adjustment, the spatial focusing and noise suppression ability of potential instantaneous crying signals are further enhanced, ensuring that the sound source can be quickly locked and the signal-to-noise ratio can be strengthened when the credibility of the target signal increases, effectively suppressing the surrounding noise interference, ensuring the clear acquisition and reliable identification of the abnormal signals of the victim, and improving the real-time detection, precise early warning and rapid response capabilities of the system for campus bullying incidents. Brief Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0017] Figure 1 It is a module schematic diagram of an intelligent prevention and control Internet of Things system for campus bullying with grid management according to the present invention. Detailed Embodiments
[0018] Now, the example embodiments will be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these example embodiments are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the example embodiments to those skilled in the art.
[0019] The present invention provides an intelligent prevention and control Internet of Things system for campus bullying with grid management as Figure 1 shown, including: By arranging a high-precision microphone array with multi-point distribution in a large scene, multi-channel synchronous acquisition of audio signals in the environment is carried out to obtain sound information from different directions and different distances, realizing the comprehensive audio perception coverage of the target area in the time and space dimensions; the large scene can include typical campus public or semi-public areas such as playgrounds, classrooms, corridors, dormitories, etc. A high-precision microphone should have good directivity, sensitivity, and signal-to-noise ratio to ensure that sufficient clear original sound signals can be obtained in scenarios with complex noise and dense crowds.
[0020] The audio data in the original large scene collected is divided into multiple consecutive time windows according to the sliding window mode. Each time window can be 1 to 3 seconds, and it can be flexibly adjusted according to the scene and hardware configuration. By dividing the time windows, the audio detection is decomposed into multiple local time-domain analysis tasks, highlighting the detectability of weak and short abnormal signals, so that subsequent local audio features can be detected on a finer time scale; Compared with directly analyzing global audio features, time window division can better capture instantaneous and short abnormal crying signals in complex backgrounds, improving the sensitivity and detection accuracy of the system to abnormal signals.
[0021] The specific steps to divide the audio data in the original large scene collected into multiple consecutive time windows are as follows: First, set the fixed length (such as 1 second) and sliding step (such as 0.5 second) of the time window; then, starting from the starting moment of the audio data, intercept the first time window with the set window length; then, advance forward with the sliding step and intercept the next time window, that is, there is partial overlap between the new window and the previous window; and so on, repeating the sliding and intercepting operations until the entire audio stream is covered. This method can achieve local fine-grained analysis of continuous audio, maintaining both temporal continuity and the ability to capture weak or instantaneous abnormal signals that may occur in a short time.
[0022] Within each time window, feature extraction and machine learning analysis are performed on the audio signal, and the credibility within micro time segments (such as sub-segments within 0.5 seconds and 1 second) is calculated. Whether there are potential instantaneous crying signals in the audio signal is judged through the credibility data, which serves as a reference basis for subsequent abnormal signal determination; Feature extraction is performed on the audio signal. Among them, the extracted features include the amplitude of the rapid change of the spectral centroid over time and the proportion of non-linear acoustic components in the audio signal. Within each time window, after feature engineering processing on the extracted features, a spectral flutter reference value and a non-linear component reference value are generated respectively, and the potential instantaneous crying signals in the audio signal are preliminarily quantified through the spectral flutter reference value and the non-linear component reference value.
[0023] Rapid, time-varying changes in the spectral center of gravity (CMG) of an audio signal with large amplitudes typically indicate the presence of a potential transient crying signal. Compared to general speech, noise, or laughter, crying signals are typically unstable and emotionally driven. During crying, victims often experience physiological changes such as emotional fluctuations, respiratory disturbances, and vocal cord tremors, causing the spectral center of gravity of their voice to fluctuate rapidly and irregularly. This rapid change is reflected in the signal as a significant increase in the amplitude of the spectral center of gravity's change over time. In contrast, ordinary noise, conversation, and even laughter, despite their high energy, typically exhibit slow and stable changes in their spectral center of gravity, representing relatively stable vocal behavior. Therefore, detecting significantly large fluctuations in the spectral center of gravity within a short time window often indicates the presence of a short, unstable, and emotionally abnormal crying signal, which can serve as an effective indicator of potential anomaly.
[0024] The specific steps for generating spectrum jitter reference values after feature engineering the amplitude of the spectrum center of gravity that changes rapidly over time in each time window are as follows: First, within the current time window, extract the spectral center of each frame of audio (e.g., every 10ms frame), which is recorded as sequence C. , n represents the total number of audio frames in the time window, where Represents the spectrum center of gravity of the i-th frame, and then the spectrum center of gravity fluctuation rate sequence D is obtained through the first-order difference absolute change rate of the spectrum center of gravity sequence. , the spectrum center of gravity volatility series is used to express the amplitude of the relative change of the spectrum center of gravity, and the calculation expression is: ,in: and is the spectral center of adjacent frames, is the amplitude enhancement factor of the relative change of the center of gravity of the spectrum of the i-th frame; The spectral center of gravity is extracted for each frame of audio (e.g., every 10ms frame). This spectral representation is obtained by performing a short-time Fourier transform (STFT) on the audio signal, and then calculating the amplitude-weighted average frequency value of each frequency component in the frequency domain. Specifically, the amplitude of each frequency component is considered to be the weight of that frequency, and the spectral center of gravity is the ratio of the sum of the products of all frequencies and their amplitudes to the total amplitude, thus reflecting the "center position" of the audio energy in the frequency spectrum for that frame. This method can effectively capture auditory characteristics such as sound brightness and sharpness, and is highly valuable for analyzing the shifts in the spectrum of crying sounds caused by emotional fluctuations.
[0025] The purpose of this step is to extract the instantaneous fluctuation characteristics of the spectrum center of gravity as a ratio indicator with the ability to amplify mutations, thereby enhancing the ability to identify violent spectrum swings in sudden crying signals. It is particularly suitable for identifying short-term irregular jitters rather than trend changes.
[0026] After obtaining the spectral centroid volatility sequence D, a logarithmic weighted integral strategy is introduced to generate a spectral flutter reference value, and the generated expression is: , where: represents the spectral flutter reference value, is the flutter sensitivity factor (such as taking 5 - 10), which is used to adjust the sensitivity to slight changes, is used to enhance the influence of non - linear jumps and suppress the influence of small - amplitude fluctuations, is the time decay coefficient (such as 0.95), which is used to assign higher weights to fluctuations closer to the current moment. The exponential form ensures that recent fluctuations dominate the overall flutter perception and is suitable for identifying short - time bursty spectral perturbations in crying; The spectral flutter reference value combines non - linear response and time sensitivity, and can effectively capture and quantify typical crying signal features such as rapid spectral jitter and instantaneous perturbations within a short time window, suppress errors caused by background steady changes, and highlight the instantaneous occurrence of structural crying features at the same time.
[0027] From the spectral flutter reference value, it can be seen that within each time window, the larger the performance value of the spectral flutter reference value generated after feature engineering of the amplitude of the rapid change of the spectral centroid over time, the greater the risk of potential instantaneous crying signals in the audio signal. The reason is that crying signals essentially belong to an abnormal sound with non - stationarity and emotion - driven characteristics, usually manifested as violent fluctuations and rapid changes of the spectral centroid within a short time. Therefore, after feature engineering, the spectral flutter reference value can directly reflect the strength of this spectral fluctuation. Specifically, when a potential instantaneous crying signal appears in the audio signal, due to vocal cord tremors, emotional instability, and irregular energy release during crying, the spectral centroid between adjacent frames shows a large - amplitude non - stationary jump, which in turn increases the change rate of the first - order difference, and finally results in a relatively large spectral flutter reference value after calculation; conversely, if the audio signal source is ordinary noise, talking, or environmental noise, the spectral centroid usually changes smoothly with a small fluctuation amplitude, and the spectral flutter reference value will also be significantly reduced. Therefore, this spectral flutter reference value can effectively quantify the risk level of potential instantaneous crying signals.
[0028] A significant increase in the proportion of non-linear acoustic components in an audio signal usually indicates the presence of a potential instantaneous crying signal, which is determined by the physiological and emotional characteristics of crying behavior. Specifically, when crying, the vocal cords of the victim tend to vibrate involuntarily, tensely, disorderly, or under strong emotional drive, with irregular glottal opening and closing and unstable airflow, which easily leads to obvious non-linear acoustic components in the audio signal, such as creaking sounds, unstable vocal cord oscillations, spectral jumps, and enhanced non-periodic components. These non-linear components rarely appear in common group noise environments such as normal conversations, laughter, and heckling, and thus can be used as significant indicators for judging abnormal crying signals. Especially in instantaneous crying signals, although the energy level may be low, the proportion of their non-linear components is often relatively high, becoming an important feature that differentiates them from general noise or normal speech. Therefore, detecting an abnormal increase in the proportion of non-linear components in the audio often serves as an effective basis for predicting potential instantaneous crying signals.
[0029] The specific steps for generating a reference value for non-linear components by performing feature engineering on the proportion of non-linear acoustic components in the audio signal within each time window are as follows: Within each time window, the audio signal is framed, and the frequency spectrum of each frame of the signal is decomposed to extract the frequency-domain energy distribution of non-linear components. By calculating the spectral structure entropy (structural complexity) of this energy distribution, the degree of disorder of non-linear components in the signal is measured, obtaining an intermediate variable for subsequent quantification. The calculation expression is: , where: is the non-linear acoustic feature entropy, used to represent the structural complexity of the distribution of non-linear components in the frequency domain, is the normalized energy density of non-linear components in the k-th frequency band (which can be calculated through CEPSTRUM inverse transform, harmonic regression residuals, or Teager-Kaiser energy operator), and N is the total number of frequency bands; Measures the "discreteness" and "mutability" of non-linear energy in the frequency domain in the form of entropy. The non-periodicity, high-frequency perturbations, and irregular components in crying sounds will result in higher entropy values, thus having strong discriminative ability.
[0030] Compare the obtained feature entropy value with the overall spectral energy ratio of the signal to form a normalized reference value for non-linear components, which is used to identify the relative proportion and significance degree of non-linear components within the current time window. The generation expression of the reference value for non-linear components is: , where: is the reference value for non-linear components, , are hyperparameters for adjusting the entropy sensitivity and non-linear contrast amplitude (which can take values of 1.2 and 1.0 respectively), is the signal background suppression term, which controls the sudden increase of the reference value for non-linear components in a low-entropy environment to prevent false alarms. It can avoid the denominator failure caused by extremely small entropy values; The reference value of the non-linear component is obtained by non-linearly mapping and normalizing the characteristic entropy, enabling the system to stably identify audio segments with a high proportion of non-linearity under different noise intensities and background complexities, thereby preliminarily quantifying potential instantaneous crying signals.
[0031] From the reference value of the non-linear component, it can be seen that within each time window, the larger the performance value of the reference value of the non-linear component generated after feature engineering of the proportion of the non-linear acoustic component in the audio signal, the greater the risk of potential instantaneous crying signals in the audio signal. The main reason is the significant non-linear performance of crying signals in acoustic characteristics. As an emotional and involuntary sound, crying often shows irregular, disordered, intermittent, and ruptured phenomena in vocal cord vibration, resulting in a large number of non-periodic and non-stationary components in the audio signal, making the signal exhibit obvious non-linear characteristics in terms of spectrum, phase, or energy structure. Relatively speaking, although ordinary speech, noise, laughter, etc. may have high energy, their periodic and harmonic structures are relatively stable, and the proportion of non-linear components is relatively low. Therefore, when the reference value of the non-linear component calculated after feature engineering of the audio signal shows a large value, it often indicates that there are a large number of irregular and non-linear transient changes in the signal, which exactly conforms to the typical characteristics of crying signals in actual acoustic performance, and thus can be used as an important criterion for potential instantaneous crying signals.
[0032] The processed spectral jitter reference value and the reference value of the non-linear component are used as feature vectors and input into a pre-trained support vector machine model. Based on the output credibility coefficient of the support vector machine model, it is judged whether there are potential instantaneous crying signals in the microscopic time segment.
[0033] A pre-trained support vector machine model refers to a model that has undergone a series of offline or preliminary training steps before actual system deployment. In this process, a large amount of labeled data is used to train the Support Vector Machine (SVM) algorithm and optimize its parameters. Specifically, researchers first collect audio samples in various scenarios, including normal environmental sounds, noisy voices of multiple people, and various types of crying audio, whether artificially or realistically recorded. For each sample, the possible instantaneous crying signals are labeled. Through this labeling process, a training dataset containing "feature-label" pairs is obtained. Next, after feature extraction and feature engineering on these original audio data (for example, integrating spectral jitter reference values, non-linear component reference values, and other possible audio features), labels indicating "whether there is a crying signal" are assigned to different audio segments. Subsequently, researchers use these labeled feature samples to train the SVM model, aiming to enable it to learn which feature distributions in the high-dimensional feature space correspond to the target category of "instantaneous crying", and which feature distributions correspond to the categories of "no crying" or "background noise". During the training process, the algorithm continuously iteratively adjusts the position and shape of the hyperplane to maximize the margin between different categories, achieving optimal discrimination between crying and non-crying features. After training, the model needs to be evaluated and fine-tuned using a validation set to ensure that the model can still effectively distinguish instantaneous crying signals from other environmental sounds in complex environments. Only when the model achieves a sufficiently high accuracy and robustness on the validation or test set will it be solidified as the final "pre-trained support vector machine model" and loaded into the actual system during the deployment phase to perform real-time audio discrimination tasks.
[0034] After completing pre-training and passing the verification, this support vector machine model will be integrated into the online discrimination process of the system. When the real-time audio stream is segmented into multiple time windows, the system extracts relevant features in each time window, such as the spectral jitter reference value and the non-linear component reference value, and combines them into a feature vector to input into the trained support vector machine model. At this time, the support vector machine model will map and calculate the input vector based on the hyperplane or decision boundary it has learned, and output a classification result of the target category and the corresponding support degree value or confidence level. To meet the requirements of crying detection, the system will convert the output of the support vector machine model into a readable credibility coefficient, and the higher the value, the more likely there is an instantaneous crying signal in the current time window (or even more microscopic sub-segments). Compared with the traditional volume or spectral detection scheme based on a fixed threshold, the pre-trained support vector machine model can better adapt to diverse and complex noise environments, and has higher discriminative ability in distinguishing real crying sounds from other similar but non-crying sounds (such as screams, laughter, interference noises, etc.), thus significantly reducing the risks of missed detection and false detection.
[0035] The machine learning model is not specifically limited here, and any deep learning model that can implement comprehensive analysis of the spectral jitter reference value and the non-linear component reference value to generate a credibility coefficient is acceptable. To implement the technical solution of the present invention, the present invention provides a specific implementation method; the expression for generating the credibility coefficient is: , where, , are respectively the preset proportionality coefficients of the spectral jitter reference value and the non-linear component reference value , and , are both greater than 0. The preset proportionality coefficient refers to the weight coefficients given to different feature components (i.e., the spectral jitter reference value and the non-linear component reference value ) in the overall score when calculating the credibility coefficient, denoted as and . Since the spectral jitter reference value and the non-linear component reference value are two features with different sources and physical meanings, and their contribution degrees to the final credibility coefficient may vary, so it is necessary to artificially preset (or determine through experiments) their weights in synthesizing the credibility coefficient The proportion at a certain time is used to avoid an unreasonable dominant effect of a certain feature on the result due to differences in the value range or sensitivity. In other words, the role of the preset proportional coefficient is to perform weighted balancing on the relative importance, signal-to-noise ability, and recognition contribution of the two features to ensure that the finally generated credibility coefficient is more in line with the abnormal detection requirements in actual applications. Usually, these two coefficients can be adjusted according to historical data, offline training, or empirical formulas to ensure that in the actual environment, the joint discrimination ability of the two types of features for identifying abnormal crying signals can be comprehensively exerted.
[0036] It can be seen from the credibility coefficient that within each time window, the larger the performance value of the spectral tremor reference value generated after feature engineering processing of the amplitude of the rapid change of the spectral centroid over time, and the larger the performance value of the nonlinear component reference value generated after feature engineering processing of the proportion of the nonlinear acoustic component in the audio signal. That is, when predicting potential instantaneous crying signals in micro time segments through a pre-trained support vector machine model, the larger the performance value of the generated credibility coefficient, the greater the risk that there is a potential instantaneous crying signal in the audio signal, and vice versa, it indicates that the risk of a potential instantaneous crying signal in the audio signal is smaller.
[0037] Compare and analyze the credibility coefficient generated when predicting potential instantaneous crying signals in micro time segments through a pre-trained support vector machine model with a pre-set credibility coefficient reference threshold to determine whether there is a potential instantaneous crying signal in the audio signal. The specific judgment steps are as follows: If the credibility coefficient is greater than the credibility coefficient reference threshold, an abnormal signal is generated, indicating that there is a potential instantaneous crying signal in the audio signal. If the credibility coefficient is less than or equal to the credibility coefficient reference threshold, a normal signal is generated, indicating that there is no potential instantaneous crying signal in the audio signal.
[0038] When a potential instantaneous crying signal is detected in a certain time window, control the beamforming unit of the microphone array to automatically direct the main lobe of the beam (i.e., the auditory focus) to the direction of the potential signal, so that subsequent audio acquisition can spatially focus on the sound source direction of the potential instantaneous crying signal. In a complex acoustic environment, actively control the beam to focus on the area where the potential victim is located, effectively enhancing the target signal energy and suppressing background interference such as noise and heckling from other directions; When a potential instantaneous crying signal is detected in a certain time window, that is, when the credibility coefficient generated within the time window is greater than the credibility coefficient reference threshold, locate the spatial position of the potential signal source through the generalized cross-correlation technique, and control the beamforming unit of the microphone array to automatically direct the main lobe of the beam (i.e., the auditory focus) to the direction of the potential signal. The specific steps are as follows: First, based on the features extracted within the time window and the prediction results, when the system detects a potential instantaneous crying signal, it immediately triggers the sound source localization process; Secondly, the Generalized Cross-Correlation (GCC) technique is adopted. By combining the audio signals received by each channel in the microphone array and calculating the time delay differences between channels, the specific azimuth angle or three-dimensional coordinates of the potential instantaneous crying signal in space are determined; Next, the beamforming unit in the microphone array is controlled. According to the located sound source position, the beamforming weights are adjusted to direct the main lobe of the beam (i.e., the auditory focus of the array) towards the potential signal source in real time; Finally, the system continues to perform high-precision acquisition and subsequent processing on the audio signals in the target area in the newly formed beam direction.
[0039] Through the above steps, the system can focus on collecting audio signals only in the target direction in a large-scale, multi-source background, effectively enhancing the energy and feature expression of potential instantaneous crying signals, while suppressing non-target noises such as noise and heckling from other directions, providing audio data with a higher signal-to-noise ratio for subsequent signal analysis, classification, and anomaly discrimination.
[0040] When the main lobe of the beam (i.e., the auditory focus) automatically points to the direction of the potential signal, the beam width formed by the microphone array is adaptively adjusted according to the credibility of the potential instantaneous crying signal. When the credibility of the identified potential instantaneous crying signal is relatively high, the beam width is narrowed to form a highly directional beam to further focus on collecting abnormal signals in this direction; when the credibility is relatively low or the signal is still unstable, a wide beam is adopted to maintain continuous attention to the surrounding area and prevent missed detections caused by signal transients or target position offsets; The beam width formed by the microphone array is adaptively adjusted according to the credibility of the potential instantaneous crying signal. The specific steps are as follows: Obtain the credibility coefficient of the potential instantaneous crying signal within the current time window, compare it with the pre-set credibility coefficient reference threshold, and according to the comparison result, adaptively adjust the beam width using the following adaptive beam width control formula: , where: represents the current beam width of the beam, is the maximum beam width preset by the system (i.e., the widest beam); is the beam narrowing rate control coefficient, which controls the narrowing speed of the beam width as the credibility coefficient increases when the credibility coefficient exceeds the credibility coefficient reference threshold; is the credibility coefficient of the instantaneous crying signal; is the credibility coefficient reference threshold; This step dynamically calculates the beam width by monitoring the relative relationship between the credibility coefficient and the credibility coefficient reference threshold in real time When the confidence coefficient is less than or equal to the confidence coefficient reference threshold, , that is, maintain the wide beam; when the confidence coefficient is greater than the confidence coefficient reference threshold, the beam width begins to narrow, and the higher the confidence coefficient, the narrower the beam and the stronger the directivity; the role of this mechanism is: to maintain large-range detection in the low-confidence stage to prevent the loss of the target caused by unstable signals or position drift in the initial stage; in the high-confidence stage, actively narrow the beam to enhance the focusing ability and signal-to-noise ratio of potential abnormal signals and improve the subsequent recognition accuracy.
[0041] To avoid frequent oscillations of the system caused by fluctuations in the instantaneous confidence coefficient of the beam width, a smoothing control mechanism is further introduced to dynamically update the beam width. Specifically: , where: is the actual beam width used in the current frame (or time window), is the beam width of the previous frame, is the beam width smoothing coefficient, ; This mechanism can suppress the drastic change in the beam width caused by short-term fluctuations in the confidence coefficient, enabling the beam width to be adjusted smoothly and continuously with the change of the confidence. The finally obtained actual beam width will be directly used in the beamformer to control the directivity of the beam formed by the array. Its role is: when the signal confidence changes rapidly, to avoid the phenomenon of drastic beam contraction or expansion of the system due to short-term high confidence or low confidence, and ensure the smoothness and practicality of the beam width adaptive adjustment process. At the same time, the smoothing control enables the system to have better robustness and anti-misjudgment ability when facing weak signals, displacement signals or unstable signals, ensuring that the system can more reliably track and enhance potential crying signals.
[0042] On the basis of forming a high-directivity beam, the system focuses on collecting audio signals from the target direction and performs real-time in-depth analysis on the audio data in this direction to further confirm whether there is a campus bullying incident and output an alarm signal when the trigger condition is met; The high-directivity beam can greatly improve the energy ratio of the target signal to the background noise, enabling the subsequent analysis link to accurately identify short crying signals even in a noisy group environment, ensuring the accuracy and timeliness of the system's anomaly detection. Based on the high signal-to-noise ratio audio input with dynamic focusing, it provides a solid signal foundation for the real-time monitoring, anomaly warning, intelligent analysis, etc. of bullying incidents, significantly improving the recognition effect of the system in a real and complex environment.
[0043] Through the above solution, it is possible to effectively solve the problem that the existing audio monitoring system for campus bullying is insensitive to weak and short crying signals in a noisy and complex background, and achieve high-precision recognition and precise focusing on abnormal signals of potential victims in a noisy environment. By introducing a high-precision microphone array with multi-point distribution, it realizes omni-directional and multi-channel perception of audio information in a large scene; through the division of sliding time windows and the analysis of microscopic time segments, it significantly improves the capture ability for short-time and weak-energy abnormal signals; combined with beam automatic pointing based on credibility and adaptive adjustment of beam width, it further enhances the spatial focusing and noise suppression ability for potential instantaneous crying signals, ensuring that the sound source can be quickly locked and the signal-to-noise ratio can be enhanced when the credibility of the target signal increases, effectively suppressing the surrounding noise interference, guaranteeing the clear acquisition and reliable recognition of the abnormal signals of the victim, and significantly improving the real-time detection, precise early warning and rapid response capabilities of the system for campus bullying incidents.
[0044] Only some exemplary embodiments of the present invention have been described above by way of illustration. Undoubtedly, for those of ordinary skill in the art, the described embodiments can be modified in various different ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
Claims
1. An intelligent prevention and control Internet of Things system for campus bullying with grid management, characterized in that, It includes an audio acquisition module, an audio data windowing module, a feature extraction and determination module, a beam pointing control module, a beam width adjustment module, and a depth analysis and alarm module; The audio acquisition module synchronously acquires multi-channel audio signals in the environment through a high-precision microphone array to obtain sound information from different directions and different distances; The audio data windowing module divides the acquired audio data in the original large scene into multiple consecutive time windows according to the sliding window mode, decomposing the audio detection into multiple local time-domain analysis tasks; The feature extraction and determination module extracts features and performs machine learning analysis on the audio signal in each time window, calculates the credibility within the micro time segment, and determines whether there is a potential instantaneous crying signal in the audio signal through the credibility data; When a potential instantaneous crying signal is identified in a certain time window, the beam pointing control module controls the beamforming unit of the microphone array to automatically point the main lobe of the beam in the direction of the potential instantaneous crying signal, enabling subsequent audio acquisition to focus spatially on the sound source direction of the potential instantaneous crying signal; When the main lobe of the beam is automatically pointed in the direction of the potential instantaneous crying signal, the beam width adjustment module adaptively adjusts the beam width formed by the microphone array according to the credibility of the potential instantaneous crying signal; The depth analysis and alarm module acquires the audio signal from the target direction, performs real-time depth analysis on the audio data in this direction, further confirms whether there is a campus bullying incident, and outputs an alarm signal when the triggering condition is met.
2. The intelligent prevention and control Internet of Things system for campus bullying with grid management according to claim 1, characterized in that The specific steps for dividing the acquired audio data in the original large scene into multiple consecutive time windows according to the sliding window mode are as follows: Set the fixed length and sliding step of the time window; Starting from the starting moment of the audio data, intercept the first time window with the set window length; Advance forward with the sliding step to intercept the next time window, that is, there is partial overlap between the new window and the previous window; And so on, repeating the sliding and intercepting operations until the entire audio stream is covered.
3. The intelligent prevention and control Internet of Things system for campus bullying with grid management according to claim 1, characterized in that, Feature extraction is performed on the audio signal. Among them, the extracted features include the amplitude of the rapid change of the spectral centroid over time and the proportion of the non-linear acoustic component in the audio signal. In each time window, after feature engineering processing of the extracted features, a spectral flutter reference value and a non-linear component reference value are respectively generated, and the potential instantaneous crying signal in the audio signal is preliminarily quantified through the spectral flutter reference value and the non-linear component reference value.
4. The intelligent prevention and control Internet of Things system for campus bullying with grid management according to claim 3, characterized in that, The processed spectral flutter reference value and non-linear component reference value are used as feature vectors and input into a pre-trained support vector machine model. Based on the output of the support vector machine model, a credibility coefficient is obtained, and it is determined whether there is a potential instantaneous crying signal in the micro time segment through the credibility coefficient.
5. The intelligent prevention and control Internet of Things system for campus bullying with grid management according to claim 4, characterized in that, Compare and analyze the credibility coefficient generated when predicting the potential instantaneous crying signal in the micro time segment by the pre-trained support vector machine model with the pre-set credibility coefficient reference threshold to determine whether there is a potential instantaneous crying signal in the audio signal. The specific judgment steps are as follows: If the credibility coefficient is greater than the credibility coefficient reference threshold, an abnormal signal is generated, indicating that there is a potential instantaneous crying signal in the audio signal; If the credibility coefficient is less than or equal to the credibility coefficient reference threshold, a normal signal is generated, indicating that there is no potential instantaneous crying signal in the audio signal.
6. The intelligent prevention and control Internet of Things system for campus bullying with grid management according to claim 5, characterized in that, When a potential instantaneous crying signal is identified within a certain time window, that is, when the credibility coefficient generated within the time window is greater than the credibility coefficient reference threshold, the spatial position of the potential signal source is located through the generalized cross-correlation technique, and the beamforming unit of the microphone array is controlled to automatically direct the main lobe of the beam towards the direction of the potential instantaneous crying signal. The specific steps are as follows: Based on the features extracted within the time window and the prediction results, when a potential instantaneous crying signal is detected, the sound source localization process is immediately triggered; The generalized cross-correlation technique is used, combined with the audio signals received by each channel in the microphone array. By calculating the time delay difference between each channel, the specific azimuth angle or three-dimensional coordinates of the potential instantaneous crying signal in space are determined; The beamforming unit in the microphone array is controlled. According to the located sound source position, the beamforming weight is adjusted to direct the main lobe of the beam towards the potential signal source in real time; In the newly formed beam direction, the audio signals in the target area are continuously collected with high precision.
7. The intelligent prevention and control Internet of Things system for campus bullying with grid management according to claim 5, characterized in that, According to the credibility of the potential instantaneous crying signal, the beam width formed by the microphone array is adaptively adjusted. The specific steps are as follows: Obtain the credibility coefficient of the potential instantaneous crying signal within the current time window, compare it with the pre-set credibility coefficient reference threshold, and adaptively adjust the beam width according to the comparison result using the following adaptive beam width control formula: , where: represents the current beam width of the beam, is the maximum beam width preset by the system; is the beam narrowing rate control coefficient, which controls the narrowing speed of the beam width as the credibility increases when the credibility coefficient exceeds the credibility coefficient reference threshold; is the credibility coefficient of the instantaneous crying signal; is the credibility coefficient reference threshold; Introduce a smoothing control mechanism to dynamically update the beam width, specifically: , where: is the actually used beam width of the current frame, is the beam width of the previous frame, is the beam width smoothing coefficient, .
8. The intelligent prevention and control Internet of Things system for campus bullying with grid management according to claim 3, characterized in that, In each time window, the specific steps for generating the spectral tremor reference value by performing feature engineering on the amplitude of the spectral centroid that changes rapidly with time are as follows: In the current time window, the spectral centroid of each frame of audio is extracted. The spectral centroid volatility sequence is obtained through the absolute change rate of the first-order difference of the spectral centroid sequence. The spectral centroid volatility sequence is used to represent the relative change amplitude of the spectral centroid; After obtaining the spectral centroid volatility sequence, a logarithmic weighted integration strategy is introduced to generate the spectral tremor reference value.
9. The intelligent prevention and control Internet of Things system for campus bullying with grid management according to claim 3, characterized in that, In each time window, the specific steps for generating the non-linear component reference value by performing feature engineering on the proportion of non-linear acoustic components in the audio signal are as follows: In each time window, the audio signal is framed, and the spectrum of each frame of the signal is decomposed. The frequency-domain energy distribution of the non-linear components is extracted. By calculating the spectral structure entropy of this energy distribution, the disorder degree of the non-linear components in the signal is measured, and an intermediate variable for subsequent quantization is obtained; The obtained feature entropy value is compared with the overall spectral energy ratio of the signal to form a normalized non-linear component reference value, which is used to identify the relative proportion and significance degree of the non-linear components within the current time window.