Speech anti-bullying specific timbre template voiceprint feature extraction system

By constructing a voiceprint feature extraction system based on specific timbre templates and combining deep neural networks and multimodal information fusion, we have achieved accurate identification and robust detection of school bullying behavior. This solves the problems of detection generalization, high false alarm rate and environmental interference in existing technologies, and provides real-time early warning and evidence retention capabilities.

CN122454984APending Publication Date: 2026-07-24SHENZHEN JULONG CHUANGSHI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN JULONG CHUANGSHI TECH CO LTD
Filing Date
2026-06-01
Publication Date
2026-07-24

Smart Images

  • Figure CN122454984A_ABST
    Figure CN122454984A_ABST
Patent Text Reader

Abstract

The present application relates to the technical fields of speech signal processing, pattern recognition and intelligent security and protection, and discloses a specific timbre template voiceprint feature extraction system for speech anti-bullying, comprising: a multi-channel speech signal acquisition and preprocessing module; a speech activity detection and key segment segmentation module; a specific timbre voiceprint feature extraction engine; a bullying-related specific timbre template library construction and management module; a real-time timbre matching and bullying risk assessment module; a multi-modal information fusion and behavior confirmation module; a real-time early warning and evidence preservation module; a self-learning and template adaptive updating module. The present application realizes the leap of the detection target from the generalized abnormal sound to the precise bullying-related timbre by constructing the bullying-related specific timbre template library and designing a special specific timbre voiceprint feature extraction engine, and solves the core problems of detection target generalization, insufficient pertinence and weak voiceprint feature universality and distinguishability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of speech signal processing, pattern recognition and intelligent security technology, specifically to a system for extracting voiceprint features from specific timbre templates for anti-bullying speech. Background Technology

[0002] Traditional campus security monitoring mainly relies on video cameras. However, in areas with sensitive privacy or physical limitations, such as restrooms, dormitory corridors, and stairwell corners, video surveillance is often impossible or inconvenient to install, creating "blind spots" that make these areas hotspots for bullying incidents. To compensate for the shortcomings of video surveillance, audio-based anomaly detection technology is being explored for application in anti-bullying scenarios.

[0003] The existing related technical solutions mainly have the following problems: The detection targets are too generalized and lack specificity: Existing voice security systems mostly detect "abnormal sounds" (such as screams or loud noises) or perform simple keyword recognition. The former has a high false alarm rate (easily mistaking normal playfulness for abnormality), while the latter is difficult to cover due to the diversity and variability of bullying language (such as subtle threats or mocking tones). They lack the ability to specifically model and identify specific vocal patterns strongly associated with bullying behavior (such as trembling cries for help due to fear, sobbing with a crying tone, or aggressive threatening tones).

[0004] Voiceprint features are general but lack discriminative power: Traditional voiceprint recognition technology aims to distinguish the identities of different speakers, and the extracted features focus on individual uniqueness. However, in anti-bullying scenarios, the key is not "who is speaking," but "whether the manner of speaking (timbre) reveals bullying or victimization characteristics." The timbre characteristics of the same person differ significantly when in normal conversation, in angry threats, or in fearful crying, but general voiceprint features have limited ability to capture and distinguish timbre changes in such contexts.

[0005] Severe environmental interference and poor robustness: The campus environment, especially blind spots such as toilets and corridors, is characterized by complex background noise (flushing sounds, echoes, distant noises). Existing systems exhibit inaccurate voice activity detection (VAD) and degraded feature extraction quality in complex acoustic environments, leading to a sharp decline in detection performance. There is a lack of effective speech front-end processing and robust feature extraction methods for high-noise, high-reverberation environments.

[0006] The decision-making basis is singular, resulting in a high false alarm rate: relying solely on audio monomodal information for judgment makes it highly susceptible to environmental noise interference and intense, non-bullying conversations (such as debates or shouts during sports competitions), leading to alarm fatigue, reduced system reliability, and an inability to be truly put into practical use.

[0007] The system is rigid and inflexible, lacking adaptability: the pre-set detection models and thresholds are difficult to adapt to the differences in speech characteristics and usage habits among students from different schools and age groups. After deployment, the system cannot learn and optimize itself based on actual cases and false alarms, resulting in a decline in detection effectiveness or the need for frequent manual adjustments after long-term use.

[0008] In view of this, we propose a voiceprint feature extraction system based on specific timbre templates for voice anti-bullying. Summary of the Invention

[0009] The purpose of this invention is to provide a specific timbre template voiceprint feature extraction system for anti-bullying voice, so as to solve the problems existing in the above-mentioned background technology.

[0010] To achieve the above objectives, the present invention provides the following technical solution: A specific timbre template voiceprint feature extraction system for voice anti-bullying, the system comprising: The multi-channel speech signal acquisition and preprocessing module is used to acquire multiple environmental audio signals in real time, and to perform noise reduction, echo cancellation, speech enhancement and frame windowing preprocessing on the acquired raw audio signals to obtain a clean speech frame sequence. The speech activity detection and key segmentation module is used to perform endpoint detection on the preprocessed speech frame sequence, distinguish speech segments from non-speech segments, and further segment key speech segments containing suspected bullying semantics or strong emotional features from continuous speech segments based on energy, zero-crossing rate and spectral entropy features. A specific timbre voiceprint feature extraction engine is used to extract high-dimensional voiceprint feature vectors that can characterize the speaker's timbre from the key speech segments. The engine has a built-in deep neural network model that maps time-frequency domain speech features to a highly discriminative voiceprint feature space through multi-layer nonlinear transformation. A module for constructing and managing a bullying-related specific voiceprint template library is used to store and manage specific voiceprint feature templates that are highly related to school bullying behavior. The template library includes at least: a call for help voice template, a crying voice template, a threatening and intimidating voice template, and a mocking and abusive voice template. Each template is composed of voiceprint feature vector cluster centers extracted from a large number of historical bullying incident related speech, and is accompanied by behavioral labels and confidence weights. The real-time timbre matching and bullying risk assessment module is used to perform similarity matching calculations between the real-time extracted voiceprint feature vector of the test speech and various specific timbre templates in the template library. Based on the matching results and preset thresholds, it generates the probability that the current speech segment belongs to various bullying-related timbres and comprehensively calculates the bullying risk level of the segment. The multimodal information fusion and behavior confirmation module is used to trigger and fuse auxiliary information from other sensors when the voice matching indicates a high risk. The auxiliary information includes, but is not limited to: infrared thermal imaging information of the same area, vibration sensor data, and preset emergency button trigger signals, to cross-verify bullying behavior and reduce the false alarm rate. The real-time early warning and evidence retention module is used to immediately send tiered early warning information to the terminals of the campus security center, relevant class teachers and administrators when a high-probability bullying incident is confirmed. At the same time, it automatically saves the complete audio stream, matching results and timestamps for a period of time before and after the early warning is triggered, forming a structured chain of evidence. The self-learning and template adaptive update module is used to incrementally learn and dynamically optimize a specific timbre template library based on confirmed bullying events (including true positives and false positives) and their corresponding voice data. It adjusts the template feature vector and confidence weights so that the system can adapt to timbre changes in different campus environments and student groups.

[0011] Preferably, the deep neural network model used in the specific timbre voiceprint feature extraction engine is an improved architecture that integrates a time-delay neural network and an attention mechanism; this model takes the pre-processed Mel spectrogram as input and processes it sequentially through: a) A group of convolutional layers is used to extract local time-frequency pattern features of speech signals; b) A time-delay neural network layer is used to capture long-term contextual dependencies in speech signals; c) Multi-head self-attention layer, used to focus on the keyframe region that contributes the most to timbre differentiation; d) Statistical pooling layer, used to aggregate variable-length frame-level feature sequences into a fixed-dimensional global feature vector; e) The fully connected layer and the feature normalization layer ultimately output a unit-length, highly discriminative timbre-specific voiceprint feature vector.

[0012] Preferably, in the real-time timbre matching and bullying risk assessment module, the feature vector to be tested is calculated. With the i-th type of timbre template in the template library (by K cluster centers { , ,..., } represents the similarity At that time, a matching algorithm based on weighted minimum-maximum distance is adopted, and its calculation formula is as follows: ; in, The speaker feature vector of the speech segment to be tested; Let K represent the feature vector of the k-th cluster center in the i-th timbre template, where K is the total number of cluster centers in this type of template; For vectors and The cosine distance between them; For scaling parameters (normal values); For the k-th cluster center Confidence weights; The probability that the speech segment belongs to the i-th type of bullying-related timbre is a natural exponential function. Depend on It is obtained by normalization using the Softmax function.

[0013] Preferably, when the module comprehensively calculates the bullying risk level R of the speech segment, it does not simply take the highest probability, but considers the joint occurrence pattern of multiple timbre features. The calculation formula is as follows: ; in, M represents the highest matching probability among all bullying-related voice categories; M is the set of voice categories with strong negative associations (such as threats and crying occurring simultaneously). B is the sum of the probabilities of these associated categories, used to enhance the identification of complex bullying behaviors; B is the set of timbre categories with confusing or anti-associated characteristics (such as normal conversation, laughter). The sum of these class probabilities is used to suppress false alarms caused by normal noise. , , This is an adjustment factor used to balance the contribution of various factors to the final risk level.

[0014] Preferably, the construction process of the bullying-related specific timbre template library construction and management module includes two stages: offline training and online initialization. a) Offline training phase: Collect a large number of labeled campus scene speech databases. The database should contain clear audio samples and variations of cries for help, crying, threats, and mockery; use the feature extraction engine to extract the voiceprint features of all samples; for the feature vector set of each timbre, use a clustering algorithm (such as DBSCAN or K-means++) to generate multiple cluster centers to cover the diversity of this timbre under different speakers and different emotional intensities; each cluster center is initialized as a template prototype and assigned an initial confidence weight; b) Online initialization phase: In the newly deployed campus environment, the system runs in low-sensitivity mode during the initial monitoring period (e.g., one week), mainly recording environmental background noise and common student voices; through semi-supervised learning, the feature vectors that frequently appear and have a very low matching degree with the offline template library are clustered to form environmentally specific background noise templates for the campus, and these templates are added to the template library for background interference filtering in subsequent matching calculations.

[0015] Preferably, the fusion decision logic of the multimodal information fusion and behavior confirmation module adopts an improved method based on DS evidence theory: the probability of various timbres output by the timbre matching module is... As one source of evidence E1; the detection results of "abnormal gathering of multiple people" from infrared thermal imaging, the detection results of "violent running or pushing" from vibration sensors, and the emergency button signal are used as other independent sources of evidence E2, E3, ...; a basic probability value is assigned to each source of evidence; all evidence is synthesized using Dempster's combination rule to calculate the confidence interval between the propositions "bullying occurred" and "bullying did not occur"; only when the confidence of the proposition "bullying occurred" exceeds the preset decision threshold, and the difference between its confidence and the confidence of "bullying did not occur" is sufficiently large, is the bullying event finally confirmed and an early warning is triggered.

[0016] Preferably, the tiered early warning strategy executed by the real-time early warning and evidence retention module is as follows: Level 1 Warning (Low Risk Alert): When the bullying risk level R exceeds the first threshold but multimodal confirmation is not triggered, it will only be recorded in the system background log and a silent text alert will be sent to the portable devices of security guards patrolling the relevant area, suggesting that patrols be strengthened. Level 2 Warning (Medium Risk Alert): When the risk level R exceeds the higher second threshold, or when it is preliminarily confirmed by multimodal information, a real-time warning window will pop up on the campus security center monitoring screen, displaying the location of the incident, the risk type (such as "suspected verbal threat"), and automatically retrieving available video footage from the vicinity of the area (if any). Level 3 Early Warning (High-Risk Emergency Response): When the risk level R is extremely high and the multimodal information fusion is highly confirming (such as simultaneously detecting distress sounds, heat sources with multiple people gathering, and violent vibrations), an audible and visual alarm will be immediately sent to the terminals of the security center, the school leader on duty, and the class teacher of the class involved. The on-site emergency broadcast will be automatically activated for voice intervention (such as playing a warning sound), and the complete evidence package will be uploaded to the cloud management platform in real time.

[0017] Preferably, the workflow of the system self-learning and template adaptive update module includes: 1. Feedback Collection: Collect the results of manual review after the warning (confirmed as real bullying, false alarm, or inconclusive) and the corresponding original audio data; 2. Sample screening and labeling: For samples confirmed as genuine bullying, their voiceprint feature vectors are added to the positive sample pool of the corresponding timbre category; for samples confirmed as false alarms, the main timbre features that caused the false alarms are analyzed and added to the "easily confused negative sample pool". 3. Incremental Template Update: Periodically use new data from the positive sample pool to fine-tune the cluster centers of the existing template or add new cluster centers, and recalculate the confidence weights based on the confirmation of the new samples. ; 4. Model fine-tuning: After accumulating enough new samples, the deep neural network model in the feature extraction engine is fine-tuned using incremental learning techniques to improve its ability to extract the timbre features of the current environment.

[0018] Preferably, the acquisition device deployed in the multi-channel speech signal acquisition and preprocessing module is a linear microphone array with directional pickup function. This array uses beamforming technology to physically enhance the speech signal from a specific area of ​​interest (such as inside a toilet cubicle or at a stairwell corner) while suppressing background noise and reverberation from other directions. This improves the clarity and intelligibility of the speech signal in complex acoustic environments, laying the foundation for subsequent high-precision voiceprint feature extraction.

[0019] By employing the above technical solution, this invention provides a specific timbre template voiceprint feature extraction system for voice anti-bullying. It possesses at least the following beneficial effects: This invention addresses the core issues of generalized detection targets, insufficient targeting, and weak distinguishability of voiceprint features by constructing a bullying-related specific timbre template library and designing a dedicated specific timbre voiceprint feature extraction engine. It achieves a leap from generalized abnormal sounds to precise bullying-related timbre detection, solving these problems. Furthermore, by employing a weighted minimum-maximum distance matching algorithm and a multi-timbre joint risk assessment model, it upgrades the assessment from single-probability judgment to multi-dimensional, context-related comprehensive risk assessment, effectively solving the misjudgment problem caused by a single decision-making basis and enhancing the ability to identify complex bullying scenarios. Finally, through multimodal information fusion and behavior confirmation modules and a tiered early warning strategy, it constructs an audio triggering mechanism. The collaborative workflow of multi-source verification and hierarchical response overcomes the challenges of high false alarm rates and crude early warning mechanisms. By deploying a microphone array with directional sound pickup capabilities and a corresponding front-end preprocessing algorithm, and introducing a system self-learning and template adaptive update module, the robustness and adaptability of the system in complex real-world environments are significantly improved, solving the problems of severe environmental interference, poor robustness, system rigidity, and lack of adaptability. By achieving real-time early warning and structured evidence retention, not only is immediate intervention achieved during the incident, but a complete monitoring-early warning-evidence collection closed loop is also constructed, providing objective and powerful technical evidence for the post-incident handling of school bullying and solving the pain points of difficulty in detection and evidence collection by traditional methods. Attached Figure Description

[0020] The accompanying drawings, which are provided to further illustrate the invention, constitute a part of this application: Figure 1 This is a schematic diagram of the structure of the specific timbre template voiceprint feature extraction system for anti-voice bullying according to the present invention. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] See Figure 1 This invention provides a specific timbre template voiceprint feature extraction system for voice anti-bullying, comprising: The multi-channel speech signal acquisition and preprocessing module 101 is used to acquire multiple environmental audio signals in real time, and to perform noise reduction, echo cancellation, speech enhancement and frame windowing preprocessing on the acquired raw audio signals to obtain a clean speech frame sequence. It should be noted that this module acts as the "front-end ear" of the system's perception, and its performance directly affects the accuracy of all subsequent analyses. In blind spots such as campus restrooms and stairwell corners, environmental noise (flushing sounds, echoes, distant noises) is complex. The aforementioned "multi-channel" acquisition allows the system to utilize spatial information from the microphone array. Noise reduction typically employs spectral subtraction or deep learning-based models to suppress steady-state and non-steady-state noise; echo cancellation is crucial for eliminating feedback from the device's own playback (such as broadcasts); speech enhancement aims to improve the signal-to-noise ratio and clarity of the speech signal. Framing and windowing is a standard procedure for converting continuous audio signals into a frame sequence suitable for short-time analysis, typically with a frame length of 20-30ms, a frame shift of 10-15ms, and a Hamming window to reduce spectral leakage. The output of this module is a high-quality, short-time speech frame prepared for subsequent processing, laying the foundation for high-precision voiceprint feature extraction.

[0023] The speech activity detection and key segment segmentation module 102 is used to perform endpoint detection on the preprocessed speech frame sequence, distinguish speech segments from non-speech segments, and further segment key speech segments containing suspected bullying semantics or strong emotional features from continuous speech segments based on energy, zero-crossing rate and spectral entropy features. It should be noted that this module is responsible for "preliminary screening," aiming to accurately locate potentially effective segments requiring in-depth analysis from the continuous audio stream. Voice Activity Detection (VAD) effectively filters out silent and purely noisy segments using a dual-threshold method (combining short-time energy and zero-crossing rate) or machine learning-based methods, saving computational resources. Key segment segmentation is the core innovation of this module. Its goal is not to segment all speech, but to focus on segments that "may contain bullying-related features." Energy features can capture sudden screams or shouts; the zero-crossing rate reflects, to some extent, the voiced / unvoiced characteristics of speech, which may change due to abnormal crying or abrupt threats; spectral entropy measures the flatness of the spectrum, with pure vowels or stable crying sounds exhibiting lower spectral entropy, while noisy or complex speech has higher spectral entropy. By calculating the joint criteria of these features in real time, the system can keenly capture short segments (typically 2-5 seconds) that exhibit "abnormal" or "strong emotional" fluctuations in energy, pitch, or timbre, thus concentrating valuable subsequent computational resources on high-value analysis targets.

[0024] The specific timbre voiceprint feature extraction engine 103 is used to extract high-dimensional voiceprint feature vectors that can characterize the speaker's timbre from the key speech segments; the engine has a built-in deep neural network model that maps time-frequency domain speech features to a highly discriminative voiceprint feature space through multi-layer nonlinear transformation. It should be noted that this engine is the system's "feature processing center," and its goal is to extract a new, specialized voiceprint feature. Unlike traditional speaker recognition, which focuses on "who this person is," this feature focuses on "the current timbre state of this person's speech," that is, it focuses on timbre change patterns driven by emotions and intentions (bullying / victimization). The deep neural network model takes Mel spectrograms (simulating human auditory characteristics) of key segments as input. Convolutional layers capture patterns such as local formants and harmonic structures in the spectrum; the temporal delay neural network (TDNN) layer spans a long time window to model the temporal evolution of sound, which is crucial for recognizing patterns such as prolonged intonation and trembling caused by emotional fluctuations; the multi-head self-attention layer enables the model to autonomously focus on the frequency bands and time regions with the highest discriminative power for the current timbre state (for example, attention may be focused on the high-frequency region representing vocal cord tension caused by fear). Finally, after statistical pooling and normalization, a fixed-dimensional, unit-length deep timbre embedding vector is output. This vector is weakly correlated with speaker identity, but strongly correlated with emotional / intentional timbre states such as "threat," "crying," and "calling for help."

[0025] The bullying-related specific timbre template library construction and management module 104 is used to store and manage specific timbre voiceprint feature templates that are highly related to school bullying behavior; the template library includes at least: a call for help timbre template, a crying timbre template, a threatening and intimidating timbre template, and a mocking and abusive timbre template; each template is composed of voiceprint feature vector cluster centers extracted from a large number of historical bullying incident related speech, and is attached with behavioral labels and confidence weights; It should be noted that this module is the system's "knowledge base" or "memory core," and its quality determines the accuracy of pattern matching. The construction process consists of two steps: First, in the offline phase, a well-labeled, large-scale campus voice database needs to be established, containing clean recordings of various bullying-related scenarios. All samples are processed using the aforementioned feature extraction engine to generate a high-dimensional feature vector set for each type of timbre (e.g., "crying"). Since the same type of timbre varies among different people and at different intensities, a single center cannot represent it. Therefore, a clustering algorithm (e.g., DBSCAN) is used to discover multiple "sub-patterns" or "prototypes" for this type of timbre, with each cluster center becoming a template prototype. Each prototype is assigned an initial confidence weight. Second, in the online initialization phase, the system runs in learning mode in the new environment, automatically learning common local background sounds and speaking habits to form an "environment-specific background sound template," which is added to the database as a special "counterexample" category to suppress false responses to regular environmental sounds during matching.

[0026] The real-time timbre matching and bullying risk assessment module 105 is used to perform similarity matching calculations between the real-time extracted voiceprint feature vector of the speech to be tested and various specific timbre templates in the template library. Based on the matching results and preset thresholds, it generates the probability that the current speech segment belongs to various bullying-related timbres and comprehensively calculates the bullying risk level of the segment. It should be noted that this module is the system's "real-time analytical brain." Its input is the deep timbre embedding vector extracted from the segment to be tested. The matching algorithm does not simply calculate the distance to a single center of each template class, but instead employs a weighted minimum-maximum distance strategy: for the i-th timbre template (containing K cluster centers { , ,..., }),calculate The similarity score is calculated by taking the distances to all centers of that class and selecting the one with the highest similarity (i.e., the max operation). From the formula: Provided.

[0027] here, This refers to the confidence weight of the cluster center, dynamically updated based on historical confirmation data, making repeatedly validated prototypes more influential in decision-making. The Softmax function is used to combine the confidence weights of each cluster center. Normalization to probability Then, the risk assessment model is activated, and its calculation formula is as follows: This reflects advanced analytical logic: it goes beyond just looking at the highest probability. It also positively weights frequently co-occurring negative timbre combinations (M set, such as "threat" + "crying") and negatively suppresses normal or anti-correlated timbres that often lead to false alarms (B set, such as "laughing"). Coefficients , , Adjust the contributions of each part.

[0028] The multimodal information fusion and behavior confirmation module 106 is used to trigger and fuse auxiliary information from other sensors when the voice matching prompts a high risk. The auxiliary information includes, but is not limited to: infrared thermal imaging information of the same area, vibration sensor data, and preset emergency button trigger signals, to cross-verify bullying behavior and reduce false alarm rate. It should be noted that this module serves as the system's "cross-validation defense," designed to address the high false alarm rate of a single audio sensor. This module is activated when the bullying risk level R derived from audio analysis exceeds a certain threshold. It utilizes infrared thermal imaging sensors deployed in the same physical area (to determine if there is a large, close gathering of more people than normal), vibration sensors (to detect abnormal physical vibrations such as running, pushing, or impacts), and emergency button signals (the highest-priority manual trigger). Information fusion is performed using Dempster's evidence theory: the audio analysis results (derived from various timbre probabilities and risk levels) are used as evidence source E1, while the infrared, vibration, and button detection results are used as independent evidence sources E2, E3, ... . Basic probability assignments (BPAs) are assigned to the propositions "bullying occurred" and "bullying did not occur" for each evidence source, and then all evidence is synthesized using Dempster's combination rules. Finally, the system calculates the overall reliability of the "bullying occurred" proposition. Only when this reliability exceeds a preset decision threshold and differs sufficiently from the reliability of "bullying did not occur" is a bullying event definitively confirmed. This mathematically based fusion method is more scientific than simple logical AND / OR, effectively characterizing the uncertainty of evidence and significantly improving the reliability of decision-making.

[0029] The real-time early warning and evidence retention module 107 is used to immediately send graded early warning information to the terminals of the campus security center, relevant class teachers and management personnel when a high-probability bullying incident is confirmed. At the same time, it automatically saves the complete audio stream, matching results and timestamps for a period of time before and after the early warning is triggered, forming a structured chain of evidence. It should be noted that this module is the system's "execution and recording terminal," realizing a closed loop from analysis to action. The early warning system employs a tiered strategy to balance response efficiency and avoid alarm fatigue: Level 1 (low risk) only records in the background and provides a silent alert, used for suspicious but unconfirmed situations; Level 2 (medium risk) triggers an alarm in the monitoring center, suitable for higher risks or when there is preliminary multimodal evidence; Level 3 (high risk) triggers audible and visual alarms and multi-party notifications, suitable for emergency events with high multimodal confirmation. The evidence preservation function is equally crucial. The system automatically archives the original audio, feature extraction results, matching probabilities, risk levels, multimodal fusion data, and precise timestamps from tens of seconds before the warning trigger to a period afterward (e.g., 30 seconds before and 60 seconds after), packaging them into an immutable, structured evidence file. This file can not only be used for post-event review, but more importantly, it provides objective and powerful technical evidence for the investigation, identification, and handling of school bullying incidents, solving the problem of "no evidence beyond verbal testimony" in traditional school disputes.

[0030] The self-learning and template adaptive update module 108 is used to incrementally learn and dynamically optimize a specific timbre template library based on confirmed bullying events (including true positives and false alarms) and their corresponding voice data, and adjust the template feature vector and confidence weight so that the system can adapt to the timbre changes of different campus environments and student groups. It's important to note that this module endows the system with the ability to "continuously evolve," which is key to its long-term high efficiency. After the system is running, each alert is manually reviewed (to confirm whether it's a genuine bullying incident, a false alarm, or an inconclusive case). This feedback, along with the corresponding original audio, is collected. For genuine bullying samples, their feature vectors are added to the positive sample pool of the corresponding timbre category for subsequent incremental template updates—this allows for fine-tuning the positions of existing cluster centers or adding new cluster centers when new "sub-patterns" appear in the feature space. Simultaneously, the confidence weight of the center is dynamically increased based on the frequency with which samples corresponding to that center are confirmed. For false positive samples, their characteristics are analyzed, and they are added to the "easily confused negative sample pool" to better define the anti-association set B in the risk assessment formula, or to suppress similar patterns in future matching. After accumulating a sufficient number of new samples, incremental learning techniques can be used to fine-tune the deep neural network model of the feature extraction engine, making its extracted features more adaptable to the specific acoustic environment of the school and the timbre characteristics of the student group, achieving increasingly accurate personalized adaptation over time.

[0031] The deep neural network model used in the specific timbre voiceprint feature extraction engine is an improved architecture that integrates a time-delay neural network and an attention mechanism; this model takes the pre-processed Mel spectrogram as input and processes it sequentially as follows: a) A group of convolutional layers is used to extract local time-frequency pattern features of speech signals; b) A time-delay neural network layer is used to capture long-term contextual dependencies in speech signals; c) Multi-head self-attention layer, used to focus on the keyframe region that contributes the most to timbre differentiation; d) Statistical pooling layer, used to aggregate variable-length frame-level feature sequences into a fixed-dimensional global feature vector; e) The fully connected layer and the feature normalization layer ultimately output a unit-length, highly discriminative timbre-specific voiceprint feature vector; It's important to note that this model architecture is specifically tailored for the "timbre state recognition" task. Convolutional layers extract fundamental spectral local features, such as formant contours. The Temporal Delay Neural Network (TDNN) layer, through its unique sparse connections, captures long-range contextual dependencies (e.g., hundreds of milliseconds) between speech frames with a small number of parameters. This is crucial for perceiving slow changes in tone or periodic tremolos caused by emotion. The multi-head self-attention layer is a core innovation, allowing the model to autonomously weigh the contribution weight of each speech frame to the final timbre judgment in different sub-feature spaces (i.e., different "heads"). For example, one "head" might focus more on the mid-to-high frequency regions containing strong emotional information, while another "head" might focus more on the harmonic structure stability of the speech. By weighted aggregation of this information focused on different aspects, the model can more robustly and comprehensively represent complex timbre states. Statistical pooling (such as mean and standard deviation) converts variable-length sequences into fixed-length vectors, fully connected layers perform high-dimensional nonlinear combinations, and finally feature normalization (such as length normalization) ensures that the output vector lies on the hypersphere, which facilitates subsequent cosine distance calculation.

[0032] In the real-time timbre matching and bullying risk assessment module, the feature vector to be tested is calculated. With the i-th type of timbre template in the template library (by K cluster centers { , ,..., } represents the similarity At that time, a matching algorithm based on weighted minimum-maximum distance is adopted, and its calculation formula is as follows: ; in, The speaker feature vector of the speech segment to be tested; Let K represent the feature vector of the k-th cluster center in the i-th timbre template, where K is the total number of cluster centers in this type of template; For vectors and The cosine distance between them; For scaling parameters (normal values); For the k-th cluster center Confidence weights; The probability that the speech segment belongs to the i-th type of bullying-related timbre is a natural exponential function. Depend on Obtained by normalization using the Softmax function; It should be noted that this matching algorithm is ingeniously designed, balancing computational efficiency and matching accuracy. The operation means that for each type of timbre, only the distance between the vector to be tested and the cluster prototype most similar to it is considered. This is intuitive: as long as the timbre to be tested is highly similar to a typical sub-pattern of that type of timbre, it can be considered a match. Map distance to similarity Controlling the decay rate. The key lies in introducing weights. In the initial stage of system operation, all weights can be initialized to be equal. As the system learns, the weights of cluster centers that have been repeatedly manually identified as genuine bullying in historical alerts will be adjusted. The center weights associated with false positives will be increased; conversely, the center weights associated with false positives may be decreased. This makes the matching results not only dependent on geometric distance but also incorporate the "credibility" judgment based on historical experience, making the system's decision-making more intelligent and reliable. Finally, Softmax normalization converts the similarity scores of each category into probabilities, providing standardized input for subsequent risk assessment.

[0033] When calculating the bullying risk level R of the speech segment, the module does not simply take the highest probability, but considers the joint occurrence pattern of multiple timbre features. The calculation formula is as follows: ; in, M represents the highest matching probability among all bullying-related voice categories; M is the set of voice categories with strong negative associations (such as threats and crying occurring simultaneously). B is the sum of the probabilities of these associated categories, used to enhance the identification of complex bullying behaviors; B is the set of timbre categories with confusing or anti-associated characteristics (such as normal conversation, laughter). The sum of these class probabilities is used to suppress false alarms caused by normal noise. , , This is an adjustment factor used to balance the contribution of each factor to the final risk level; It should be noted that this risk assessment formula is one of the core logics behind the high accuracy and low false alarm rate of this invention. It simulates the human thought process of comprehensive judgment: This item represents the strength of the "dominant evidence." Any bullying tone (such as a threat) with a sufficiently high probability inherently carries a basic risk.

[0034] This term introduces an enhancement effect of "evidence association." Set M defines timbre categories that frequently co-occur in real-world bullying scenarios; for example, the simultaneous occurrence of "threat" and "crying" significantly increases the likelihood of bullying compared to either timbre appearing alone. This term weights and enhances these co-occurrence patterns, enabling the system to effectively identify complex bullying behaviors.

[0035] The item implements an "evidence exclusion" suppression mechanism. Set B defines tone categories that are common in normal or friendly interactions but are generally incompatible with bullying scenarios or even reduce the likelihood of bullying, such as "normal conversation" or "happy laughter." When the probability of these categories is high, the item will actively reduce the overall risk level R, thereby effectively suppressing false alarms caused by normal noise such as heated debates or cheering at sports.

[0036] Adjustment coefficient , , The specific definitions of sets M and B need to be calibrated based on domain knowledge (such as campus psychology and security expert experience) and experimental data before system deployment. Together, they constitute a set of adjustable "judgment rules" that enable the system to flexibly adapt to the needs of different campus cultures and security strategies.

[0037] The construction process of the bullying-related specific timbre template library construction and management module includes two stages: offline training and online initialization. a) Offline training phase: Collect a large number of labeled campus scene speech databases. The database should contain clear audio samples and variations of cries for help, crying, threats, and mockery; use the feature extraction engine to extract the voiceprint features of all samples; for the feature vector set of each timbre, use a clustering algorithm (such as DBSCAN or K-means++) to generate multiple cluster centers to cover the diversity of this timbre under different speakers and different emotional intensities; each cluster center is initialized as a template prototype and assigned an initial confidence weight; b) Online initialization phase: In the newly deployed campus environment, the system runs in low-sensitivity mode during the initial monitoring period (e.g., one week), mainly recording environmental background noise and common student voices; through semi-supervised learning, the feature vectors that frequently appear and have a very low matching degree with the offline template library are clustered to form a background noise template specific to the campus environment, and added to the template library for background interference filtering in subsequent matching calculations. It's important to note that this two-stage construction strategy balances universality and specificity, which is crucial for ensuring the system's practicality. The offline training phase utilizes a large-scale, well-annotated general campus speech database to establish a "seed template library" with good generalization ability, covering various basic variations in bullying-related timbre. This library forms the foundation for system startup and operation. The online initialization phase involves personalized adaptation for each specific deployment environment. During the learning period after installation in a new school, the system operates in a "record only, no alarm" or low-sensitivity mode. Its primary purpose is not to detect bullying, but to learn local "sound fingerprints": including the spectral characteristics of background noise, the common fundamental frequency range of students' speech at the school, dialect characteristics, and common non-bullying high-energy sounds (such as specific school bells or sports activity slogans). Through semi-supervised clustering, these frequently occurring features that differ significantly from the offline templates are aggregated into "environment-specific background sound templates" and added to the template library. In subsequent real-time matching, if the current audio highly matches these local background templates, its risk is significantly suppressed, thereby greatly improving the system's initial adaptability and resistance to local interference in new environments.

[0038] The fusion decision logic of the multimodal information fusion and behavior confirmation module adopts an improved method based on DS evidence theory: the timbre probabilities output by the timbre matching module are... As one source of evidence E1; the detection results of "abnormal gathering of multiple people" from infrared thermal imaging, the detection results of "violent running or pushing" from vibration sensors, and the emergency button signal are used as other independent sources of evidence E2, E3, ...; a basic probability value is assigned to each source of evidence; all evidence is synthesized using Dempster's combination rule to calculate the confidence interval between the propositions "bullying occurred" and "bullying did not occur"; only when the confidence of the proposition "bullying occurred" exceeds the preset decision threshold, and the difference between its confidence and the confidence of "bullying did not occur" is sufficiently large, is the bullying event finally confirmed and an early warning is triggered; It is worth noting that using Dempster's evidence theory for multimodal fusion has a natural advantage in handling uncertain and conflicting information compared to simple threshold logic "AND / OR". In the complex environment of a campus, each sensor may provide incomplete or uncertain evidence. For example, an infrared sensor may be unable to determine the exact number of people due to occlusion, and a vibration sensor may be unable to distinguish between running and falling objects. Dempster's theory allows for the assignment of confidence (basic probability value) to each source of evidence for "bullying", "not bullying", and "uncertain". When multiple sources of evidence simultaneously support "bullying", the confidence of the proposition will be significantly enhanced through synthesis using Dempster's rules; when there is conflict between the evidence (such as an audio warning of high risk but an infrared display showing no one), the synthesized result will exhibit high "uncertainty", which may suppress alarms or trigger a more cautious response (such as a level 1 warning). This mathematically based fusion framework makes the system's decision-making process more rigorous and transparent, and provides a clear mathematical model for subsequent optimization (such as adjusting BPA allocation), which is the core of achieving high-reliability confirmation.

[0039] The tiered early warning strategy implemented by the real-time early warning and evidence retention module is as follows: Level 1 Warning (Low Risk Alert): When the bullying risk level R exceeds the first threshold but multimodal confirmation is not triggered, it will only be recorded in the system background log and a silent text alert will be sent to the portable devices of security guards patrolling the relevant area, suggesting that patrols be strengthened. Level 2 Warning (Medium Risk Alert): When the risk level R exceeds the higher second threshold, or when it is preliminarily confirmed by multimodal information, a real-time warning window will pop up on the campus security center monitoring screen, displaying the location of the incident, the risk type (such as "suspected verbal threat"), and automatically retrieving available video footage from the vicinity of the area (if any). Level 3 Early Warning (High-Risk Emergency Response): When the risk level R is extremely high and the multimodal information fusion is highly confirmed (such as the simultaneous detection of distress call timbre, heat source of multiple people gathering, and violent vibration), an audible and visual alarm will be immediately sent to the terminals of the security center, the school leader on duty, and the class teacher of the class involved, and the on-site emergency broadcast will be automatically activated for voice intervention (such as playing a warning sound). At the same time, the complete evidence package will be uploaded to the cloud management platform in real time. It is worth noting that this tiered early warning strategy profoundly embodies the principles of "precise response" and "minimal interference." It does not escalate all suspicious events to the highest alert level, but rather takes differentiated measures based on the level of risk confidence. Level 1 alerts are equivalent to "increased attention," raising the local alert level without alarming the majority of people; this is suitable for early signs or unclear situations. Level 2 alerts initiate the formal response process, visualizing information for professional security personnel to facilitate rapid judgment and decision-making. Level 3 alerts are equivalent to activating the emergency response plan, involving multiple parties to immediately prevent potential harm. This tiered mechanism has two major advantages: first, it greatly reduces the disruption to normal teaching order caused by unavoidable false alarms, avoiding the "boy who cried wolf" effect and maintaining the system's credibility; second, it ensures that in truly critical moments, response resources can be mobilized most quickly and effectively, forming a complete emergency chain of "suspicion-attention-confirmation-emergency response," improving the refinement and intelligence of campus security management.

[0040] The workflow of the system's self-learning and template adaptive update module includes: 1. Feedback Collection: Collect the results of manual review after the warning (confirmed as real bullying, false alarm, or inconclusive) and the corresponding original audio data; 2. Sample screening and labeling: For samples confirmed as genuine bullying, their voiceprint feature vectors are added to the positive sample pool of the corresponding timbre category; for samples confirmed as false alarms, the main timbre features that caused the false alarms are analyzed and added to the "easily confused negative sample pool". 3. Incremental Template Update: Periodically use new data from the positive sample pool to fine-tune the cluster centers of the existing template or add new cluster centers, and recalculate the confidence weights based on the confirmation of the new samples. ; 4. Model fine-tuning: After accumulating enough new samples, the deep neural network model in the feature extraction engine is fine-tuned using incremental learning techniques to improve its ability to extract the timbre features of the current environment. It is worth noting that this module achieves closed-loop feedback and continuous optimization, ensuring the system's long-term viability. Its core idea is to iteratively optimize itself using production data (i.e., the warnings and verification results generated during actual system operation). This solves the common "data distribution drift" problem in machine learning models—that is, the distribution of training data (offline library) differs from the data distribution in the real production environment. By continuously feeding back confirmed positive samples (genuine bullying) and negative samples (false alarms) to the system, the template library is continuously refined, better covering the real timbre variations in the current environment; the feature extraction model, through fine-tuning, can better extract discriminative features from the speech in the current environment. More importantly, the confidence weight... The dynamic updates transform the template library from a static set of prototypes into an "intelligent agent" carrying historical judgment experience, whose decision-making authority (weight) increases with the number of times it is verified. This design allows the system to truly integrate into the campus, and over time, its detection performance will increasingly align with the actual needs of the campus, achieving personalized performance improvements.

[0041] The acquisition device deployed in the multi-channel speech signal acquisition and preprocessing module is a linear microphone array with directional sound pickup function. This array uses beamforming technology to physically enhance the speech signal from a specific area of ​​interest (such as inside a toilet cubicle or at a stair corner) while suppressing background noise and reverberation from other directions. This improves the clarity and intelligibility of the speech signal in complex acoustic environments, laying the foundation for subsequent high-precision voiceprint feature extraction. It is worth noting that in harsh acoustic environments with high noise and strong reverberation, such as blind spots in campus surveillance, the quality of front-end sound pickup is the bottleneck determining the performance of the entire system. Employing linear microphone arrays and beamforming technology is an effective way to solve this problem at the physical hardware level. The array calculates the direction of sound wave arrival by analyzing the phase difference of signals received by multiple microphones. The beamforming algorithm synthesizes a highly directional "virtual microphone" in the digital domain by adjusting the weights and delays of the signals in each channel. Its main lobe (the direction with the strongest receiving capability) can be electronically scanned or fixedly aligned with the focal area of ​​the space to be monitored (such as a toilet stall or a stair landing). In this way, the speech signal from the focal area is enhanced by in-phase superposition, while noise and reverberation from other directions (usually from reflections from walls and ceilings) are suppressed to varying degrees. This is equivalent to installing an "auditory telescope" for the system in a noisy environment, directly improving the signal-to-noise ratio and speech quality of the input signal from the source. It provides the "cleanest" possible raw material for the complex back-end voiceprint feature extraction and pattern matching algorithms, serving as the first and most important guarantee for the stable operation of the entire system in real-world complex environments.

[0042] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A specific timbre template voiceprint feature extraction system for voice anti-bullying, characterized in that, The system includes: The multi-channel speech signal acquisition and preprocessing module is used to acquire multiple environmental audio signals in real time, and to perform noise reduction, echo cancellation, speech enhancement and frame windowing preprocessing on the acquired raw audio signals to obtain a clean speech frame sequence. The speech activity detection and key segmentation module is used to perform endpoint detection on the preprocessed speech frame sequence, distinguish speech segments from non-speech segments, and further segment key speech segments containing suspected bullying semantics or strong emotional features from continuous speech segments based on energy, zero-crossing rate and spectral entropy features. A specific timbre voiceprint feature extraction engine is used to extract high-dimensional voiceprint feature vectors that can characterize the speaker's timbre from the key speech segments. The engine has a built-in deep neural network model that maps time-frequency domain speech features to a highly discriminative voiceprint feature space through multi-layer nonlinear transformation. A module for constructing and managing a bullying-related specific voiceprint template library is used to store and manage specific voiceprint feature templates that are highly related to school bullying behavior. The template library includes at least: a call for help voice template, a crying voice template, a threatening and intimidating voice template, and a mocking and abusive voice template. Each template is composed of voiceprint feature vector cluster centers extracted from a large number of historical bullying incident related speech, and is accompanied by behavioral labels and confidence weights. The real-time timbre matching and bullying risk assessment module is used to perform similarity matching calculations between the real-time extracted voiceprint feature vector of the test speech and various specific timbre templates in the template library. Based on the matching results and preset thresholds, it generates the probability that the current speech segment belongs to various bullying-related timbres and comprehensively calculates the bullying risk level of the segment. The multimodal information fusion and behavior confirmation module is used to trigger and fuse auxiliary information from other sensors when the voice matching indicates a high risk. The auxiliary information includes, but is not limited to: infrared thermal imaging information of the same area, vibration sensor data, and preset emergency button trigger signals, to cross-verify bullying behavior and reduce the false alarm rate. The real-time early warning and evidence retention module is used to immediately send tiered early warning information to the terminals of the campus security center, relevant class teachers and administrators when a high-probability bullying incident is confirmed. At the same time, it automatically saves the complete audio stream, matching results and timestamps for a period of time before and after the early warning is triggered, forming a structured chain of evidence. The self-learning and template adaptive update module is used to incrementally learn and dynamically optimize a specific timbre template library based on confirmed bullying events (including true positives and false positives) and their corresponding voice data. It adjusts the template feature vector and confidence weights so that the system can adapt to timbre changes in different campus environments and student groups.

2. The specific timbre template voiceprint feature extraction system for voice anti-bullying according to claim 1, characterized in that, The deep neural network model used in the specific timbre voiceprint feature extraction engine is an improved architecture that integrates a time-delay neural network and an attention mechanism; this model takes the pre-processed Mel spectrogram as input and processes it sequentially as follows: a) A group of convolutional layers is used to extract local time-frequency pattern features of speech signals; b) A time-delay neural network layer is used to capture long-term contextual dependencies in speech signals; c) Multi-head self-attention layer, used to focus on the keyframe region that contributes the most to timbre differentiation; d) Statistical pooling layer, used to aggregate variable-length frame-level feature sequences into a fixed-dimensional global feature vector; e) The fully connected layer and the feature normalization layer ultimately output a unit-length, highly discriminative timbre-specific voiceprint feature vector.

3. The specific timbre template voiceprint feature extraction system for voice anti-bullying according to claim 1, characterized in that, In the real-time timbre matching and bullying risk assessment module, the feature vector to be tested is calculated. With the i-th type of timbre template in the template library (by K cluster centers { , ,..., } represents the similarity At that time, a matching algorithm based on weighted minimum-maximum distance is adopted, and its calculation formula is as follows: ; in, The speaker feature vector of the speech segment to be tested; Let K represent the feature vector of the k-th cluster center in the i-th timbre template, where K is the total number of cluster centers in this type of template; For vectors and The cosine distance between them; For scaling parameters (normal values); For the k-th cluster center Confidence weights; The probability that the speech segment belongs to the i-th type of bullying-related timbre is a natural exponential function. Depend on It is obtained by normalization using the Softmax function.

4. The specific timbre template voiceprint feature extraction system for voice anti-bullying according to claim 1, characterized in that, When calculating the bullying risk level R of the speech segment, the module does not simply take the highest probability, but considers the joint occurrence pattern of multiple timbre features. The calculation formula is as follows: ; in, M represents the highest matching probability among all bullying-related voice categories; M is the set of voice categories with strong negative associations (such as threats and crying occurring simultaneously). B is the sum of the probabilities of these associated categories, used to enhance the identification of complex bullying behaviors; B is the set of timbre categories with confusing or anti-associated characteristics (such as normal conversation, laughter). The sum of these class probabilities is used to suppress false alarms caused by normal noise. , , This is an adjustment factor used to balance the contribution of various factors to the final risk level.

5. The specific timbre template voiceprint feature extraction system for voice anti-bullying according to claim 1, characterized in that, The construction process of the bullying-related specific timbre template library construction and management module includes two stages: offline training and online initialization. a) Offline training phase: Collect a large number of labeled campus scene speech databases. The database should contain clear audio samples of calls for help, crying, threats, mockery, etc. and their variations; use the feature extraction engine to extract the voiceprint features of all samples; For each set of feature vectors for a timbre, a clustering algorithm (such as DBSCAN or K-means++) is used to generate multiple cluster centers to cover the diversity of that timbre under different speakers and different emotional intensities; each cluster center is initialized as a template prototype and assigned an initial confidence weight. b) Online initialization phase: In the newly deployed campus environment, the system runs in low-sensitivity mode during the initial monitoring period (e.g., one week), mainly recording environmental background noise and common student voices; through semi-supervised learning, the feature vectors that frequently appear and have a very low matching degree with the offline template library are clustered to form environmentally specific background noise templates for the campus, and these templates are added to the template library for background interference filtering in subsequent matching calculations.

6. The specific timbre template voiceprint feature extraction system for voice anti-bullying according to claim 1, characterized in that, The fusion decision logic of the multimodal information fusion and behavior confirmation module adopts an improved method based on DS evidence theory: the timbre probabilities output by the timbre matching module are... As one source of evidence E1; the detection results of "abnormal gathering of multiple people" from infrared thermal imaging, the detection results of "violent running or pushing" from vibration sensors, and the emergency button signal are used as other independent sources of evidence E2, E3, ...; basic probability values ​​are assigned to each source of evidence; all evidence is synthesized using Dempster's combination rule to calculate the confidence interval of the propositions "bullying occurred" and "bullying did not occur"; A bullying event is finally confirmed and an alert is triggered only when the reliability of the proposition "bullying occurred" exceeds a preset decision threshold and the difference between its reliability and the reliability of "bullying did not occur" is large enough.

7. The specific timbre template voiceprint feature extraction system for voice anti-bullying according to claim 1, characterized in that, The tiered early warning strategy implemented by the real-time early warning and evidence retention module is as follows: Level 1 Warning (Low Risk Alert): When the bullying risk level R exceeds the first threshold but multimodal confirmation is not triggered, it will only be recorded in the system background log and a silent text alert will be sent to the portable devices of security guards patrolling the relevant area, suggesting that patrols be strengthened. Level 2 Warning (Medium Risk Alert): When the risk level R exceeds the higher second threshold, or when it is preliminarily confirmed by multimodal information, a real-time warning window will pop up on the campus security center monitoring screen, displaying the location of the incident, the risk type (such as "suspected verbal threat"), and automatically retrieving available video footage from the vicinity of the area (if any). Level 3 Early Warning (High-Risk Emergency Response): When the risk level R is extremely high and the multimodal information fusion is highly confirming (such as simultaneously detecting distress sounds, heat sources with multiple people gathering, and violent vibrations), an audible and visual alarm will be immediately sent to the terminals of the security center, the school leader on duty, and the class teacher of the class involved. The on-site emergency broadcast will be automatically activated for voice intervention (such as playing a warning sound), and the complete evidence package will be uploaded to the cloud management platform in real time.

8. The specific timbre template voiceprint feature extraction system for voice anti-bullying according to claim 1, characterized in that, The workflow of the system's self-learning and template adaptive update module includes: 1) Feedback collection: Collect the results of manual review after the warning (confirmed as real bullying, false alarm, or cannot be determined) and the corresponding original audio data; 2) Sample screening and labeling: For samples confirmed as genuine bullying, their voiceprint feature vectors are added to the positive sample pool of the corresponding timbre category; for samples confirmed as false alarms, the main timbre features that caused the false alarms are analyzed and added to the "easily confused negative sample pool". 3) Incremental template update: Periodically use new data from the positive sample pool to fine-tune the cluster centers of the existing template or add new cluster centers, and recalculate the confidence weights based on the confirmation of the new samples. ; 4) Model fine-tuning: After accumulating enough new samples, the deep neural network model in the feature extraction engine is fine-tuned using incremental learning techniques to improve its ability to extract the timbre features of the current environment.

9. The specific timbre template voiceprint feature extraction system for voice anti-bullying according to claim 1, characterized in that, The acquisition device deployed in the multi-channel speech signal acquisition and preprocessing module is a linear microphone array with directional pickup function. This array uses beamforming technology to physically enhance the speech signal from a specific area of ​​interest (such as inside a toilet cubicle or at a stairwell corner) while suppressing background noise and reverberation from other directions. This improves the clarity and intelligibility of the speech signal in complex acoustic environments, laying the foundation for subsequent high-precision voiceprint feature extraction.