Anti-spoofing alarm method and system based on AI voice recognition

By deploying anti-bullying equipment on campus and utilizing AI voice recognition and voiceprint technology, bullying behavior can be monitored and located in real time. This solves the problems of real-time monitoring and coverage of traditional monitoring methods, enabling rapid and accurate detection and prevention of bullying behavior and ensuring student safety.

CN120932367AActive Publication Date: 2025-11-11SHANGHAI LIQING INTELLIGENT TECH CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202511443483.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2025-11-11
Estimated Expiration
2045-10-10

AI Technical Summary

Technical Problem

Existing methods for monitoring and preventing school bullying suffer from poor real-time performance, limited coverage, reliance on manual labor, and low efficiency, making it difficult to detect and stop bullying behavior in a timely manner.

Method used

By deploying anti-bullying devices in specific areas of the campus, AI voice recognition technology is used to collect sound data in real time, identify key information about bullying, generate alarm information, locate the target area, output stop-the-bullying information, and identify bullying suspects by combining facial recognition and voiceprint feature recognition.

Benefits of technology

It enables real-time and comprehensive monitoring of school bullying, allowing for rapid and accurate detection and timely intervention, improving the efficiency and accuracy of information transmission, ensuring student safety, and creating a healthy campus environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932367A_ABST
    Figure CN120932367A_ABST
Patent Text Reader

Abstract

The invention discloses an AI voice recognition-based anti-spoofing alarm method and system, and the method comprises the steps: deploying anti-spoofing equipment in at least one preset spoofing monitoring region in a campus, and obtaining to-be-recognized sound data collected by each piece of anti-spoofing equipment; performing voice recognition on the to-be-recognized sound data, and when any target sound data in the to-be-recognized sound data is recognized to contain preset spoofing key information, generating anti-spoofing alarm information according to the target sound data and the anti-spoofing equipment identifier corresponding to the target sound data; the anti-deception alarm information is reported to a management terminal, so that the management terminal outputs anti-deception prompt information based on a target deception monitoring area and target sound data corresponding to the anti-deception equipment identifier; and receiving anti-deception stopping information of the management terminal, and performing voice output of the anti-deception stopping information through target anti-deception equipment corresponding to the anti-deception equipment identifier.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voice analysis technology, and in particular to an anti-bullying alarm method and system based on AI voice recognition. Background Technology

[0002] School bullying is a serious problem facing the global education sector, severely harming students' physical and mental health and development. In recent years, with increasing societal attention to school safety, how to effectively prevent and promptly stop school bullying has become a focus of common concern for education administrators, parents, and all sectors of society.

[0003] Traditional methods for monitoring and preventing school bullying primarily rely on manual patrols and student reports. While manual patrols can detect some bullying behavior to a certain extent, they have significant limitations. Schools are vast, and bullying can occur in various hidden locations such as classrooms, corridors, and corners of the playground. Manual patrols cannot provide comprehensive and real-time coverage, causing many bullying incidents to go undetected in their early stages. Furthermore, manual patrols require substantial manpower and resources, resulting in low efficiency; they often fail to detect and stop sudden and highly concealed bullying incidents in a timely manner.

[0004] Students' initiative to report bullying incidents also faces many problems. On the one hand, some bullied students are afraid to report their experiences to teachers or parents due to fear, shame, or other reasons, fearing more severe retaliation. On the other hand, some bystanders may remain silent for fear of being implicated or feeling it's none of their business, failing to report any bullying they witness. This results in many bullying incidents going undetected for a long time after they occur, causing ongoing psychological and physical harm to the bullied students.

[0005] Furthermore, while existing campus surveillance systems can record visual information on campus, their ability to capture and analyze audio information is limited. Bullying behavior is often accompanied by specific verbal content, such as threats and insults. It is difficult to accurately determine whether bullying has occurred based solely on video surveillance, nor can it obtain crucial information about the bullying process in a timely manner. Moreover, the video data volume of surveillance systems is enormous, and manually reviewing and analyzing this data requires a significant amount of time and effort, making real-time monitoring and rapid response difficult.

[0006] In summary, existing methods for monitoring and preventing school bullying suffer from problems such as poor real-time performance, limited coverage, reliance on manual labor, and low efficiency, making it impossible to detect and stop school bullying in a timely and effective manner. Summary of the Invention

[0007] In view of this, the embodiments of this application provide an anti-bullying alarm method and system based on AI voice recognition, which realizes real-time and comprehensive monitoring of school bullying behavior, can quickly and accurately detect bullying behavior and stop it in time, overcomes the shortcomings of traditional monitoring and prevention methods such as poor real-time performance, limited coverage, reliance on manual labor and low efficiency, effectively protects students' campus safety and creates a healthy and harmonious campus environment.

[0008] According to one aspect of this application, an anti-bullying alarm method based on AI voice recognition is provided, the method comprising: Acquire the voice data to be identified by anti-bullying devices deployed in at least one pre-set bullying monitoring area on campus; The voice data to be identified is subjected to speech recognition, and when any target voice data in the voice data to be identified contains preset bullying key information, anti-bullying alarm information is generated based on the target voice data and the anti-bullying device identifier corresponding to the target voice data. The anti-bullying alarm information is reported to the management terminal, so that the management terminal outputs anti-bullying prompt information based on the target bullying monitoring area corresponding to the anti-bullying device identifier and the target sound data; Receive anti-bullying information from the management terminal, and output the anti-bullying information via voice through the target anti-bullying device corresponding to the anti-bullying device identifier; The method further includes: When any target sound data in the sound data to be identified contains preset bullying key information, the target bullying monitoring area corresponding to the anti-bullying device identifier is determined according to the anti-bullying device identifier corresponding to the target sound data, and the surrounding monitoring device corresponding to the target bullying monitoring area is determined, and the monitoring data captured by the surrounding monitoring device is obtained. The monitoring data is subjected to facial recognition. Based on the facial recognition results, the person to be investigated corresponding to the target voice data is determined, and the individual voiceprint features corresponding to the person to be investigated are obtained. Obtain the contextual audio data corresponding to the target sound data, and perform speaker voice separation on the contextual audio data to obtain multiple speaker voice data corresponding to the contextual audio data; For any speaker's voice data, a pre-trained voiceprint model is used to identify the speaker's voice features, and the bullying suspect is identified based on the individual voiceprint features that match the identified speaker's voiceprint features.

[0009] According to another aspect of this application, an anti-bullying alarm system based on AI voice recognition is provided. The system includes an anti-bullying device and a control device. The anti-bullying device is deployed in at least one pre-set bullying monitoring area on campus. The anti-bullying device and the control device are used to implement the above-mentioned anti-bullying alarm method based on AI voice recognition.

[0010] According to another aspect of this application, a storage medium is provided that stores a computer program thereon, which, when executed by a processor, implements the above-described anti-bullying alarm method based on AI voice recognition.

[0011] According to another aspect of this application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the program to implement the above-described anti-bullying alarm method based on AI voice recognition.

[0012] By employing the above technical solution, this application provides an anti-bullying alarm method and system based on AI voice recognition. This method collects sound data by deploying anti-bullying devices in a pre-defined area on campus, uses AI voice recognition technology to identify bullying behavior, promptly reports alarm information to a management terminal and locates the target area, and finally outputs a stop-the-bullying message through the target device. This series of operations achieves real-time and comprehensive monitoring of bullying behavior on campus, enabling rapid and accurate detection and timely intervention. It overcomes the shortcomings of traditional monitoring and prevention methods, such as poor real-time performance, limited coverage, reliance on manual labor, and low efficiency, effectively ensuring student safety on campus and creating a healthy and harmonious campus environment.

[0013] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0014] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating an anti-bullying alarm method based on AI voice recognition provided in an embodiment of this application is shown. Figure 2 A flowchart illustrating another anti-bullying alarm method based on AI voice recognition provided in an embodiment of this application is shown. Figure 3The diagram shows a structural schematic of an anti-bullying alarm system based on AI voice recognition provided in an embodiment of this application. Detailed Implementation

[0015] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present application can be combined with each other.

[0016] This embodiment provides an anti-bullying alarm method based on AI voice recognition, such as... Figure 1 As shown, the method includes: Step 101: Obtain the sound data to be identified by anti-bullying devices deployed in at least one pre-set bullying monitoring area on campus; Step 102: Perform speech recognition on the sound data to be identified, and when any target sound data in the sound data to be identified contains preset bullying key information, generate anti-bullying alarm information based on the target sound data and the anti-bullying device identifier corresponding to the target sound data; Step 103: Report the anti-bullying alarm information to the management terminal, so that the management terminal outputs anti-bullying prompt information based on the target bullying monitoring area corresponding to the anti-bullying device identifier and the target sound data; Step 104: Receive the anti-bullying information from the management terminal, and output the anti-bullying information via voice through the target anti-bullying device corresponding to the anti-bullying device identifier.

[0017] In this embodiment, traditional school bullying monitoring relies on manual patrols and student reports, which suffers from limited coverage and difficulty in real-time detection of covert bullying behavior. Therefore, this invention deploys anti-bullying devices in specific areas of the school campus to collect sound data in real-time and extensively. These devices can be installed in classrooms, corridors, playgrounds, and other locations where bullying may occur, without strict time and space limitations. They can comprehensively capture sound information from every corner of the campus, overcoming the limitations of manual patrols and providing fundamental data support for accurate identification of bullying behavior. Next, AI speech recognition technology is used to analyze the sound data. By using preset bullying key information (such as specific words and phrases like threats and insults), potential bullying behavior can be quickly and accurately identified. Once target sound data containing preset bullying key information is identified, anti-bullying alarm information is generated by combining it with the corresponding anti-bullying device identifier, enabling real-time and accurate identification of bullying behavior. Furthermore, to avoid the problem that even if bullying behavior is detected using traditional monitoring methods, it is difficult to promptly and accurately transmit the information to relevant management personnel for processing. This embodiment reports the generated anti-bullying alarm information to the management terminal. The management terminal can quickly locate the target bullying monitoring area based on the anti-bullying device identifier and output anti-bullying prompts based on the target's sound data. This allows managers to understand the specific location and general situation of the bullying behavior as soon as possible, improving the efficiency and accuracy of information transmission and ensuring timely measures to stop the bullying behavior. Finally, to address the problem that in the past, due to untimely information transmission and processing, bullying behavior was often difficult to stop quickly after it was discovered, this embodiment outputs voice information through the corresponding target anti-bullying device after receiving the anti-bullying stop information from the management terminal. This allows for direct sound warnings at the scene of the bullying, deterring the bullying student and preventing further escalation of the bullying behavior, thus avoiding more serious harm to the bullied student and effectively solving the problem of traditional methods being unable to quickly stop bullying behavior.

[0018] By applying the technical solution of this embodiment, the anti-bullying alarm method based on AI voice recognition collects sound data by deploying anti-bullying devices in a preset area on campus, identifies bullying behavior using AI voice recognition technology, promptly reports alarm information to the management terminal and locates the target area, and finally outputs stop-the-bullying information through the target device. This series of operations achieves real-time and comprehensive monitoring of bullying behavior on campus, enabling rapid and accurate detection and timely intervention. It overcomes the shortcomings of traditional monitoring and prevention methods, such as poor real-time performance, limited coverage, reliance on manual labor, and low efficiency, effectively ensuring student safety on campus and creating a healthy and harmonious campus environment.

[0019] Furthermore, as a refinement and extension of the specific implementation methods of the above embodiments, and to fully illustrate the specific implementation process of this embodiment, another anti-bullying alarm method based on AI voice recognition is provided, such as... Figure 2 As shown, the method includes: Step 201: When it is found that any target sound data in the sound data to be identified contains preset bullying key information, the target bullying monitoring area corresponding to the anti-bullying device identifier is determined according to the anti-bullying device identifier corresponding to the target sound data, and the surrounding monitoring device corresponding to the target bullying monitoring area is determined, and the monitoring data captured by the surrounding monitoring device is obtained. Step 202: Perform facial recognition on the monitoring data, determine the person to be investigated corresponding to the target voice data based on the facial recognition results, and obtain the individual voiceprint features corresponding to the person to be investigated. Step 203: Obtain the contextual audio data corresponding to the target sound data, and perform speaker voice separation on the contextual audio data to obtain multiple speaker voice data corresponding to the contextual audio data; Step 204: For any speaker's voice data, perform voiceprint feature recognition on the speaker's voice data using a pre-trained voiceprint model, and determine the bullying suspect based on the individual voiceprint features that match the recognized speaker's voiceprint features.

[0020] In this embodiment, after AI voice recognition detects that the target voice data contains preset bullying key information, the anti-bullying device identifier can be used to accurately locate the target bullying monitoring area. Next, the surrounding monitoring devices (e.g., several monitoring devices closest to the area) are identified and their monitoring data is acquired. Further, facial recognition is performed on the monitoring data to identify individuals from the image, thereby identifying potential suspects related to the target voice data. Simultaneously, the individual voiceprint features of these suspects are acquired. Voiceprint features are unique and stable, like fingerprints, and differ between individuals. Acquiring voiceprint features provides crucial biometric evidence for further accurate identification of bullying suspects, establishing a link between individual identity information and voice feature information, thus improving the accuracy of identifying individuals related to bullying behavior. Next, the contextual audio data of the target voice data is acquired to gain a more comprehensive understanding of the linguistic environment in which the bullying occurred. Then, speaker voice separation is performed on the contextual audio data to distinguish the voice data of different speakers. This allows for clear identification of the specific vocalizations of each speaker during the bullying incident, avoiding the analysis difficulties caused by the mixing of multiple voices. Finally, after speaker voice separation, suspects related to the bullying behavior can be identified. Specifically, a pre-trained voiceprint model is used to identify the voiceprint features of each speaker's voice data, extracting the speaker's voiceprint features. The identified voiceprint features are then matched with the previously acquired individual voiceprint features of the individuals to be investigated. Through voiceprint feature comparison, suspects related to the bullying behavior can be accurately identified, improving the efficiency and accuracy of identifying bullying suspects. Compared to traditional methods of monitoring and preventing school bullying, this method can detect bullying behavior more promptly and accurately, identify relevant responsible parties, and solve the problems of limited information acquisition and subjective inaccurate judgment in traditional methods, providing more scientific and efficient technical support for the prevention and handling of school bullying.

[0021] Optionally, in this embodiment, determining a bullying suspect based on individual voiceprint features matching the identified speaker's voiceprint features includes: calculating the similarity between the speaker's voiceprint features and each individual voiceprint feature, and determining the target individual voiceprint feature with the highest similarity; if the target individual voiceprint feature is greater than or equal to a preset similarity threshold, then the person to be investigated corresponding to the target individual voiceprint feature is determined as a bullying suspect; if the target individual voiceprint feature is less than the preset similarity threshold, then based on a preset number of supplementary investigations, supplementary monitoring devices are determined around the target bullying monitoring area, the monitoring data captured by the supplementary monitoring devices are acquired as new monitoring data, and the process returns to the step of performing facial recognition on the monitoring data to re-determine the bullying suspect.

[0022] In this embodiment, after identifying the speaker's voiceprint features using a voiceprint model, these features are compared with the previously acquired individual voiceprint features of the person to be investigated. A similarity calculation method is used here, which quantifies the closeness between two voiceprint features. The calculation process iterates through the individual voiceprint features of each person to be investigated, performing a similarity calculation with the current speaker's voiceprint features. Finally, the individual voiceprint feature with the highest similarity among all calculation results is selected as the target individual voiceprint feature. A preset similarity threshold is a pre-defined standard value used to determine the reliability of voiceprint feature matching. When the calculated similarity between the target individual voiceprint feature and the speaker's voiceprint feature is greater than or equal to this threshold, it indicates a high degree of similarity in voiceprint features, providing sufficient reason to believe that the person to be investigated corresponding to the target individual voiceprint feature is the person uttering the current speaker's voice. Therefore, this person to be investigated is identified as a bullying suspect, providing a clear target for subsequent processing and investigation. When the similarity between the target individual's voiceprint features and the speaker's voiceprint features is less than a preset similarity threshold, it indicates that the match between the person to be investigated and the speaker is not high, and there may be omissions or insufficient initial investigation scope. In this case, the number of surrounding supplementary monitoring devices to be called is determined according to a preset supplementary investigation quantity. The preset supplementary investigation quantity is a parameter set according to the actual situation to control the scope and extent of supplementary investigation. The monitoring data captured by these surrounding supplementary monitoring devices is acquired as new monitoring data, and then the process returns to the step of performing facial recognition on the monitoring data. Through facial recognition, new persons to be investigated are discovered, and voiceprint feature matching is performed again to identify bullying suspects. In this embodiment, when the initial investigation fails to accurately identify a suspect, it does not directly abandon the investigation or draw incorrect conclusions, but instead expands the investigation scope by calling surrounding supplementary monitoring devices. This method can collect as much relevant information as possible, avoid missing potential suspects, improve the accuracy and completeness of the entire anti-bullying alarm system in identifying suspects, ensure that the limitations of the initial investigation do not cause the real bullies to be missed, and further protect campus safety.

[0023] In this embodiment of the application, optionally, the step of performing voiceprint feature recognition on the speaker's voice data using a pre-trained voiceprint model includes: extracting speech features from the speaker's voice data using a speech feature extraction network in the pre-trained voiceprint model to obtain speaker speech features; extracting frame-level features from the speaker's speech features using a frame-level feature extraction network in the pre-trained voiceprint model; and converting the extracted frame-level features into sentence-level features using a sentence-level feature extraction network in the pre-trained voiceprint model to obtain speaker voiceprint features corresponding to the speaker's voice data.

[0024] In this embodiment, speech feature extraction is the first step in voiceprint recognition. The network extracts key features that reflect the speaker's identity from the original sound signal. For example, it extracts spectral features (such as Mel-frequency cepstral coefficients, MFCC), fundamental frequency features, and energy features. These features capture the physical properties of the sound, providing basic data for subsequent feature processing. After obtaining the speech features, the frame-level feature extraction network further processes the features of each time frame. It uses techniques such as convolutional neural networks (CNNs) or recurrent neural networks (RNNs) to mine local patterns and capture temporal information of the features in each frame. For example, CNNs can extract local feature patterns within a frame, while RNNs can capture temporal dependencies between frames. Furthermore, although frame-level features contain rich local information, in order to obtain voiceprint features that can represent the entire sentence, frame-level features can be integrated and transformed. The sentence-level feature extraction network uses techniques such as pooling operations (such as global average pooling) and fully connected layers to fuse frame-level features into a fixed-dimensional feature vector that can represent the entire sentence. This feature vector is the speaker's voiceprint feature, which integrates the sound information from the entire sentence, making the voiceprint feature more robust and representative. This feature can better cope with interference factors such as noise and pitch distortion in the sound, improving the accuracy and stability of voiceprint recognition, and providing a more reliable basis for subsequently identifying bullying suspects.

[0025] Optionally, in this embodiment of the application, before performing voiceprint feature recognition on the speaker's voice data using a pre-trained voiceprint model, the method further includes: acquiring regional noise data corresponding to each bullying monitoring area, and adding noise with corresponding bullying monitoring area features to the clean sound samples based on the regional noise data corresponding to each bullying monitoring area to obtain noise sound samples corresponding to each bullying monitoring area; constructing voiceprint model training samples corresponding to each bullying monitoring area based on the clean sound samples and the noise sound samples corresponding to each bullying monitoring area; and performing enhancement training on the pre-trained original voiceprint model based on the voiceprint model training samples corresponding to each bullying monitoring area for each bullying monitoring area to obtain the pre-trained voiceprint model corresponding to the bullying monitoring area. Accordingly, the step of performing voiceprint feature recognition on the speaker's voice data using a pre-trained voiceprint model includes: performing voiceprint feature recognition on the speaker's voice data using a pre-trained voiceprint model corresponding to the target bullying monitoring area.

[0026] In this embodiment, actual environmental noise data (such as corridor echoes, playground noise, classroom equipment noise, etc.) is collected from various bullying monitoring areas on campus. Clean sound samples are then superimposed with the noise from the corresponding areas to generate noise sound samples with regional characteristics. For example, background noise from sports scenes is added to the playground area, and specific noises such as the friction of desks and chairs and the sound of projector fans are added to the classroom area. Traditional voiceprint model training typically uses standard noise libraries, which are difficult to adapt to the complex and ever-changing real environment of a campus. This embodiment constructs training samples using customized regional noise, allowing the model to adapt to the actual acoustic environment of the target area in advance, effectively reducing the interference of environmental noise on voiceprint recognition and improving the robustness of the model in specific scenarios. Based on clean sound samples and noise sound samples from the corresponding areas, dedicated training sample sets for each bullying monitoring area are constructed. Each sample set contains mixed data of "clean speech + regional noise," covering various speech scenarios such as different speakers, different speaking speeds, and different emotional states. This allows the model to learn the unique noise patterns and speech feature associations specific to that region through training with region-specific sample sets. This achieves deep "scene-voiceprint" binding, avoiding recognition biases caused by environmental differences and improving the model's localization accuracy within the target area. Furthermore, for each bullying monitoring region, the initial voiceprint model is enhanced using its specific training samples. During training, the focus is on learning the fusion features of noise and clean speech in that region, ultimately generating a pre-trained voiceprint model adapted to the corresponding region. In application, the system automatically calls the corresponding pre-trained model for voiceprint recognition based on the target bullying monitoring region. Through region-enhanced training, the model can specifically strengthen its noise suppression and feature extraction capabilities in the target region. For example, in the playground area model, the filtering of wind noise and movement sounds is enhanced; in the classroom area model, the noise reduction processing of desk and chair friction sounds is optimized, ultimately achieving "one model per region" adaptation and improving the accuracy of voiceprint recognition in complex environments. During the voiceprint feature recognition stage, the system automatically calls the pre-trained voiceprint model corresponding to the target bullying monitoring area to extract and match features from the speaker's voice data. For example, if the bullying behavior occurs in the playground area, a playground-specific model is used for voiceprint analysis, rather than a general model, to ensure that voiceprint recognition is always performed under the optimal model.

[0027] Optionally, in this embodiment of the application, the step of enhancing the original voiceprint model after preliminary training based on the voiceprint model training samples corresponding to the bullying monitoring area includes: The original voiceprint model is copied to obtain a copied voiceprint model; The pure speech features are obtained by extracting speech features from the pure sound sample through the speech feature extraction network in the original voiceprint model; the pure speech features are obtained by extracting frame-level features through the frame-level feature extraction network in the original voiceprint model; and the pure frame-level features are converted into pure sentence-level features through the sentence-level feature extraction network in the original voiceprint model. The noise speech features are obtained by extracting speech features from the noise sound sample through the speech feature extraction network in the replicated voiceprint model; the noise speech features are obtained by extracting frame-level features through the frame-level feature extraction network in the replicated voiceprint model; and the noise frame-level features are converted into noise sentence-level features through the sentence-level feature extraction network in the replicated voiceprint model. Calculate a first loss between the clean frame-level features and the noisy frame-level features, calculate a second loss between the clean sentence-level features and the noisy sentence-level features, and optimize the replicated voiceprint model based on the first loss and the second loss to achieve enhanced training of the original voiceprint model that has been initially trained.

[0028] In this embodiment, the original voiceprint model, after initial training, is replicated to generate an identical "replicated voiceprint model." This retains the core structure and parameters of the original model, serving as the foundational framework for subsequent enhancement training and avoiding the loss of training results due to direct modification of the original model. The original voiceprint model is used to perform a three-stage feature extraction process on clean sound samples: speech feature extraction: acquiring physical features such as basic spectrum and fundamental frequency through a speech feature extraction network; frame-level feature extraction: mining local temporal patterns (such as phoneme transitions and intonation changes) through a frame-level network; and sentence-level feature extraction: integrating global features through a sentence-level network to generate a voiceprint feature vector representing clean speech. The replicated voiceprint model is then used to perform a feature extraction process on noisy sound samples that is completely symmetrical to that of clean samples, generating speech features, frame-level features, and sentence-level features in a noisy environment. Next, the feature differences between clean and noisy samples at the frame level (first loss) and sentence level (second loss) are calculated: the first loss measures the degree of distortion of frame-level features under noise interference; the second loss measures the loss of semantic integrity of sentence-level features in a noisy environment.

[0029] The parameters of the replicated voiceprint model are optimized based on the dual loss method, and the original voiceprint model is enhanced by backpropagating the model parameters.

[0030] Optionally, in this embodiment, optimizing the replicated voiceprint model based on the first loss and the second loss includes: if the first loss is greater than the second loss, then optimizing the frame-level feature extraction network in the replicated voiceprint model separately; otherwise, jointly optimizing the frame-level feature extraction network and the sentence-level feature extraction network in the replicated voiceprint model.

[0031] In this embodiment, an optimization strategy is dynamically selected by comparing the differences in frame-level features (first loss) and sentence-level features (second loss). Specifically, when the first loss is greater than the second loss, it indicates that the distortion of frame-level features under noise interference is higher than that of sentence-level features. In this case, the frame-level feature extraction network is optimized separately to address the noise resistance problem of local temporal patterns (such as phoneme transitions and intonation fluctuations), avoiding interference from sentence-level optimization on the frame level. When the first loss is less than or equal to the second loss, it indicates that the semantic integrity loss of sentence-level features in noise is more prominent. In this case, the frame-level and sentence-level feature extraction networks are jointly optimized, and the ability to preserve local details and global semantic stability are improved synchronously through end-to-end training, ensuring the consistency of the two levels of features in a noisy environment. Thus, precise optimization is achieved through loss value comparison. In addition, this embodiment avoids the risk of "over-adjustment" of different levels of features by using the same optimization strategy through dynamic switching between separate optimization and joint optimization. For example, if joint optimization is forced when frame-level issues are prominent, it may lead to overfitting of sentence-level features to noise patterns; conversely, if only frame-level optimization is performed when sentence-level issues are prominent, it may fail to address the global semantic loss. This strategy ensures the model's stable performance in complex campus acoustic environments through contextual processing. Noise patterns differ significantly in different areas of the campus (e.g., the friction sound of desks and chairs in classrooms, echoes in corridors, and voices on the playground), and this optimization strategy can dynamically adapt to these differences. For example, in the classroom area, if frame-level features are more severely affected by projector fan noise, then frame-level networks are optimized first; in the playground area, if sentence-level features are more significantly affected by shouts, then joint optimization is initiated to avoid optimization conflicts and over-adjustment, ensuring the model's continued stable performance in complex environments.

[0032] Optionally, after reporting the anti-bullying alarm information to the management terminal, the method further includes: obtaining the terminal location information corresponding to the security terminal worn by the security personnel; determining the target security terminal based on the distance between the terminal location information and the target bullying monitoring area; and sending the anti-bullying alarm information to the target security terminal so that the target security terminal displays inspection prompt information based on the anti-bullying alarm information, prompting the security personnel to go to the target bullying monitoring area for inspection.

[0033] In this embodiment, real-time location data from security personnel's devices is collected via GPS, Bluetooth beacons, or campus Wi-Fi positioning systems to create a dynamic location coordinate database. For example, smart bracelets worn by security patrols can report latitude and longitude information every 2 seconds. Based on the geographic coordinates of the target bullying monitoring area (such as the northeast corner of the playground), the straight-line distance or path distance between all security terminals and that area is calculated. Anti-bullying alarm information, including the specific location, time, and sound characteristic summary of the bullying incident, is pushed to the target security terminal via 4G / 5G networks or a dedicated campus communication network. The terminal device simultaneously alerts the security personnel through vibration, voice broadcast, and pop-up windows, ensuring that security personnel are informed immediately and arrive at the scene before the bullying escalates, effectively preventing the situation from worsening.

[0034] Optionally, in this embodiment of the application, the method further includes: acquiring monitoring image information from various monitoring devices on campus, and performing image recognition on the monitoring image information to identify bullying-suspected images in the monitoring image information; determining the bullying monitoring area corresponding to the location of the monitoring device as a warning area based on the location of the monitoring device corresponding to the bullying-suspected image; and outputting anti-bullying warning information through the anti-bullying device corresponding to the warning area to remind users to prohibit bullying behavior.

[0035] In this embodiment, video stream data from various areas is collected in real time through a network of surveillance cameras deployed on campus. Deep learning-based image recognition algorithms (such as YOLOv7 object detection + 3D convolutional behavior analysis) are used to identify bullying behaviors in the videos, such as physical conflict, blocking, and pushing. For example, the system can detect abnormal proximity of two or more people, abnormal frequency of physical contact, and facial expressions of pain, combining these features with time-series analysis to determine if bullying is suspected. Based on the surveillance device ID of the identified suspected bullying image, the coordinates of the surveillance device are mapped to the corresponding bullying monitoring area through a pre-built campus GIS geographic information system. Anti-bullying devices (such as smart broadcasting and electronic fences) in the alert area are triggered via IoT protocols (such as MQTT) to output anti-bullying alert information in real time. For example, the device may play a pre-recorded warning message, "Abnormal behavior detected, please stop immediately," and may also activate the electronic fence's audio-visual alarm function, creating a dual deterrent effect of visual and auditory means.

[0036] Furthermore, embodiments of this application provide an anti-bullying alarm system based on AI voice recognition, such as... Figure 3 As shown, the system includes anti-bullying devices and control devices. The anti-bullying devices are deployed in at least one pre-designated bullying monitoring area on campus. The anti-bullying devices and the control devices are used to implement the aforementioned AI-based voice recognition-based anti-bullying alarm method. For a corresponding description of the system, please refer to the corresponding description in the above method; it will not be repeated here.

[0037] This application also provides a computer device, specifically a personal computer, server, network device, etc. The computer device includes a bus, processor, memory, and communication interface, and may also include input / output interfaces and a display device. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores location information. The network interface of the computer device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps in the various method embodiments.

[0038] Those skilled in the art will understand that the structure of the computer device described above is only a partial structure related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. A specific computer device may include more or fewer components, or combine certain components, or have different component arrangements.

[0039] In one embodiment, a computer-readable storage medium is provided, which may be non-volatile or volatile, having stored thereon a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0040] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0041] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0042] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, graphics processors, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0043] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0044] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A bullying alarm method based on AI voice recognition, characterized in that, The method includes: Acquire the voice data to be identified by anti-bullying devices deployed in at least one pre-set bullying monitoring area on campus; The voice data to be identified is subjected to speech recognition, and when any target voice data in the voice data to be identified contains preset bullying key information, anti-bullying alarm information is generated based on the target voice data and the anti-bullying device identifier corresponding to the target voice data. The anti-bullying alarm information is reported to the management terminal, so that the management terminal outputs anti-bullying prompt information based on the target bullying monitoring area corresponding to the anti-bullying device identifier and the target sound data; Receive anti-bullying information from the management terminal, and output the anti-bullying information via voice through the target anti-bullying device corresponding to the anti-bullying device identifier; The method further includes: When any target sound data in the sound data to be identified contains preset bullying key information, the target bullying monitoring area corresponding to the anti-bullying device identifier is determined according to the anti-bullying device identifier corresponding to the target sound data, and the surrounding monitoring device corresponding to the target bullying monitoring area is determined, and the monitoring data captured by the surrounding monitoring device is obtained. The monitoring data is subjected to facial recognition. Based on the facial recognition results, the person to be investigated corresponding to the target voice data is determined, and the individual voiceprint features corresponding to the person to be investigated are obtained. Obtain the contextual audio data corresponding to the target sound data, and perform speaker voice separation on the contextual audio data to obtain multiple speaker voice data corresponding to the contextual audio data; For any speaker's voice data, a pre-trained voiceprint model is used to identify the speaker's voice features, and the bullying suspect is identified based on the individual voiceprint features that match the identified speaker's voiceprint features.

2. The anti-bullying alarm method based on AI voice recognition according to claim 1, characterized in that, The process of identifying a suspected bully based on individual voiceprint features matched with the identified speaker's voiceprint features includes: Calculate the similarity between the speaker's voiceprint features and the voiceprint features of each individual, and determine the target individual's voiceprint features with the highest similarity. If the voiceprint feature of the target individual is greater than or equal to a preset similarity threshold, then the person to be investigated corresponding to the voiceprint feature of the target individual is identified as a bullying suspect. If the voiceprint feature of the target individual is less than a preset similarity threshold, then based on the preset number of supplementary screenings, supplementary monitoring devices around the target bullying monitoring area are determined, the monitoring data captured by the supplementary monitoring devices are acquired as new monitoring data, and the process returns to the step of performing facial recognition on the monitoring data to re-identify the bullying suspect.

3. The anti-bullying alarm method based on AI voice recognition according to claim 1, characterized in that, The process of recognizing speaker voice features using a pre-trained voiceprint model includes: The speaker's voice data is processed by a speech feature extraction network in a pre-trained voiceprint model to extract speech features, thereby obtaining speaker speech features; the speaker's speech features are then processed by a frame-level feature extraction network in a pre-trained voiceprint model to extract frame-level features; and the extracted frame-level features are then converted into sentence-level features by a sentence-level feature extraction network in a pre-trained voiceprint model to obtain speaker voiceprint features corresponding to the speaker's voice data.

4. The anti-bullying alarm method based on AI voice recognition according to claim 3, characterized in that, Before performing voiceprint feature recognition on the speaker's voice data using a pre-trained voiceprint model, the method further includes: Obtain the regional noise data corresponding to each bullying monitoring area, and add noise with the characteristics of the corresponding bullying monitoring area to the clean sound samples based on the regional noise data corresponding to each bullying monitoring area to obtain the noise sound samples corresponding to each bullying monitoring area. Based on the clean sound samples and the noise sound samples corresponding to each bullying monitoring area, voiceprint model training samples corresponding to each bullying monitoring area are constructed respectively. For each bullying monitoring area, the original voiceprint model that has been initially trained is enhanced based on the voiceprint model training samples corresponding to the bullying monitoring area to obtain the pre-trained voiceprint model corresponding to the bullying monitoring area. Accordingly, the process of recognizing speaker voice features using a pre-trained voiceprint model includes: Voiceprint feature recognition is performed on the speaker's voice data using a pre-trained voiceprint model corresponding to the target bullying monitoring area.

5. The anti-bullying alarm method based on AI voice recognition according to claim 4, characterized in that, The enhancement training of the original voiceprint model, which has undergone preliminary training, based on the voiceprint model training samples corresponding to the bullying monitoring area includes: The original voiceprint model is copied to obtain a copied voiceprint model; The pure speech features are obtained by extracting speech features from the pure sound sample through the speech feature extraction network in the original voiceprint model; the pure speech features are obtained by extracting frame-level features through the frame-level feature extraction network in the original voiceprint model; and the pure frame-level features are converted into pure sentence-level features through the sentence-level feature extraction network in the original voiceprint model. The noise speech features are obtained by extracting speech features from the noise sound sample through the speech feature extraction network in the replicated voiceprint model; the noise speech features are obtained by extracting frame-level features through the frame-level feature extraction network in the replicated voiceprint model; and the noise frame-level features are converted into noise sentence-level features through the sentence-level feature extraction network in the replicated voiceprint model. Calculate a first loss between the clean frame-level features and the noisy frame-level features, calculate a second loss between the clean sentence-level features and the noisy sentence-level features, and optimize the replicated voiceprint model based on the first loss and the second loss to achieve enhanced training of the original voiceprint model that has been initially trained.

6. The anti-bullying alarm method based on AI voice recognition according to claim 5, characterized in that, The optimization of the replicated voiceprint model based on the first loss and the second loss includes: If the first loss is greater than the second loss, then the frame-level feature extraction network in the replicated voiceprint model is optimized separately; otherwise, the frame-level feature extraction network and the sentence-level feature extraction network in the replicated voiceprint model are jointly optimized.

7. The anti-bullying alarm method based on AI voice recognition according to any one of claims 1 to 6, characterized in that, After reporting the anti-bullying alarm information to the management terminal, the method further includes: Obtain the terminal location information corresponding to the security terminal worn by the security personnel, and determine the target security terminal based on the distance between the terminal location information and the target bullying monitoring area; The anti-bullying alarm information is sent to the target security terminal, so that the target security terminal displays inspection prompt information based on the anti-bullying alarm information, prompting security personnel to go to the target bullying monitoring area for inspection.

8. The anti-bullying alarm method based on AI voice recognition according to any one of claims 1 to 6, characterized in that, The method further includes: Acquire monitoring image information from various monitoring devices on campus, and perform image recognition on the monitoring image information to identify images suspected of bullying in the monitoring image information; Based on the location of the monitoring device corresponding to the suspected bullying image, the bullying monitoring area corresponding to the location of the monitoring device is determined as the alert area; The anti-bullying device corresponding to the prompt area outputs anti-bullying prompt information to remind users to refrain from bullying behavior.

9. An anti-bullying alarm system based on AI voice recognition, characterized in that, The system includes anti-bullying equipment and control equipment. The anti-bullying equipment is deployed in at least one pre-set bullying monitoring area on campus. The anti-bullying equipment and the control equipment are used to implement the anti-bullying alarm method based on AI voice recognition as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Suspect recognition method and device

    CN110309744A

  • Campus student deception behavior identification method

    CN113128383A

  • Voice quality evaluation model training method and device and storage medium

    CN113870899A

  • Intelligent campus student behavior analysis system based on artificial intelligence

    CN117237155A

  • Intelligent AI real-time dialogue alarm system for campus management

    CN117409532A