Anti-bullying alarm method and system based on AI voice recognition

By deploying anti-bullying equipment on campus and utilizing AI voice recognition and voiceprint technology, bullying behavior can be monitored and located in real time. This solves the problems of real-time monitoring and coverage of traditional monitoring methods, enabling rapid and accurate detection and prevention of bullying and ensuring student safety.

CN120932367BActive Publication Date: 2026-02-13SHANGHAI LIQING INTELLIGENT TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511443483.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2026-02-13
Estimated Expiration
2045-10-10

AI Technical Summary

Technical Problem

Existing methods for monitoring and preventing school bullying suffer from poor real-time performance, limited coverage, reliance on manual labor, and low efficiency, making it difficult to detect and stop bullying behavior in a timely manner.

Method used

By deploying anti-bullying devices in specific areas of the campus, AI voice recognition technology is used to collect sound data in real time, identify key information about bullying, generate alarm information, locate the target area, output stop-the-bullying information, and identify bullying suspects by combining facial recognition and voiceprint feature recognition.

Benefits of technology

It enables real-time and comprehensive monitoring of school bullying, allowing for rapid and accurate detection and timely intervention, improving the efficiency and accuracy of information transmission, ensuring student safety, and creating a healthy and harmonious campus environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932367B_ABST
    Figure CN120932367B_ABST
Patent Text Reader

Abstract

The application discloses an anti-bullying alarm method and system based on AI voice recognition. The method comprises the following steps: deploying anti-bullying devices in at least one bullying monitoring area in the campus, and acquiring sound data to be recognized collected by each anti-bullying device; performing voice recognition on the sound data to be recognized, and when it is recognized that any target sound data in the sound data to be recognized contains preset bullying key information, generating anti-bullying alarm information according to the target sound data and the anti-bullying device identifier corresponding to the target sound data; reporting the anti-bullying alarm information to a management terminal, so that the management terminal outputs anti-bullying prompt information based on the target bullying monitoring area corresponding to the anti-bullying device identifier and the target sound data; receiving anti-bullying suppression information of the management terminal, and performing voice output of the anti-bullying suppression information through the target anti-bullying device corresponding to the anti-bullying device identifier.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of voice analysis, in particular to a bullying prevention alarm method and system based on AI voice recognition. BACKGROUND

[0002] Campus bullying, as a serious problem in the global education field, seriously endangers the physical and mental health and growth of students. In recent years, with the increasing attention to campus safety, how to effectively prevent and promptly stop campus bullying behavior has become the focus of attention of educators, parents and the community.

[0003] Traditional campus bullying monitoring and prevention methods mainly rely on manual patrols and student self-reporting. Although manual patrols can discover some bullying behavior to some extent, they have obvious limitations. The campus is extensive, and bullying behavior can occur in classrooms, corridors, playground corners and other hidden places, making it difficult for manual patrols to achieve comprehensive and real-time coverage, resulting in many bullying incidents being unable to be detected in the early stages. Moreover, manual patrols require a large amount of manpower and resources, and are inefficient, making it difficult to discover and stop some sudden and highly concealed bullying behavior in the first instance.

[0004] Student self-reporting of bullying incidents also faces many problems. On the one hand, some bullied students are afraid to report their bullying to teachers or parents due to fear, shame, and other reasons, fearing that they will be retaliated against. On the other hand, some bystanders may choose to remain silent and not report the bullying they have witnessed, either because they are afraid of being implicated or because they feel it is not their business. This makes many bullying incidents go undiscovered for a long time after they occur, causing continuous physical and psychological harm to the bullied students.

[0005] In addition, existing campus monitoring systems can record pictures within the campus, but have limited ability to capture and analyze sound information. Bullying behavior often involves specific language content, such as threats and insults, making it difficult to accurately determine whether bullying has occurred through picture monitoring alone, and making it impossible to obtain key information during the bullying process in a timely manner. Moreover, the video data volume of the monitoring system is large, and manual viewing and analysis of this data requires a large amount of time and effort, making it difficult to achieve real-time monitoring and rapid response.

[0006] In summary, existing methods of monitoring and preventing campus bullying have poor real-time performance, limited coverage, rely on manual work and are inefficient, and are unable to timely and effectively discover and stop campus bullying behavior. SUMMARY

[0007] Therefore, the application provides a bullying prevention alarm method and system based on AI voice recognition, which realizes real-time and comprehensive monitoring of bullying behaviors in a campus, can quickly and accurately find bullying behaviors and timely stop them, overcomes the defects of traditional monitoring and prevention methods, such as poor real-time performance, limited coverage, dependence on manual work and low efficiency, effectively guarantees the safety of students in the campus, and creates a healthy and harmonious campus environment.

[0008] According to an aspect of the application, a bullying prevention alarm method based on AI voice recognition is provided, which comprises:

[0009] obtaining sound data to be identified collected by bullying prevention devices respectively arranged in at least one bullying monitoring area preset in a campus;

[0010] performing voice recognition on the sound data to be identified, and when it is identified that any target sound data in the sound data to be identified contains preset bullying key information, generating bullying prevention alarm information according to the target sound data and a bullying prevention device identifier corresponding to the target sound data;

[0011] reporting the bullying prevention alarm information to a management terminal, so that the management terminal outputs bullying prevention prompt information based on a target bullying monitoring area corresponding to the bullying prevention device identifier and the target sound data;

[0012] receiving bullying prevention stop information of the management terminal, and performing voice output of the bullying prevention stop information through a target bullying prevention device corresponding to the bullying prevention device identifier;

[0013] The method further comprises:

[0014] when it is identified that any target sound data in the sound data to be identified contains preset bullying key information, determining a target bullying monitoring area corresponding to the bullying prevention device identifier according to the bullying prevention device identifier, and determining a surrounding monitoring device corresponding to the target bullying monitoring area, and obtaining monitoring data photographed by the surrounding monitoring device;

[0015] performing face recognition on the monitoring data, determining a person to be investigated corresponding to the target sound data according to a face recognition result, and obtaining an individual voiceprint feature corresponding to the person to be investigated;

[0016] obtaining context audio data corresponding to the target sound data, performing speaker voice separation on the context audio data to obtain a plurality of speaker voice data corresponding to the context audio data;

[0017] For any speaker voice data, the voiceprint feature of the speaker voice data is recognized by a pre-trained voiceprint model, and a bullying suspect is determined according to an individual voiceprint feature matched with the recognized speaker voiceprint feature.

[0018] According to another aspect of the present application, there is provided an anti-bullying alarm system based on AI voice recognition, comprising an anti-bullying device and a control device, the anti-bullying device being deployed in at least one bullying monitoring area preset in a campus, and the anti-bullying device and the control device being used to implement the anti-bullying alarm method based on AI voice recognition described above.

[0019] According to still another aspect of the present application, there is provided a storage medium having a computer program stored thereon, the program being executed by a processor to implement the anti-bullying alarm method based on AI voice recognition described above.

[0020] According to still another aspect of the present application, there is provided a computer device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, the processor implementing the anti-bullying alarm method based on AI voice recognition described above when executing the program.

[0021] By means of the technical solutions described above, the anti-bullying alarm method and system based on AI voice recognition provided by the embodiments of the present application achieve real-time and comprehensive monitoring of bullying behavior in a campus by deploying an anti-bullying device in a preset area of the campus to collect voice data, using AI voice recognition technology to identify bullying behavior, timely reporting alarm information to a management terminal and locating a target area, and finally outputting a stop message through a target device. This series of operations realizes real-time and comprehensive monitoring of bullying behavior in a campus, can quickly and accurately discover bullying behavior and timely stop it, overcomes the defects of traditional monitoring and prevention methods such as poor real-time performance, limited coverage, dependence on manual work, and low efficiency, effectively safeguards the campus safety of students, and creates a healthy and harmonious campus environment.

[0022] The above description is only a summary of the technical solutions of the present application, in order to enable a clearer understanding of the technical means of the present application, and the technical solutions can be implemented in accordance with the content of the description, and in order to enable the above and other purposes, features and advantages of the present application to be more apparent and easy to understand, the following specific embodiments of the present application are described. BRIEF DESCRIPTION OF DRAWINGS

[0023] The drawings described herein are used to provide further understanding of the present application, and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation on the present application. In the drawings:

[0024] Figure 1 Fig. 1 shows a flowchart of an anti-bullying alarm method based on AI voice recognition provided by an embodiment of the present application;

[0025] Figure 2 A flowchart of another anti-bullying alarm method based on AI voice recognition provided by an embodiment of the present application is shown.

[0026] Figure 3 A structural diagram of an anti-bullying alarm system based on AI voice recognition provided by an embodiment of the present application is shown. DETAILED DESCRIPTION

[0027] Hereinafter, the present application will be described in detail with reference to the accompanying drawings and in conjunction with embodiments. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.

[0028] An anti-bullying alarm method based on AI voice recognition is provided in the present embodiment, as shown in the figure, the method comprises: Figure 1

[0029] Step 101, acquiring sound data to be identified collected by anti-bullying devices respectively deployed in at least one bullying monitoring area preset in a campus;

[0030] Step 102, performing voice recognition on the sound data to be identified, and when it is identified that any target sound data in the sound data to be identified contains preset bullying key information, generating anti-bullying alarm information according to the target sound data and an anti-bullying device identifier corresponding to the target sound data;

[0031] Step 103, reporting the anti-bullying alarm information to a management terminal, so that the management terminal outputs anti-bullying prompt information based on a target bullying monitoring area corresponding to the anti-bullying device identifier and the target sound data;

[0032] Step 104, receiving anti-bullying stop information of the management terminal, and performing voice output of the anti-bullying stop information through a target anti-bullying device corresponding to the anti-bullying device identifier.

[0033] ​In the embodiments of the present application, since the traditional campus bullying monitoring relies on manual patrol and student active reporting, there are problems of limited coverage and difficulty in discovering hidden bullying behavior in real time. Therefore, by deploying anti-bullying devices in specific areas of the campus, these devices are used to collect sound data in the campus in real time and extensively. The anti-bullying devices can be installed in classrooms, corridors, playgrounds and other places where bullying behavior may occur, and are not strictly limited by time and space, so they can fully capture sound information in every corner of the campus, overcoming the defect that manual patrol cannot fully cover the campus, and providing basic data support for subsequent accurate identification of bullying behavior. Then, the AI voice recognition technology is used to analyze the sound data to be identified, and through the pre-set bullying key information (such as specific words and sentence patterns such as threats and insults), the possible bullying behavior can be quickly and accurately identified. Once the target sound data containing the pre-set bullying key information is identified, the anti-bullying alarm information is generated in combination with the corresponding anti-bullying device identifier, so as to realize real-time and accurate identification of bullying behavior. Further, in order to avoid the problem that even if bullying behavior is found in the traditional monitoring method, it is difficult to timely and accurately transmit the information to the relevant management personnel for processing. The anti-bullying alarm information generated in the embodiments of the present application is reported to the management terminal, and the management terminal can quickly locate the target bullying monitoring area according to the anti-bullying device identifier, and output the anti-bullying prompt information in combination with the target sound data. This enables the management personnel to learn about the specific location and general situation of the bullying behavior in the first time, improves the efficiency and accuracy of information transmission, and provides a guarantee for timely taking measures to stop bullying behavior. Finally, in order to solve the problem that after discovering bullying behavior, due to the delay in information transmission and processing, it is often difficult to quickly stop the bullying behavior, after receiving the anti-bullying stop information of the management terminal, the corresponding target anti-bullying device is used to output voice. In this way, the stop sound can be directly emitted at the scene of the bullying behavior, which can deter the students who implement bullying and timely stop the further development of bullying behavior, avoid the bullying students from suffering more serious harm, and effectively solve the problem that the traditional method is difficult to quickly stop bullying behavior.

[0034] By applying the technical solutions of the embodiments, the anti-bullying alarm method based on AI voice recognition collects sound data by deploying anti-bullying devices in pre-set areas of the campus, identifies bullying behavior by using AI voice recognition technology, timely reports alarm information to the management terminal and locates the target area, and finally outputs stop information through the target device. This series of operations realizes real-time and comprehensive monitoring of bullying behavior in the campus, can quickly and accurately discover bullying behavior and timely stop it, overcomes the defects of traditional monitoring and prevention methods such as poor real-time performance, limited coverage, dependence on manual work and low efficiency, effectively guarantees the safety of students in the campus, and creates a healthy and harmonious campus environment.

[0035] Further, as a refinement and expansion of the above embodiment, in order to complete the specific implementation process of the embodiment, another anti-bullying alarm method based on AI voice recognition is provided, as shown in Figure 2 The method comprises the following steps:

[0036] Step 201, when it is identified that any target sound data in the to-be-identified sound data contains preset bullying key information, determining a target bullying monitoring area corresponding to the anti-bullying device according to the anti-bullying device identifier corresponding to the target sound data, and determining a surrounding monitoring device corresponding to the target bullying monitoring area, and acquiring monitoring data photographed by the surrounding monitoring device;

[0037] Step 202, performing face recognition on the monitoring data, determining a to-be-investigated person corresponding to the target sound data according to the face recognition result, and acquiring individual voiceprint features corresponding to the to-be-investigated person;

[0038] Step 203, acquiring context audio data corresponding to the target sound data, performing speaker voice separation on the context audio data to obtain a plurality of speaker voice data corresponding to the context audio data;

[0039] Step 204, for any speaker voice data, performing voiceprint feature recognition on the speaker voice data through a pre-trained voiceprint model, and determining a bullying suspect according to individual voiceprint features matched with the recognized speaker voiceprint features.

[0040] In the embodiments of the present application, after the target sound data is found to contain the preset bullying key information through AI voice recognition, the target bullying monitoring area can be accurately located by using the anti-bullying device identifier. Then the surrounding monitoring devices of the area (for example, the nearest monitoring devices to the area) are determined and the monitoring data is obtained. Further, face recognition is performed on the monitoring data to identify the identities of the personnel in the picture, and then the personnel to be investigated who may be related to the target sound data are determined. At the same time, the individual voiceprint features of these personnel to be investigated are obtained. The voiceprint feature has uniqueness and stability, just like everyone's fingerprint. The voiceprint features of different people are different. By obtaining the voiceprint feature, a key biometric basis is provided for further accurate identification of bullying suspects in the future. The identity information of the personnel is linked with the sound feature information, and the accuracy of identifying the personnel related to the bullying behavior is improved. Then, the context audio data of the target sound data is obtained to better understand the language environment when the bullying behavior occurs. Then, speaker voice separation is performed on the context audio data to distinguish the sound data of different speakers. In this way, the specific sound emission of each speaker during the bullying behavior can be clearly distinguished, and the analysis difficulty caused by the mixing of multiple voices is avoided. Finally, after completing the speaker voice separation, the suspect related to the bullying behavior can also be identified. Specifically, the voiceprint feature of each speaker is identified by using a pre-trained voiceprint model, and the voiceprint feature of the speaker is extracted by the voiceprint model. Then, the identified voiceprint feature is matched with the individual voiceprint features of the personnel to be investigated obtained before. Through the comparison of the voiceprint features, the suspect related to the bullying behavior can be accurately determined, and the efficiency and accuracy of determining the bullying suspect are improved. Compared with the traditional campus bullying monitoring and prevention methods, this method can more timely and accurately discover bullying behavior and determine the relevant responsible person, solving the problems of single information acquisition and subjective inaccuracy in traditional methods, and providing more scientific and efficient technical support for the prevention and handling of campus bullying.

[0041] In the embodiments of the present application, the bullying suspect is determined according to the individual voiceprint feature matched with the identified speaker voiceprint feature, including: calculating the similarity between the speaker voiceprint feature and each individual voiceprint feature respectively, determining the target individual voiceprint feature with the highest similarity; if the target individual voiceprint feature is greater than or equal to a preset similarity threshold, the personnel to be investigated corresponding to the target individual voiceprint feature is determined as the bullying suspect; if the target individual voiceprint feature is less than the preset similarity threshold, the surrounding supplementary monitoring devices of the target bullying monitoring area are determined based on a preset investigation supplementary quantity, the monitoring data shot by the surrounding supplementary monitoring devices is obtained as new monitoring data, and the step of performing face recognition on the monitoring data is returned to, so as to redetermine the bullying suspect.

[0042] In this embodiment, after identifying the speaker's voiceprint features using a voiceprint model, these features are compared with the previously acquired individual voiceprint features of the person to be investigated. A similarity calculation method is used here, which quantifies the closeness between two voiceprint features. The calculation process iterates through the individual voiceprint features of each person to be investigated, performing a similarity calculation with the current speaker's voiceprint features. Finally, the individual voiceprint feature with the highest similarity among all calculation results is selected as the target individual voiceprint feature. A preset similarity threshold is a pre-defined standard value used to determine the reliability of voiceprint feature matching. When the calculated similarity between the target individual voiceprint feature and the speaker's voiceprint feature is greater than or equal to this threshold, it indicates a high degree of similarity in voiceprint features, providing sufficient reason to believe that the person to be investigated corresponding to the target individual voiceprint feature is the person uttering the current speaker's voice. Therefore, this person to be investigated is identified as a bullying suspect, providing a clear target for subsequent processing and investigation. When the similarity between the target individual's voiceprint features and the speaker's voiceprint features is less than a preset similarity threshold, it indicates that the match between the person to be investigated and the speaker is not high, and there may be omissions or insufficient initial investigation scope. In this case, the number of surrounding supplementary monitoring devices to be called is determined according to a preset supplementary investigation quantity. The preset supplementary investigation quantity is a parameter set according to the actual situation to control the scope and extent of supplementary investigation. The monitoring data captured by these surrounding supplementary monitoring devices is acquired as new monitoring data, and then the process returns to the step of performing facial recognition on the monitoring data. Through facial recognition, new persons to be investigated are discovered, and voiceprint feature matching is performed again to identify bullying suspects. In this embodiment, when the initial investigation fails to accurately identify a suspect, it does not directly abandon the investigation or draw incorrect conclusions, but instead expands the investigation scope by calling surrounding supplementary monitoring devices. This method can collect as much relevant information as possible, avoid missing potential suspects, improve the accuracy and completeness of the entire anti-bullying alarm system in identifying suspects, ensure that the limitations of the initial investigation do not cause the real bullies to be missed, and further protect campus safety.

[0043] In this embodiment of the application, optionally, the step of performing voiceprint feature recognition on the speaker's voice data using a pre-trained voiceprint model includes: extracting speech features from the speaker's voice data using a speech feature extraction network in the pre-trained voiceprint model to obtain speaker speech features; extracting frame-level features from the speaker's speech features using a frame-level feature extraction network in the pre-trained voiceprint model; and converting the extracted frame-level features into sentence-level features using a sentence-level feature extraction network in the pre-trained voiceprint model to obtain speaker voiceprint features corresponding to the speaker's voice data.

[0044] In this embodiment, the voice feature extraction is the first step of voiceprint recognition. The network extracts key features that can reflect the identity of the speaker from the original sound signal. For example, the spectral features of the sound (such as Mel Frequency Cepstral Coefficients MFCC), the fundamental frequency features, the energy features, etc. These features can capture the physical properties of the sound and provide basic data for subsequent feature processing. After obtaining the voice features, the frame-level feature extraction network further processes the features of each time frame. It uses technologies such as Convolutional Neural Network (CNN) or Recurrent Neural Network (RNN) to mine local patterns and capture time sequence information for each frame. For example, CNN can extract local feature patterns within the frame, and RNN can capture the time sequence dependence between frames. Further, although the frame-level features contain rich local information, in order to obtain voiceprint features that can represent the entire sentence, the frame-level features can be integrated and converted. The sentence-level feature extraction network uses pooling operations (such as global average pooling), fully connected layers, etc. to fuse the frame-level features into a fixed-dimensional feature vector that can represent the entire sentence. This feature vector is the voiceprint feature of the speaker, which integrates the sound information in the entire sentence, making the voiceprint feature more robust and representative. This feature can better cope with noise, pitch variation and other interference factors in the sound, improving the accuracy and stability of voiceprint recognition and providing a more reliable basis for subsequent determination of bullying suspects.

[0045] In the embodiments of the present application, before the voiceprint feature recognition of the speaker voice data by the pre-trained voiceprint model, the method further comprises: obtaining regional noise data corresponding to each bullying monitoring area, and adding noise with corresponding bullying monitoring area characteristics to the pure sound samples based on the regional noise data corresponding to each bullying monitoring area to obtain noise sound samples corresponding to each bullying monitoring area; based on the pure sound samples and the noise sound samples corresponding to each bullying monitoring area, voiceprint model training samples corresponding to each bullying monitoring area are constructed respectively; for each bullying monitoring area, the preliminary trained original voiceprint model is enhanced and trained based on the voiceprint model training samples corresponding to the bullying monitoring area to obtain the pre-trained voiceprint model corresponding to the bullying monitoring area.

[0046] Correspondingly, the voiceprint feature recognition of the speaker voice data by the pre-trained voiceprint model comprises: performing voiceprint feature recognition on the speaker voice data by the pre-trained voiceprint model corresponding to the target bullying monitoring area.

[0047] In this embodiment, by collecting the actual environmental noise data of each bullying monitoring area on campus (such as corridor echo, playground human voice, classroom equipment noise, etc.), the pure sound sample is superimposed with the noise of the corresponding area to generate a noise sound sample with regional characteristics. For example, add the background noise of the sports scene in the playground area, and add the specific noise such as desk and chair friction sound and projector fan sound in the classroom area. Traditional voiceprint model training usually uses a standard noise library, which is difficult to adapt to the complex and variable real environment on campus. The embodiments of the present application construct training samples by customizing regional noise, so that the model can adapt to the actual acoustic environment of the target area in advance, effectively reduce the interference of environmental noise on voiceprint recognition, and improve the robustness of the model in a specific scene. Based on the pure sound sample and the noise sound sample corresponding to the area, a dedicated training sample set is constructed for each bullying monitoring area. Each sample set contains mixed data of "pure voice + regional noise", covering various speech scenes such as different speakers, different speech rates, and different emotional states, so that through regional exclusive sample set training, the model can learn the noise mode and voice feature association rule specific to the region, realize the deep binding of "scene-voiceprint", avoid recognition bias caused by environmental differences, and improve the positioning accuracy of the model in the target area. Further, for each bullying monitoring area, the initial voiceprint model is enhanced and trained using its exclusive training samples. During the training process, the fusion features of the regional noise and the pure voice are mainly learned, and finally a pre-trained voiceprint model adapted to the corresponding region is generated. In application, the system automatically calls the corresponding pre-trained model for voiceprint recognition according to the target bullying monitoring area. Through regional enhancement training, the model can specifically strengthen the noise suppression ability and feature extraction ability of the target region. For example, the filter for wind noise and sports sound is enhanced in the playground area model, and the noise reduction processing for desk and chair friction sound is optimized in the classroom area model, finally realizing the adaptation of "one region one model" and improving the voiceprint recognition accuracy in complex environment. In the voiceprint feature recognition stage, the system automatically calls the pre-trained voiceprint model of the corresponding region according to the target bullying monitoring area, and performs feature extraction and matching on the speaker voice data. For example, if the bullying behavior occurs in the playground area, the playground exclusive model is used for voiceprint analysis, rather than the general model, to ensure that voiceprint recognition is always performed under the optimal model.

[0048] In the embodiments of the present application, the enhanced training of the initial voiceprint model based on the voiceprint model training samples corresponding to the bullying monitoring area includes:

[0049] The original voiceprint model is copied to obtain a copied voiceprint model;

[0050] extracting network in the original voiceprint model to obtain pure noise speech features; extracting frame-level features from the pure noise speech features through a frame-level feature extraction network in the original voiceprint model to obtain pure noise frame-level features; and converting the pure noise frame-level features into pure noise sentence-level features through a sentence-level feature extraction network in the original voiceprint model;

[0051] extracting network in the original voiceprint model to obtain pure noise speech features; extracting frame-level features from the pure noise speech features through a frame-level feature extraction network in the original voiceprint model to obtain pure noise frame-level features; and converting the pure noise frame-level features into pure noise sentence-level features through a sentence-level feature extraction network in the original voiceprint model;

[0052] calculating a first loss between the pure noise frame-level features and the noise frame-level features, calculating a second loss between the pure noise sentence-level features and the noise sentence-level features, and optimizing the copied voiceprint model based on the first loss and the second loss to achieve enhanced training of the preliminarily trained original voiceprint model.

[0053] In this embodiment, the preliminarily trained original voiceprint model is copied to generate an identical "copied voiceprint model" that retains the core structure and parameters of the original model as a basic framework for subsequent enhanced training, avoiding loss of training results caused by directly modifying the original model. The original voiceprint model is used to perform three-stage feature extraction on the pure sound samples: speech feature extraction: basic frequency spectrum, fundamental frequency, and other physical features are obtained through a speech feature extraction network; frame-level feature extraction: local time sequence patterns such as phoneme transition and intonation change are mined through a frame-level network; and sentence-level feature extraction: global features are integrated through a sentence-level network to generate a voiceprint feature vector representing pure speech. The copied voiceprint model is used to perform a feature extraction process completely symmetrical to that of the pure samples on the noise sound samples to generate speech features, frame-level features, and sentence-level features in a noise environment. Then, the feature differences between the pure and noise samples at the frame level (first loss) and the sentence level (second loss) are calculated: the first loss measures the distortion degree of the frame-level features under noise interference; and the second loss measures the semantic integrity loss of the sentence-level features in a noise environment.

[0054] Based on the double losses, the parameters of the copied voiceprint model are optimized, and finally the original voiceprint model is enhanced through model parameter back propagation.

[0055] In the embodiments of the present application, optionally, the optimization of the cloned voiceprint model based on the first loss and the second loss comprises: if the first loss is greater than the second loss, separately optimizing the frame-level feature extraction network in the cloned voiceprint model, otherwise, jointly optimizing the frame-level feature extraction network and the sentence-level feature extraction network in the cloned voiceprint model.

[0056] In this embodiment, by comparing the size relationship between the frame-level feature difference (first loss) and the sentence-level feature difference (second loss), the optimization strategy is dynamically selected. Specifically, when the first loss is greater than the second loss, it indicates that the distortion degree of the frame-level feature under noise interference is higher than that of the sentence-level. At this time, the frame-level feature extraction network is separately optimized to solve the noise resistance problem of local timing patterns (such as phoneme transition and intonation fluctuation), and to avoid interference of sentence-level optimization on frame-level. When the first loss is less than or equal to the second loss, it indicates that the semantic integrity loss of the sentence-level feature in the noise is more prominent. At this time, the frame-level and sentence-level feature extraction networks are jointly optimized to simultaneously improve the local detail retention capability and global semantic stability through end-to-end training, so as to ensure the consistency of the two-level features in the noise environment. Thus, the "problem-oriented" precise optimization is realized through loss value comparison. In addition, the dynamic switching of separate optimization and joint optimization in the embodiments of the present application avoids the risk of "over-adjustment" of the same optimization strategy for different levels of features. For example, if joint optimization is forcibly performed when the frame-level problem is prominent, it may cause the sentence-level feature to overfit the noise pattern; conversely, if only the frame-level is optimized when the sentence-level problem is prominent, it may not be able to solve the global semantic loss. This strategy ensures the stable performance of the model in the complex campus acoustic environment through situational processing. The noise patterns in different areas of the campus are significantly different (such as the sound of desk and chair rubbing in the classroom, the echo in the corridor, and the human voice in the playground), and the optimization strategy can dynamically adapt to these differences. For example, in the classroom area, if the frame-level feature is more seriously interfered by the projector fan sound, the frame-level network is preferentially optimized; in the playground area, if the sentence-level feature is more prominently affected by the shouting sound of the sports, joint optimization is started to avoid optimization conflict and over-adjustment, and to ensure the sustained stable performance of the model in the complex environment.

[0057] In the embodiments of the present application, after the anti-bullying alarm information is reported to the management terminal, the method further comprises: acquiring terminal location information corresponding to a security terminal worn by a security personnel, determining a target security terminal according to the distance between the terminal location information and the target bullying monitoring area; and sending the anti-bullying alarm information to the target security terminal, so that the target security terminal displays a patrol prompt information based on the anti-bullying alarm information, prompting the security personnel to go to the target bullying monitoring area for patrol.

[0058] In this embodiment, the terminal position data of the security personnel is collected in real time by GPS, Bluetooth beacon or campus Wi-Fi positioning system to form a dynamic position coordinate library. For example, the smart bracelet carried by the security personnel during patrol can report the latitude and longitude information every 2 seconds. Based on the geographic coordinates of the target bullying monitoring area (such as the northeast corner of the playground), the straight-line distance or path distance of all security terminals from the area is calculated. Through the 4G / 5G network or the campus special communication network, the anti-bullying alarm information containing the specific location, time and sound feature summary of the bullying is pushed to the target security terminal. The terminal device will prompt in three ways of vibration, voice broadcast and pop-up window, to ensure that the security personnel know the information in the first time and arrive at the scene before the bullying escalates, effectively preventing the situation from deteriorating.

[0059] In the embodiments of the present application, the method further comprises: acquiring monitoring image information of each monitoring device in the campus, and performing image recognition on the monitoring image information to identify bullying suspect images in the monitoring image information; determining a bullying monitoring area corresponding to a position of a monitoring device corresponding to the bullying suspect images as a prompt area; and outputting anti-bullying prompt information through an anti-bullying device corresponding to the prompt area to prompt against bullying behavior.

[0060] In this embodiment, the video stream data of each area is collected in real time through the monitoring camera network deployed in the campus. A deep learning-based image recognition algorithm (such as YOLOv7 target detection + 3D convolution behavior analysis) is used to identify bullying behavior features such as body conflict, encirclement and push in the video. For example, the system can detect features such as abnormal close proximity of more than two people, abnormal frequency of body contact, and painful facial expressions, and judge whether there is bullying suspicion through time series analysis. According to the monitoring device ID of the bullying suspect image identified, the monitoring device coordinates are mapped to the corresponding bullying monitoring area through the pre-constructed campus GIS geographic information system. The anti-bullying device (such as smart broadcast, electronic fence) corresponding to the prompt area is triggered through the Internet of Things protocol (such as MQTT) to output anti-bullying prompt information in real time. For example, the device will play the pre-recorded warning voice "abnormal behavior has been detected, please stop immediately", and at the same time, the sound and light alarm function of the electronic fence can be started to form a visual and auditory deterrent.

[0061] Further, the embodiments of the present application provide an anti-bullying alarm system based on AI voice recognition, as shown in Figure 3 The system includes an anti-bullying device and a control device, the anti-bullying device is deployed in at least one bullying monitoring area preset in the campus, and the anti-bullying device and the control device are used to implement the anti-bullying alarm method based on AI voice recognition described above. The corresponding description of the system can be referred to the corresponding description in the above method, which will not be described here.

[0062] The embodiment of the present application further provides a computer device, which can be a personal computer, a server, a network device, etc. The computer device comprises a bus, a processor, a memory and a communication interface, and can further comprise an input / output interface and a display device. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store location information. The network interface of the computer device is used to communicate with an external terminal through network connection. The computer program is executed by the processor to implement the steps in the method embodiments.

[0063] Those skilled in the art can understand that the structure of the computer device described above is only part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can comprise more or fewer components, or combine certain components, or have a different component arrangement.

[0064] In one embodiment, a computer readable storage medium is provided, which can be non-volatile or volatile, and stores a computer program. The computer program is executed by the processor to implement the steps in the method embodiments described above.

[0065] In one embodiment, a computer program product is provided, which comprises a computer program. The computer program is executed by the processor to implement the steps in the method embodiments described above.

[0066] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties.

[0067] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.

[0068] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist, it should be considered as the scope of the present application.

[0069] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A bullying alarm method based on AI voice recognition, characterized in that, The method includes: Acquire the voice data to be identified by anti-bullying devices deployed in at least one pre-set bullying monitoring area on campus; The voice data to be identified is subjected to speech recognition, and when any target voice data in the voice data to be identified contains preset bullying key information, anti-bullying alarm information is generated based on the target voice data and the anti-bullying device identifier corresponding to the target voice data. The anti-bullying alarm information is reported to the management terminal, so that the management terminal outputs anti-bullying prompt information based on the target bullying monitoring area corresponding to the anti-bullying device identifier and the target sound data; Receive anti-bullying prompts from the management terminal, and output the anti-bullying prompts via voice through the target anti-bullying device identified by the anti-bullying device identifier; The method further includes: When any target sound data in the sound data to be identified contains preset bullying key information, the target bullying monitoring area corresponding to the anti-bullying device identifier is determined according to the anti-bullying device identifier corresponding to the target sound data, and the surrounding monitoring device corresponding to the target bullying monitoring area is determined, and the monitoring data captured by the surrounding monitoring device is obtained. The monitoring data is subjected to facial recognition. Based on the facial recognition results, the person to be investigated corresponding to the target voice data is determined, and the individual voiceprint features corresponding to the person to be investigated are obtained. Obtain the contextual audio data corresponding to the target sound data, and perform speaker voice separation on the contextual audio data to obtain multiple speaker voice data corresponding to the contextual audio data; For any speaker's voice data, the speaker's voice data is processed by a speech feature extraction network in a pre-trained voiceprint model corresponding to the target bullying monitoring area to extract speech features; the speaker's voice features are then processed by a frame-level feature extraction network in the pre-trained voiceprint model to extract frame-level features; and the extracted frame-level features are then converted into sentence-level features by a sentence-level feature extraction network in the pre-trained voiceprint model to obtain the speaker's voiceprint features corresponding to the speaker's voice data. Based on the individual voiceprint features that match the identified speaker's voiceprint features, the bullying suspect is identified. The training methods for the pre-trained voiceprint model include: Obtain the regional noise data corresponding to each bullying monitoring area, and add noise with the characteristics of the corresponding bullying monitoring area to the clean sound samples based on the regional noise data corresponding to each bullying monitoring area to obtain the noise sound samples corresponding to each bullying monitoring area. Based on the clean sound samples and the noise sound samples corresponding to each bullying monitoring area, voiceprint model training samples corresponding to each bullying monitoring area are constructed respectively. For each bullying monitoring area, the original voiceprint model is replicated to obtain a replicated voiceprint model; the pure voice sample is processed by a speech feature extraction network in the original voiceprint model to obtain pure speech features; the pure speech features are then processed by a frame-level feature extraction network in the original voiceprint model to obtain pure frame-level features; and the pure frame-level features are converted into pure sentence-level features by a sentence-level feature extraction network in the original voiceprint model; the noisy voice sample is processed by a speech feature extraction network in the replicated voiceprint model to obtain noisy speech features; and the replicated voiceprint model is then processed by a frame-level feature extraction network in the original voiceprint model to obtain pure frame-level features; the replicated voiceprint model is then processed by a frame-level feature extraction network in the original voiceprint model to obtain pure frame-level features; the replicated voiceprint model is then processed by a frame-level feature extraction network in the original voiceprint model to obtain pure frame-level features; the replicated voice sample ... The frame-level feature extraction network in the voiceprint model extracts frame-level features from the noisy speech features to obtain noisy frame-level features; and the sentence-level feature extraction network in the replicated voiceprint model converts the noisy frame-level features into noisy sentence-level features; a first loss is calculated between the clean frame-level features and the noisy frame-level features, and a second loss is calculated between the clean sentence-level features and the noisy sentence-level features. Based on the first loss and the second loss, the replicated voiceprint model is optimized to enhance the training of the pre-trained original voiceprint model to obtain a pre-trained voiceprint model corresponding to the bullying monitoring area.

2. The anti-bullying alarm method based on AI voice recognition according to claim 1, characterized in that, The process of identifying a suspected bully based on individual voiceprint features matched with the identified speaker's voiceprint features includes: Calculate the similarity between the speaker's voiceprint features and the voiceprint features of each individual, and determine the target individual's voiceprint features with the highest similarity. If the voiceprint feature of the target individual is greater than or equal to a preset similarity threshold, then the person to be investigated corresponding to the voiceprint feature of the target individual is identified as a bullying suspect. If the voiceprint feature of the target individual is less than a preset similarity threshold, then based on the preset number of supplementary screenings, supplementary monitoring devices around the target bullying monitoring area are determined, the monitoring data captured by the supplementary monitoring devices are acquired as new monitoring data, and the process returns to the step of performing facial recognition on the monitoring data to re-identify the bullying suspect.

3. The anti-bullying alarm method based on AI voice recognition according to claim 1, characterized in that, The optimization of the replicated voiceprint model based on the first loss and the second loss includes: If the first loss is greater than the second loss, then the frame-level feature extraction network in the replicated voiceprint model is optimized separately; otherwise, the frame-level feature extraction network and the sentence-level feature extraction network in the replicated voiceprint model are jointly optimized.

4. The anti-bullying alarm method based on AI voice recognition according to any one of claims 1 to 3, characterized in that, After reporting the anti-bullying alarm information to the management terminal, the method further includes: Obtain the terminal location information corresponding to the security terminal worn by the security personnel, and determine the target security terminal based on the distance between the terminal location information and the target bullying monitoring area; The anti-bullying alarm information is sent to the target security terminal, so that the target security terminal displays inspection prompt information based on the anti-bullying alarm information, prompting security personnel to go to the target bullying monitoring area for inspection.

5. The anti-bullying alarm method based on AI voice recognition according to any one of claims 1 to 3, characterized in that, The method further includes: Acquire monitoring image information from various monitoring devices on campus, and perform image recognition on the monitoring image information to identify images suspected of bullying in the monitoring image information; Based on the location of the monitoring device corresponding to the suspected bullying image, the bullying monitoring area corresponding to the location of the monitoring device is determined as the alert area; The anti-bullying device corresponding to the prompt area outputs anti-bullying prompt information to remind users to refrain from bullying behavior.

6. An anti-bullying alarm system based on AI voice recognition, characterized in that, The system includes anti-bullying devices and control devices. The anti-bullying devices are deployed in at least one pre-set bullying monitoring area on campus. The anti-bullying devices and the control devices are used to implement the anti-bullying alarm method based on AI voice recognition as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Intelligent campus student behavior analysis system based on artificial intelligence

    CN117237155A

  • Intelligent voice alarm system and method for preventing spoofing

    CN118747933A

  • Self-supervised learning model-based speaker identification method in real sound field environment

    CN119207431A