Campus bullying prevention and control intelligent equipment based on voiceprint recognition and real-time feedback

By using high-sensitivity microphones and adaptive noise cancellation algorithms in campus smart devices for sound acquisition, combined with voiceprint recognition and swear word recognition technology, real-time recognition and feedback on campus language violence is achieved, and the problem of difficulty in effectively preventing and controlling campus language violence in the existing technology is solved, and the efficiency of campus safety management is improved.

CN120220690APending Publication Date: 2025-06-27曾晓明
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510151478.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify and prevent and control linguistic violence on campus, especially in complex campus environments, and lacks real-time feedback and data recording capabilities.

Method used

A high-sensitivity MEMS microphone is used to combine adaptive noise cancellation algorithm to collect sounds; a librosa library is used to extract MFCC coefficients and combine Z-score standardization to perform voiceprint recognition; a multilingual vocabulary library and dialect voice transcription and matching mechanism is established to identify swear words; real-time feedback is provided through small speakers and adaptive volume adjustment technology; data recording and transmission is used by ESP8266 Wi-Fi module and Flask server.

Benefits of technology

It realizes accurate identification of linguistic violence in complex campus environments, provides real-time feedback and detailed data records, and improves the efficiency and effectiveness of campus safety management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The system is an intelligent device for preventing and treating campus speech bullying and dirty speech scenes, sound is collected through a high-sensitivity directional MEMS microphone, a voiceprint recognition module extracts an MFCC coefficient and differential characteristics by using a librosa library, after Z-score standardization, a pre-stored voiceprint model is matched through a DTW algorithm with path constraint, and the MFCC coefficient and the differential characteristics are subjected to Z-score standardization. The integrated dirty speech recognition module constructs a dynamic multi-language lexicon (including dialects), and adopts an optimized KMP algorithm and semantic analysis to recognize improper speech. And when a dirty call is detected, the feedback module plays a warning sound through the self-adaptive volume loudspeaker and sends an alarm to the server. And the data recording module is used for locally storing data through the SD card and encrypting and transmitting the data through the Wi-Fi module. The one-key alarm physical key supports long-press triggering, pushes notifications containing positions and time and carries out voice confirmation. The equipment adopts a low-power-consumption design, the weight is less than or equal to 80g, ergonomic wearing modes such as a hanging wearing mode and a brooch type are provided, and endurance and comfort are both considered.
Need to check novelty before this filing date? Find Prior Art

Description

I. Summary of the Invention

[0002] 1. Object of the Invention

[0003] The present invention aims to provide a multi-functional intelligent device, focusing on the prevention and control of campus bullying. Through advanced voiceprint recognition technology, it can accurately locate the speaker, timely identify and feedback campus language violence behavior, and record relevant data in detail. The device has a strong ability to recognize local dialect swear words, and at the same time supports parents to track the location of their children in real time and the one-key alarm function for students in case of danger. In addition, the device adopts a simplified design, has a long battery life, is easy to wear and carry, comprehensively guarantees the safety of students at school, effectively prevents and reduces the occurrence of campus bullying incidents, and is more easily accepted by schools, students and parents.

[0004] II. Technical Solution

[0005] 2.1 Sound Acquisition

[0006] A highly sensitive directional MEMS microphone (model INMP441) is selected, with a sensitivity of up to -38dBV / Pa and a frequency response range of 20Hz - 20kHz, which can effectively collect sounds within a radius of 1 - 2 meters. An adaptive noise cancellation algorithm is built-in, which dynamically adjusts the filter parameters according to the statistical characteristics of the environmental noise. Whether in a noisy playground, a bustling cafeteria, or a quiet classroom and other complex campus environments, it can effectively filter out background noise and ensure clear collection of the target voice. The microphone is connected to the main control board through the I2S interface to ensure high-speed and stable transmission of audio data. At the same time, a preamplifier is added in front of the microphone to further improve the sensitivity and signal-to-noise ratio of sound acquisition.

[0007] 2.2 Voiceprint Recognition

[0008] · Feature Extraction: Use the librosa library to extract MFCC coefficients, set the number of Mel filters to 40, the frame length to 256 sampling points, and the frame shift to 128 sampling points. To enhance the robustness of the features, after the MFCC coefficients are extracted, first-order and second-order difference coefficients are added to form a feature vector containing static and dynamic information. Then, the Z-score normalization method is used to normalize the extracted coefficients, converting the feature coefficients into standard normal distribution data with a mean of 0 and a standard deviation of 1, significantly improving the accuracy and stability of voiceprint recognition. Establish a voiceprint model training mechanism, regularly collect new voiceprint data, optimize and update the voiceprint model, and improve the recognition accuracy.

[0009] · Model comparison: Real-time comparison is carried out with the pre-stored voiceprint model through the DTW algorithm. The DTW algorithm uses the Euclidean distance as the measurement standard. During the calculation process, path constraints are introduced, such as setting a maximum slope limit, to prevent unreasonable matching paths and improve the accuracy and stability of the matching. At the same time, the calculation process of the DTW algorithm is optimized to reduce the amount of calculation, speed up the recognition speed, and ensure that a large amount of voiceprint data can be compared in a short time.

[0010] 2.3 Swear word recognition

[0011] · Construction of a multilingual vocabulary library: In addition to the general swear word vocabulary library, a local dialect swear word vocabulary library is specifically established. It is stored using the data structure of a hash table to improve the speed of vocabulary lookup. By widely collecting dialect materials from different regions, cooperating with linguistics experts, and using web crawlers to obtain dialect swear words from local social platforms, forums, etc. The vocabulary library is updated dynamically, and the update frequency is set to once a week. For schools in different regions, according to the local actual situation, the content of the vocabulary library can be manually added or adjusted to ensure the pertinence and accuracy of the vocabulary library.

[0012] · Dialect voice transcription and matching: Combine an advanced speech recognition API with extended dialect recognition capabilities to convert the collected voice into text. When performing local matching, the optimized KMP algorithm is used. Considering the particularity of dialect vocabulary, such as homophony, liaison, etc., when constructing the longest prefix-suffix array, an improved two-pointer method is adopted and combined with dialect voice rules for judgment to reduce the computational time complexity and improve the efficiency of dialect swear word recognition. At the same time, semantic analysis technology is introduced to perform semantic judgment on the text suspected of being a swear word, reducing the misjudgment rate. A swear word recognition feedback mechanism is established to optimize and adjust the swear word recognition algorithm according to the actual usage situation, especially continuously improving the recognition effect of dialect swear words.

[0013] 2.4 Real-time feedback

[0014] When a swear word (including dialect swear words) is detected, a warning voice is played through a small speaker with a power of 0.5W and a volume range of 80 - 100dB. Adaptive volume adjustment technology is adopted to automatically adjust the volume according to the ambient noise level. In a quiet library environment, the volume is automatically reduced to avoid excessive interference; in a noisy corridor during class breaks, the volume is automatically increased to ensure that the warning information is clearly conveyed. At the same time, a warning notice containing multi-dimensional information such as the speaker's identity, time, location, and swear word content (including dialect annotation) is sent to the background server for school administrators to handle in a timely manner.

[0015] 2.5 Data recording

[0016] · Local storage: The relevant data is temporarily stored locally through a high-speed SD card with a capacity of 16GB and read / write speeds of 90MB / s and 80MB / s respectively. A cyclic storage strategy is adopted. When the storage space of the SD card is insufficient, the earliest data is automatically overwritten. A data indexing mechanism is established to facilitate the quick search and reading of specific data.

[0017] · Remote transmission: The ESP8266 Wi-Fi module is used to send JSON-formatted data to the Flask server via the HTTP protocol. During the transmission process, SSL / TLS encryption technology is adopted to prevent data from being stolen or tampered with. A data compression algorithm, such as the gzip algorithm, is used to compress the transmitted data to reduce the transmission traffic and time. Multithreading technology is used to achieve asynchronous data transmission to avoid affecting other functions of the device during the data transmission process. A data transmission monitoring mechanism is established to monitor the data transmission status in real time, and errors during the transmission process are processed and retransmitted in a timely manner to ensure the reliability of data transmission.

[0018] 2.6 Location system

[0019] An integrated high-precision GPS module (such as Ublox NEO-M8N) is combined with base station positioning and Wi-Fi positioning technologies to achieve multi-mode fusion positioning. In outdoor environments, the GPS module, relying on its high sensitivity and fast positioning ability, quickly and accurately obtains the device's location information. In indoor environments, base station positioning and Wi-Fi positioning technologies are used for assisted positioning. Using the base station signals and Wi-Fi hotspot information around the device, the approximate location of the device is calculated through the triangulation algorithm. The location information is transmitted to the background server in real time through the ESP8266 Wi-Fi module. Parents can view the location information of their children in real time through a specially developed mobile app. The app interface is simple and intuitive, and easy to operate. The Kalman filter algorithm is used to fuse the location data in different positioning modes to improve the accuracy and stability of positioning. The positioning module is equipped with a backup power supply to ensure that positioning services can still be provided for a certain period of time in case of main power failure.

[0020] 2.7 One-key alarm system

[0021] Set an independent physical button on the device as the one-key alarm button, and adopt a trigger method of long pressing for 3 seconds. When a student encounters danger, pressing the button can trigger the alarm mechanism. The device immediately sends an alarm signal to the background server, and at the same time pushes an alarm notification to the mobile apps of parents and the school security department. The notification content includes the precise location information of the student, the alarm time, etc. After triggering the alarm, the device gives a voice prompt for confirmation to prevent false alarms and ensure the authenticity and effectiveness of the alarm. Set the alarm confirmation and processing functions on the mobile app side to facilitate the timely response of parents and the school security department. Establish an alarm record and statistics mechanism, analyze the alarm data, and provide data support for campus security management.

[0022] 2.8 Simplified Design and Battery Life Optimization

[0023] · Function Simplification: Remove complex functions unrelated to campus bullying prevention and focus on the in-depth optimization of core functions. Only retain functions such as voiceprint recognition, swear word recognition, location tracking, one-key alarm, and necessary data recording and feedback functions, simplify the operation process, and reduce the device power consumption.

[0024] · Hardware Optimization: Select low-power chips and electronic components. For example, the main control board uses a low-power single-chip microcomputer model. While meeting the device performance requirements, minimize energy consumption to the greatest extent. Optimize the circuit design, reduce unnecessary circuit losses, and improve the power utilization efficiency.

[0025] · Power Management: Adopt a rechargeable lithium battery with a capacity of 800 mAh or more, and be equipped with an efficient power management circuit to achieve intelligent charging, discharging management of the battery, as well as power distribution and control of each module of the device. Adopt intelligent sleep technology. When the device has no operation or is in a low-power state for a period of time, it automatically enters the sleep mode to reduce power consumption.

[0026] 2.9 Portable Design

[0027] The overall device is made of lightweight materials, such as high-strength and lightweight plastics or alloy materials, and the weight of the device is controlled to not exceed 80 grams. In terms of appearance design, follow the ergonomic principle, and adopt a hanging, brooch-style or arc-shaped design that fits the wrist to ensure that students feel comfortable during long-term wearing and do not hinder daily study and exercise. III. Specific Implementation Modes

[0028] 3.1 Hardware Implementation

[0029] · Main control board: Select a single-chip microcomputer that meets the operating requirements of the device and has low power consumption as the main control board. It has sufficient data processing capabilities and storage capacity. The working voltage is set to 3.3V, and the working frequency is dynamically adjusted according to task requirements to reduce power consumption. The main control board is equipped with rich digital and analog input / output interfaces, facilitating connection and communication with other modules. A power filter circuit and a reset circuit are added to the main control board to effectively improve system stability.

[0030] · Microphone: A MEMS microphone (model INMP441) is used for sound collection. Its characteristics of low power consumption and high sensitivity are suitable for the long-term stable operation of this device. The microphone is connected to the main control board through the I2S interface to ensure high-speed and stable transmission of audio data. A preamplifier is added in front of the microphone to further improve the sensitivity and signal-to-noise ratio of sound collection.

[0031] · Wi-Fi module: The ESP8266 Wi-Fi module is used to achieve data transmission. This module supports the 802.11b / g / n protocol, and the maximum transmission rate can reach 72.2Mbps. An external high-gain antenna is added to enhance the Wi-Fi signal strength and stability, ensuring stable connection to the network in the campus environment. The driver program of the Wi-Fi module is optimized to improve the efficiency and reliability of data transmission, while reducing the power consumption of the module.

[0032] · Speaker: Select a small speaker with a power of 0.5W. The volume is precisely controlled through PWM (Pulse Width Modulation) technology to achieve clear playback of warning voices. An audio power amplifier is used to connect the speaker to the main control board to ensure the normal operation of the speaker. The sound quality of the speaker is optimized to make the warning voice more clearly distinguishable.

[0033] · Power management: A rechargeable lithium battery with a capacity of 800mAh is used as the power source to meet the long-term battery life requirements of the device. An efficient power management circuit is equipped to achieve charging and discharging management of the battery, as well as power distribution and control of each module of the device. Intelligent charging algorithms, such as pulse charging algorithms, are used to extend the battery life and improve the charging efficiency.

[0034] · Positioning module: Integrate a high-precision GPS module (Ublox NEO-M8N), equipped with relevant antennas and radio frequency circuits to ensure stable reception of GPS signals. Use the ESP8266 Wi-Fi module to obtain information about surrounding base stations and Wi-Fi hotspots, and perform data processing and position calculation through the main control board. A backup power source is added to the positioning module to ensure that positioning services can still be provided for a certain period of time in case of main power failure.

[0035] · Alarm Button: An independent physical button is used and connected to the interrupt pin of the main control board. The button has good tactile feedback and durability, and its surface is designed with anti-slip features to ensure that students can trigger the alarm conveniently and quickly in case of an emergency. An indicator light is added around the button. When the button is pressed, the indicator light flashes to indicate that the alarm has been triggered.

[0036] 3.2 Software Implementation

[0037] · Voiceprint Feature Extraction: The librosa library is used to extract MFCC coefficients and their differential coefficients. The extracted coefficients are normalized using the Z-score normalization method to transform the feature coefficients into data with a standard normal distribution. A voiceprint model training mechanism is established to regularly collect new voiceprint data, optimize and update the voiceprint model, and improve the recognition accuracy.

[0038] · Swear Word Matching: The optimized KMP algorithm is combined with a dynamically updated multilingual swear word vocabulary library for swear word recognition. During the software implementation process, reasonable thread management is carried out for the update of the vocabulary library, the operation of the matching algorithm, and the interaction with the speech recognition API. Multithreading technology is used to improve the system operation efficiency. A swear word recognition feedback mechanism is established to optimize and adjust the swear word recognition algorithm according to the actual usage situation, especially to continuously improve the recognition effect for dialect swear words.

[0039] · Data Transmission: The collected and processed data is encapsulated in JSON format and sent to the Flask server via the HTTP protocol. During the data transmission process, multithreading technology is used to achieve asynchronous data transmission, avoiding the impact on other functions of the device during data transmission. A data transmission monitoring mechanism is established to monitor the data transmission status in real-time, handle errors during the transmission process in a timely manner and retransmit, ensuring the reliability of data transmission.

[0040] · Location Algorithm: In the GPS positioning mode, the satellite signal data obtained by the GPS module is used to calculate the longitude and latitude coordinates of the device through the built-in location algorithm. In the base station positioning and Wi-Fi positioning modes, based on the obtained base station signal strength and Wi-Fi hotspot MAC address and other information, the position of the device is calculated through the triangulation algorithm. The position data in different positioning modes is fused and processed, and the Kalman filter algorithm is used to improve the accuracy and stability of positioning.

[0041] · Alarm mechanism: When the one-key alarm button is pressed, it triggers the interrupt service program. The device immediately sends an alarm signal to the background server and pushes notifications through the mobile app. Before sending the alarm notification, the device accurately collects and packages the current location information to ensure that the recipient can accurately obtain the student's location. Set up an alarm confirmation and processing function on the mobile app side to facilitate timely response from parents and the school security department. Establish an alarm record and statistics mechanism to analyze the alarm data and provide data support for campus security management.

[0042] · System optimization: Conduct customized development on the device's operating system, streamline the system kernel, remove unnecessary system services and functions, reduce system resource occupancy, and improve system operation efficiency. Optimize the execution efficiency of software algorithms, reduce the algorithm running time, and lower the device power consumption. Description of the drawings: Figure 1 It is a creative design drawing of the wearable appearance of the device. It can be simply hung around the student's neck by the student. The design is simple and lightweight, without interfering with daily study and exercise. Figure 2 It is a creative design drawing of the brooch clip-on appearance of the device. It can be clipped on the collar by the student. The design is simple and lightweight, without interfering with daily study and exercise. Figure 3 It is a structural diagram of the intelligent device components for preventing campus bullying based on voiceprint recognition and real-time feedback. Figure 4 It is a schematic diagram of the circuit design of the intelligent device for preventing campus bullying based on voiceprint recognition and real-time feedback. Figure 5 It is a voiceprint recognition flowchart of the device. Figure 6 It is a system architecture diagram of the device. Figure 7 It is a flowchart of the positioning system of the device. Figure 8 It is a flowchart of the one-key alarm system of the device.

Claims

1. A campus bullying prevention device, characterized in that: It includes a sound collection module (MEMS microphone), a voiceprint recognition module (MFCC feature extraction and DTW comparison), a swear word recognition module (multi-language vocabulary library + speech recognition API), a feedback module (speaker), a data recording module (SD card + Wi-Fi transmission), a positioning module (GPS + base station + Wi-Fi positioning), and a one-button alarm module. The device adopts a simplified design, has long battery life, and is easy to wear and carry.

2. The device according to claim 1, characterized in that The voiceprint recognition module realizes speaker identity positioning through the pre-stored student voiceprint model. When extracting MFCC coefficients, the number of Mel filters is set to 40, the frame length is 256 sampling points, the frame shift is 128 sampling points, and the first-order and second-order differential coefficients are added. The DTW algorithm adopts the Euclidean distance metric and introduces path constraints, and the extracted coefficients are subjected to Z-score normalization.

3. The device according to claim 1, characterized in that The swear word recognition module combines a dynamically updated multilingual swear word vocabulary library (including a general swear word library and a local dialect swear word library) with a speech recognition API with dialect recognition function expansion. The vocabulary library is stored in a hash table and updated once a week. The latest vocabulary is obtained through multiple channels. The local matching adopts the optimized KMP algorithm and is judged in combination with dialect speech rules. Semantic analysis technology is introduced to reduce the misjudgment rate.

4. The device according to claim 1, characterized in that The sound collection module adopts a sensitivity of - 38dBV / Pa, directional MEMS microphone with a frequency response range of 20Hz-20kHz (model INMP441), filters background noise through adaptive noise cancellation algorithm, and connects to the main control board through I2S interface, and adds a preamplifier at the front end of the microphone.

5. The device according to claim 1, characterized in that When swearing is detected, the feedback module plays a warning voice through a small speaker with a power of 0.5W and a volume range of 80-100dB, adopts adaptive volume adjustment technology, and sends a warning notification containing multi-dimensional information such as the speaker's identity, time, location, swearing content (including dialect annotation) to the background server.

6. The device according to claim 1, characterized in that The data recording module stores data locally via a 16GB SD card with a read and write speed of 90MB / s and 80MB / s respectively, adopts a circular storage strategy and establishes a data index mechanism, and transmits JSON format data to the server via an ESP8266 Wi-Fi module that supports 802.11b / g / n protocols and has a transmission rate of up to 72.2Mbps. The transmission process uses SSL / TLS Encryption and gzip data compression algorithms are used, and multi-threading is used to achieve asynchronous transmission, and a data transmission monitoring mechanism is established.

7. The device according to claim 1, characterized in that The positioning module integrates a high-precision GPS module (such as Ublox NEO-M8N), combines base station positioning and Wi-Fi positioning technology, obtains device location information in real time through a multi-mode fusion positioning algorithm, and transmits it to the backend server through the Wi-Fi module. The Kalman filter algorithm is used to fuse the location data under different positioning modes to improve the accuracy and stability of positioning; the positioning module is equipped with a backup power supply, which can continue to provide positioning services for a certain period of time when the main power supply fails.

8. The device according to claim 1, characterized in that The one-button alarm module is equipped with an independent physical button, which can be pressed for 3 seconds to trigger the alarm mechanism. After the device triggers the alarm, it immediately sends an alarm signal to the background server, and pushes an alarm notification containing precise location information, alarm time, etc. to the mobile app of parents and school security departments; at the same time, the device performs voice prompt confirmation to prevent false alarms. The mobile app is equipped with alarm confirmation and processing functions to facilitate timely response by relevant personnel; establish an alarm record and statistical mechanism, analyze the alarm data, and provide data support for campus safety management.

9. The device according to claim 1, characterized in that The simplified design is reflected in retaining only the core functions related to the prevention and treatment of school bullying, removing other irrelevant complex functions, simplifying the operating process, and reducing the power consumption of the equipment. By streamlining the system architecture and reducing unnecessary software running processes, device operation becomes more convenient and smoother, while reducing the usage of hardware resources and further reducing overall energy consumption.

10. The device according to claim 1, characterized in that The long battery life is achieved by selecting low-power chips and electronic components, optimizing circuit design, using rechargeable lithium batteries with a capacity of 800mAh or more, equipping with efficient power management circuits and intelligent sleep technology. Under the condition of meeting the normal operation of the equipment, the battery life of the equipment is extended and the charging frequency is reduced. The power management circuit has an intelligent adjustment function, which can dynamically allocate power according to the real-time working status of the equipment, ensuring that each module operates efficiently while saving power to the maximum extent; the intelligent sleep technology can accurately identify the low-power state of the equipment, quickly respond to enter the sleep mode, and quickly wake up when there is an operation demand, without affecting the normal use of the equipment.

11. The device according to claim 1, characterized in that Easy to wear and carry means that the entire device is made of lightweight materials and weighs no more than 80 grams. In terms of appearance design, it follows the principles of ergonomics and adopts a hanging, brooch or wrist-fitting arc design to ensure that students feel comfortable when wearing it for a long time and does not interfere with daily study and exercise.

Citation Information

Cited By

  • A bathroom anti-bullying alarm and evidence collection method based on multi-modal triggering

    CN122531161A