Intelligent toilet management system with alarm function

The intelligent toilet management system, which combines voice acquisition and recognition technology with edge computing, solves the problem of the difficulty in timely detection of bullying incidents in public toilets, achieves efficient and accurate alarm response, and improves the intelligence and safety of toilet management.

CN120977307APending Publication Date: 2025-11-18SHENZHEN LIONKING SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511162681.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Public restrooms, due to their enclosed nature and low foot traffic, are prone to becoming venues for bullying incidents. Traditional management methods struggle to detect and respond to abnormal situations in a timely manner, leading to delayed rescue efforts.

Method used

The system employs a voice acquisition unit and a voice recognition alarm unit working in tandem, combined with an edge computing unit. Through voice recognition models and acoustic scene analysis, it monitors voice information in the toilet in real time. It utilizes microphone arrays and edge computing to perform multi-dimensional acoustic analysis, quantifies the probability of bullying events, and ensures alarm accuracy by resetting reliability thresholds twice.

Benefits of technology

It enables timely and accurate alarms for emergencies in toilets, reduces false alarm and missed alarm rates, improves the intelligence and security of toilet management, and ensures the safety of users in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977307A_ABST
    Figure CN120977307A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent toilet management system with an alarm function, and relates to the field of toilet management. The system comprises a management control center, a voice acquisition unit, a voice recognition alarm unit and an edge calculation unit. The voice acquisition unit is arranged in a target toilet and is used for acquiring toilet voice information; the voice recognition alarm unit is connected to the voice acquisition unit and the management control center, recognizes voice information through a first voice recognition model, and sends alarm information to the management control center when a target keyword is recognized and the first confidence of the target keyword is greater than a preset threshold; when the first confidence coefficient does not reach the standard, the edge calculation unit receives the voice information and identifies the voice information again through a second voice identification model, and if the second confidence coefficient reaches the standard, the edge calculation unit sends alarm information; the first voice recognition model is a student model obtained through a knowledge distillation technology based on the second voice recognition model. According to the invention, the problems of lagging alarm response and untimely countermeasures in toilet management can be relieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of toilet management technology, and in particular to an intelligent toilet management system with an alarm function. Background Technology

[0002] Public restrooms, due to their specific purpose, are generally designed as relatively secluded and enclosed spaces, creating opportunities for bullying. Bullies often choose places like restrooms to harass and harm their victims because of the relatively low foot traffic and the lack of real-time monitoring and management.

[0003] At the same time, due to the relatively enclosed spatial layout of toilets, it is difficult for outsiders to detect any abnormalities inside in a timely manner, making it difficult for victims to receive timely rescue and assistance.

[0004] Traditional toilet management relies mainly on manual inspections, which is inefficient and makes it difficult to detect abnormal events in a timely manner. In particular, when bullying occurs in toilets, traditional management methods are slow to respond and cannot take swift action. Summary of the Invention

[0005] To address the aforementioned technical problems and deficiencies, the purpose of this invention is to provide an intelligent toilet management system with an alarm function, which can alleviate the problems of delayed alarm response and untimely countermeasures in toilet management.

[0006] To achieve the above objectives, this invention provides an intelligent toilet management system with an alarm function, comprising a management control center, a voice acquisition unit, a voice recognition alarm unit, and an edge computing unit. The voice acquisition unit, installed inside the target toilet, is used to acquire toilet voice information. The voice recognition alarm unit, connected to both the voice acquisition unit and the management control center, is used to recognize the toilet voice information using a preset first voice recognition model. When a target keyword is recognized and its first confidence level is greater than a preset first confidence threshold, a bullying event is determined to have occurred in the target toilet, and an alarm message is sent to the management control center. The target keyword includes "call for help" or "bullying." The edge computing unit, connected to both the voice recognition alarm unit and the management control center, is used to receive the toilet voice information transmitted by the voice recognition alarm unit when the first confidence level is less than or equal to the first confidence threshold. It then processes the toilet voice information using a preset second voice recognition model to obtain a second confidence level for the target keyword. When the second confidence level is greater than a preset second confidence threshold, a bullying event is determined to have occurred in the target toilet, and an alarm message is sent to the management control center. The first voice recognition model is a student model obtained based on the second voice recognition model using knowledge distillation technology.

[0007] This invention achieves real-time monitoring and recognition of voice information in toilets through the collaborative work of a voice acquisition unit and a voice recognition alarm unit. It can promptly and accurately send alarm information to the management and control center when bullying, calls for help, or other emergencies are detected. At the same time, it uses an edge computing unit and a second voice recognition model to perform secondary recognition on low-confidence voice information, improving the accuracy and reliability of the alarm. Furthermore, it optimizes the first voice recognition model through knowledge distillation technology, which improves operational efficiency while maintaining recognition accuracy, reduces the performance requirements of front-end equipment, facilitates deployment, and effectively solves the problems of delayed alarm response and untimely countermeasures in toilet management.

[0008] Optionally, in some embodiments, the voice acquisition unit includes a microphone array for acquiring multi-directional audio signals within the target toilet and transmitting the raw multi-channel audio signals to the edge computing unit; the edge computing unit is further configured to: construct a sound source direction energy distribution map based on the multi-channel audio signals acquired by the microphone array, the sound source direction energy distribution map being used to characterize the intensity distribution of sound energy at different directional angles, as well as the dominance of the sound source direction and the concentration index of the sound source direction; and perform time-domain analysis on the total energy of the raw multi-channel audio signals to calculate the instantaneous wave of the energy. The frequency of dynamic and energy bursts is used to quantify the intensity of the sound; frequency domain analysis is performed on the original multi-channel audio signal to calculate the proportion of high-frequency energy and the clarity of the speech harmonic structure. If the clarity of the harmonic structure decreases, it indicates that multiple people are speaking at the same time or the sound is noisy and chaotic; based on the concentration index of the sound source direction, the instantaneous fluctuation rate of the energy, the frequency of the energy bursts, the proportion of high-frequency energy, and the clarity of the speech harmonic structure, the probability of a bullying event is calculated; when the probability of a bullying event exceeds a preset probability threshold, an alarm message is sent to the management and control center.

[0009] The technical solution described above provides the system with an independent alarm capability based on acoustic scene analysis, independent of specific keywords. Utilizing a microphone array and edge computing unit, the acoustic environment within the restroom is comprehensively analyzed from multiple dimensions, including spatial distribution of sound, energy intensity, and frequency characteristics. By constructing a sound source radiation map and analyzing energy fluctuations and frequency domain characteristics, the system can quantitatively determine the presence of typical bullying behavior patterns such as multiple people surrounding someone, heated arguments, or high-pitched screams. This method effectively identifies bullying incidents that do not use explicit distress signals, serving as a powerful supplement to keyword recognition and significantly improving the system's detection coverage and accuracy in complex and covert bullying scenarios.

[0010] Optionally, in some embodiments, the edge computing unit is further specifically used to: calculate a basic risk score based on the instantaneous fluctuation rate of the energy, the frequency of the energy burst events, and the proportion of high-frequency energy; calculate a scene confidence score based on the concentration index of the sound source direction and the clarity of the speech harmonic structure; and calculate the probability of the bullying event based on the basic risk score and the scene confidence score.

[0011] The technical solution described in the above embodiments optimizes and refines the probability calculation method for bullying incidents logically, making its judgment process more structured and robust. The complex acoustic analysis is decomposed into two core evaluation dimensions: a "basic risk score" directly reflecting the intensity of the incident, and a "scenario confidence score" describing whether the environment matches a multi-person conflict scenario. By combining these two for the final probability calculation, the system can achieve more intelligent decision-making, such as using scenario confidence to adjust the weight of the basic risk. This step-by-step calculation method effectively avoids misjudgments caused by single-feature anomalies, enabling alarm decisions to capture both the key dynamics of the incident and the static corroboration of the environment, thereby improving the accuracy of the final alarm.

[0012] Optionally, in some embodiments, an environmental acoustic baseline learning module is provided on the edge computing unit to periodically perform statistical analysis on the background voice information acquired by the voice acquisition unit to generate a dynamic acoustic baseline characterizing the normal acoustic environment of the target toilet at different times; the voice recognition alarm unit is also used to call the corresponding dynamic acoustic baseline according to the current real-time time and adaptively adjust the first confidence threshold; wherein, when the environmental background noise characterized by the dynamic acoustic baseline is low, the first confidence threshold is reduced to improve detection sensitivity.

[0013] The technical solution described above enables the system to adapt to different environments, addressing the detection challenges posed by significant variations in background noise levels across different toilets and time periods. By establishing an environmental acoustic baseline learning module, the system can autonomously learn and "remember" the normal sound conditions of a specific toilet at various times. Based on this dynamic baseline, the voice recognition alarm unit can intelligently adjust its recognition sensitivity threshold. During quiet periods such as nighttime, the system automatically lowers the threshold to capture weak distress signals; during noisy periods such as peak hours, it maintains a higher threshold to prevent false alarms. This allows the system to maximize detection sensitivity while ensuring a minimum false alarm rate, achieving intelligent and precise operation.

[0014] Optionally, in some embodiments, the voice recognition alarm unit is further configured to: extract the energy change characteristics of the toilet voice information in the time domain; determine whether the energy change characteristics match a preset voice energy reference characteristic; if so, transmit the toilet voice information to the edge computing unit.

[0015] The technical solution described in the above embodiments sets up a highly efficient computing resource filtering barrier for the system, thereby reducing system power consumption and improving response efficiency. Before performing complex speech recognition, the front-end speech recognition alarm unit first performs a simplified energy feature analysis to quickly determine whether the sound possesses "human-like" dynamic characteristics. Only when the energy change pattern of the sound conforms to the rules of human vocalization will the relevant audio data be transmitted to the back-end edge computing unit for in-depth analysis. This pre-screening mechanism effectively filters out a large amount of meaningless background noise (such as continuous wind or flushing sounds), avoiding waste of back-end computing resources and ensuring that the system concentrates its computing power on the most valuable audio segments.

[0016] Optionally, in some embodiments, the voice recognition alarm unit is further configured to: analyze the real-time collected toilet voice information to identify whether it simultaneously contains human voice and non-human voice noise, wherein the non-human voice noise includes flushing sounds or door opening and closing sounds; if so, then separate the toilet voice information into human voice stream information and noise stream information in real time; and perform recognition processing on the human voice stream information.

[0017] The technical solution described above improves the accuracy of keyword recognition under strong noise interference, ensuring reliable operation even in noisy environments. It endows the front-end voice recognition alarm unit with real-time "sound source separation" capabilities. When the system detects the simultaneous presence of human voice and strong noises such as flushing water or opening / closing doors, it first uses signal processing technology to separate the mixed sound into a clean human voice stream and a noise stream. Then, the system only performs keyword recognition on this "purified" human voice stream. This "noise reduction before recognition" processing flow greatly improves the signal-to-noise ratio of the voice signal, fundamentally solving the problem of strong background noise masking and interfering with recognition, and significantly enhancing the system's robustness and reliability in harsh acoustic environments.

[0018] Optionally, in some embodiments, the edge computing unit is also connected to the voice acquisition unit and is also used to update and train the second voice recognition model using the voice samples of the toilet obtained by the voice acquisition unit.

[0019] By employing the technical solution described in the above embodiments, the connection between the edge computing unit and the voice acquisition unit enables continuous optimization of the second speech recognition model. By acquiring real speech samples, the edge computing unit can update and train the model, allowing it to continuously adapt to new speech features and environmental changes, thereby improving recognition accuracy and robustness. This online learning mechanism reduces reliance on cloud resources, lowers data transmission latency, improves system response speed and operating efficiency, and ensures the timeliness and reliability of alarms.

[0020] Optionally, in some embodiments, when the difference between the first confidence level and the second confidence level is greater than a set difference threshold, the edge computing unit is further configured to send a model update instruction to the speech recognition alarm unit, and the speech recognition alarm unit is further configured to update the first speech recognition model according to the model update instruction.

[0021] Using the technical solution of the above embodiments, when the difference between the first confidence level and the second confidence level exceeds a set threshold, the edge computing unit sends a model update command to the speech recognition alarm unit, realizing the system's self-optimization and performance improvement. By dynamically adjusting the first speech recognition model according to the actual speech recognition results, the system can correct possible recognition deviations and improve the model's adaptability and accuracy. This mechanism ensures that the system maintains efficient speech recognition capabilities under different environments and usage conditions, enhancing the system's reliability and intelligence level.

[0022] Optionally, in some embodiments, the system further includes a broadcast unit connected to the voice recognition alarm unit and the management control center, for playing background music or receiving broadcast instructions from the voice recognition alarm unit to issue emergency notifications.

[0023] The technical solution adopted in the above embodiments enables the broadcast unit to promptly play background music or emergency notifications into the restroom. Under normal circumstances, background music can enhance user comfort and relaxation, creating a more pleasant environment. In emergencies, such as bullying or evacuation, the broadcast unit can quickly play emergency notifications to guide users to take appropriate actions and ensure safety. Furthermore, the connection between the broadcast unit and the voice recognition alarm unit and the management control center enables automated emergency response, improving the system's response speed and management efficiency.

[0024] Optionally, in some embodiments, the system further includes a remote terminal and a remote communication unit; the remote terminal is connected to the management and control center and is used to provide the management personnel with the operating status of the target toilet, including disinfection status, odor concentration, and alarm records; the remote communication unit is installed inside the target toilet and is connected to the management and control center for communication, and is used to receive notifications and instructions issued by the management personnel to the target toilet through the remote terminal.

[0025] By adopting the technical solution of the above embodiments, the introduction of remote terminals and remote communication units greatly enhances the intelligence and remote monitoring capabilities of toilet management. The remote terminal enables managers to obtain real-time information on the toilet's operational status, including disinfection status, odor concentration, and alarm records, facilitating timely scheduling of cleaning and maintenance work and optimizing management efficiency. The remote communication unit allows managers to communicate directly with users inside the toilet, promptly issuing notifications and instructions, such as equipment maintenance reminders and safety alerts, enhancing interaction between users and managers and improving user experience and safety. Attached Figure Description

[0026] Figure 1 This is a schematic diagram of the architecture of an intelligent toilet management system with alarm function according to an embodiment of the present invention; Figure 2 This is a flowchart of the processing operation of an edge computing unit in an embodiment of the present invention; Figure 3 This is a functional module diagram of an intelligent toilet management system with alarm function according to an embodiment of the present invention. Detailed Implementation

[0027] The terminology used in the following embodiments of the present invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used in the specification of the invention, the singular expressions “a,” “an,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in the invention refers to any or all possible combinations comprising one or more of the listed items.

[0028] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of the present invention, unless otherwise stated, "a plurality of" means two or more.

[0029] It should also be noted that, unless otherwise explicitly specified and limited, the terms "setting" and "connection" in the embodiments of the present invention should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can be a connection within two components; it can be a wired communication connection or a wireless communication connection. Those skilled in the art can understand the specific meaning of the above terms in the present invention according to the specific circumstances. The embodiments of the present invention will be described in detail below.

[0030] This invention provides an intelligent toilet management system (hereinafter referred to as the system) with an alarm function, such as... Figure 1 As shown, it includes a management control center 101, a voice acquisition unit 102, a voice recognition alarm unit 103, and an edge computing unit 104.

[0031] The management control center 101 is the core of the entire intelligent toilet management system and can be composed of high-performance servers or industrial computers. These servers are equipped with multi-core processors, large-capacity memory, and high-speed data storage devices, capable of processing large amounts of data from multiple sensors and units simultaneously. To achieve efficient data transmission and network communication, the management control center 101 is also equipped with high-speed network interface cards and wireless communication modules such as 5G, Wi-Fi, and Bluetooth, ensuring real-time data interaction with various units. Furthermore, for visual management and monitoring of the entire system, the management control center 101 is connected to professional input / output devices such as monitors, keyboards, and mice, facilitating operation and system status viewing for administrators. In some complex systems, data backup and recovery equipment, such as tape libraries or optical disc burners, may also be provided to ensure data security and reliability.

[0032] The voice acquisition unit 102 is installed inside the target toilet to acquire the toilet's voice information. The voice acquisition unit 102 mainly consists of a high-sensitivity microphone array. These microphones are carefully designed and arranged in various areas of the target toilet to ensure comprehensive coverage and clear capture of toilet voice information. To adapt to the complex acoustic environment inside the toilet, such as high background noise, sound reflection, and reverberation, the microphone array typically has directional and noise reduction functions, effectively filtering out environmental noise and highlighting the voice signal. In addition to the microphones themselves, the voice acquisition unit 102 is also equipped with a professional audio acquisition card. This card converts the analog audio signals captured by the microphones into digital signals and performs preprocessing operations such as amplification and filtering to improve the quality and usability of the voice signal. To ensure the stability and reliability of the system, the hardware of the voice acquisition unit 102 typically has a high protection level, being dustproof, waterproof, and moisture-proof, adapting to the humid and dusty environment inside the toilet.

[0033] The voice recognition alarm unit 103 is connected to the voice acquisition unit 102 and the management control center 101 respectively. It is used to recognize toilet voice information through a preset first voice recognition model. When a target keyword is recognized and the first confidence level of the target keyword is greater than the preset first confidence threshold, it is determined that a bullying event has occurred in the target toilet and an alarm message is sent to the management control center 101. The target keyword includes calling for help or bullying.

[0034] Confidence level refers to the degree of certainty the model has about the predicted result, usually expressed as a probability value. Target keywords include, but are not limited to, words representing emergency calls for help (such as "help me," "save me," "help someone quick"), and words highly related to bullying behavior (such as phrases with strong commands, crying, or fear, such as "hit," "rob," "give it back," "don't come any closer," etc.). Furthermore, the voice recognition alarm unit 103 recognizes not only isolated words but also non-verbal sound events such as continuous crying, screaming, and struggling sounds.

[0035] The voice recognition alarm unit 103 can be equipped with a dedicated voice processing chip or accelerator card. These hardware devices are optimized for voice signal processing and analysis, enabling them to execute voice recognition algorithms quickly and efficiently. To run a primary voice recognition model, such as a deep neural network model, the voice recognition alarm unit 103 can also be equipped with a graphics processing unit (GPU) or a tensor processing unit (TPU). These processors can handle a large number of computational tasks in parallel, significantly improving the speed and accuracy of voice recognition.

[0036] Furthermore, the voice recognition alarm unit 103 also features a large-capacity memory and high-speed cache for storing voice data and intermediate calculation results, ensuring smooth data processing. To enable communication with the management control center 101 and other units, the voice recognition alarm unit 103 is equipped with multiple communication interfaces, such as Ethernet and USB interfaces, supporting both wired and wireless communication. In its hardware design, the voice recognition alarm unit 103 prioritizes heat dissipation and stability to ensure reliable continuous operation over extended periods.

[0037] For example, when the voice recognition alarm unit 103 recognizes the voice information in the toilet through the first voice recognition model, if someone shouts "Help! Someone is bullying me!" in the public toilet, the first voice recognition model quickly captures the target keywords such as "help" and "bullying" and calculates the first confidence level as 85%. If the preset first confidence level threshold is 80%, the first confidence level is greater than the threshold, and it is determined that a bullying event has occurred in the target toilet. The voice recognition alarm unit 103 immediately sends an alarm message to the management control center 101, reporting that a bullying event may have occurred in the toilet and that it needs to be dealt with in a timely manner.

[0038] If the first confidence level is not greater than the first confidence level threshold, the voice recognition alarm unit 103 sends the toilet voice information to the edge computing unit 104 for further detection and confirmation.

[0039] The edge computing unit 104 is connected to the voice recognition alarm unit 103 and the management control center 101 respectively. When the first confidence level is less than or equal to the first confidence level threshold, it receives the toilet voice information transmitted by the voice recognition alarm unit 103, and performs recognition processing on the toilet voice information through a preset second voice recognition model to obtain the second confidence level of the target keyword. When the second confidence level is greater than the preset second confidence level threshold, it determines that a bullying event has occurred in the target toilet and sends alarm information to the management control center 101.

[0040] The edge computing unit 104 can be composed of an embedded computing platform or an edge server. These devices have a compact form factor and low power consumption. The edge computing unit 104 is equipped with adequate computing resources, including a multi-core processor and a certain amount of memory, to meet the needs of real-time preprocessing and preliminary analysis of voice data.

[0041] To achieve efficient data storage and management, the edge computing unit 104 incorporates a solid-state drive or high-speed memory card, enabling rapid data reading, writing, and storage. For communication, the edge computing unit 104 is equipped with various communication modules, such as Ethernet, WiFi, and Bluetooth, allowing for flexible data transmission with the voice acquisition unit 102, the voice recognition alarm unit 103, and the management control center 101. For operational stability and reliability, the edge computing unit 104's casing typically possesses excellent protective properties, such as dustproof, waterproof, and corrosion-resistant capabilities, while also exhibiting a certain degree of electromagnetic interference resistance to ensure stable operation in complex electromagnetic environments.

[0042] In addition, the edge computing unit 104 can also be equipped with a backup power supply or power management module to cope with sudden power outages and ensure data integrity and normal system operation.

[0043] For example, suppose that in a public restroom, the voice acquisition unit 102 acquires a voice message containing the phrase "Don't bully me, or I'll call for help." The voice recognition alarm unit 103, after recognizing the message using a first voice recognition model, finds it contains the keyword "bullying," which is suspected to be a bullying target. However, the first confidence level is 78%, lower than the preset first confidence threshold of 80%. The voice recognition alarm unit 103 then sends this voice message to the edge computing unit 104. At this time, the edge computing unit 104 receives the voice message and reprocesses it using a second voice recognition model. It obtains a second confidence level of 86% for the keyword "bullying," which is higher than the preset second confidence threshold of 85%. Therefore, it determines that a bullying event has occurred in the target restroom, and the edge computing unit 104 immediately sends an alarm message to the management control center 101, reporting that bullying may exist in the restroom and requires immediate attention.

[0044] In this embodiment, the second confidence threshold is greater than the first confidence threshold. This design aims to balance the false positive rate and the false negative rate, and to optimize system resource allocation. The first speech recognition model runs on the front-end device (speech recognition alarm unit 103). Due to limitations in computing resources and time, the recognition may not be completely accurate. Therefore, the first confidence threshold is set relatively low to initially filter out speech information that may contain the target keywords. Furthermore, by setting a relatively low first confidence threshold, high sensitivity can be achieved, quickly filtering potential dangerous events to avoid missed detections.

[0045] The second speech recognition model runs in the edge computing unit 104, which typically has more powerful computing capabilities and a more complex model structure, enabling it to more accurately identify target keywords in speech. To ensure the accuracy of the alarm and reduce false alarms, the second confidence threshold is set higher. For example, if the first speech recognition model identifies a speech segment that may contain keywords such as "call for help" or "bullying," but the confidence level is only 75% (below the first confidence threshold of 80%), the edge computing unit 104 will use the second speech recognition model to reprocess the speech segment. If the confidence level calculated by the second model is 92% (above the second confidence threshold of 90%), an alarm message is sent to the management control center 101.

[0046] By using two levels of identification and setting different thresholds, the timeliness of alarms is ensured, the reliability of alarms is improved, and an optimal balance is achieved between real-time performance, accuracy, and resource efficiency.

[0047] In this embodiment, the first speech recognition model is a student model obtained based on the second speech recognition model through knowledge distillation technology.

[0048] The first and second speech recognition models can employ different network architectures to balance performance and efficiency. The second speech recognition model, acting as the teacher model, can utilize a deep convolutional neural network (CNN) architecture, such as a VGG-like network, containing multiple convolutional and pooling layers. This effectively extracts high-level features of the speech signal, making it suitable for complex speech recognition tasks. The first speech recognition model, acting as the student model, can employ a lightweight network architecture, such as a network based on deep separable convolutions, reducing the number of parameters and computational cost, making it suitable for running on edge devices. Furthermore, the first model can incorporate an attention mechanism to focus on key parts of the speech signal, improving recognition accuracy. Through knowledge distillation, the feature representations and classification knowledge of the second model are transferred to the first model, allowing it to inherit the recognition capabilities of the second model while maintaining a smaller scale.

[0049] During training, the second speech recognition model acts as the teacher model, imparting its rich "knowledge" and experience to the first speech recognition model. Knowledge distillation technology allows the first speech recognition model to learn the output distribution and feature representation of the second model, enabling it to maintain high accuracy while having a smaller model size and faster running speed. This allows the first speech recognition model to be deployed more efficiently on resource-constrained front-end devices (speech recognition alarm unit 103), meeting the needs of real-time speech recognition and providing the system with fast and accurate speech recognition services.

[0050] The training process of the second speech recognition model is as follows: First, a large-scale, diverse collection of speech data is gathered, covering different accents, ages, genders, etc. The data is then labeled, including text transcription and keyword tagging. Next, a deep neural network architecture, such as a convolutional neural network or a recurrent neural network, is constructed for feature extraction and sequence modeling. The labeled data is input into the model, and forward propagation is used to calculate the loss function between the predicted output and the true label, such as cross-entropy loss. Backpropagation and an optimizer (such as Adam) are used to update the model parameters, iteratively optimizing until the model achieves satisfactory recognition accuracy and recall on the validation set, demonstrating strong expressive power and high accuracy, providing a solid foundation for the subsequent training of the first speech recognition model.

[0051] The training of the first speech recognition model (student model) is based on knowledge distillation, with the second speech recognition model serving as the teacher model. First, the initial architecture of the student model is prepared, typically smaller and more efficient than the teacher model. During training, the same speech data is input to both the teacher and student models simultaneously. The teacher model outputs soft labels containing probability distributions for each category. The student model learns not only the hard labels of the original data but also the soft labels from the teacher model, combining both sets of information through a loss function, such as a weighted sum of distillation loss and cross-entropy loss. Optimization algorithms are used to adjust the student model's parameters, allowing it to approach the performance of the teacher model as closely as possible while maintaining a smaller scale. After multiple rounds of iterative optimization, the student model ultimately achieves high recognition accuracy while meeting the efficiency and resource requirements of practical applications, making it suitable for resource-constrained environments such as edge computing.

[0052] The workflow of this toilet management system is as follows: First, the voice acquisition unit 102 is installed inside the target toilet to acquire voice information from inside the toilet. The voice acquisition unit 102 then transmits the acquired toilet voice information to the voice recognition alarm unit 103.

[0053] The voice recognition alarm unit 103 recognizes the toilet voice information using a preset first voice recognition model. When a target keyword is recognized and its first confidence level is greater than a preset first confidence threshold, the voice recognition alarm unit 103 sends an alarm message to the management control center 101. The target keyword may include a call for help or bullying. If the first confidence level is less than or equal to the first confidence threshold, the voice recognition alarm unit 103 sends the toilet voice information to the edge computing unit 104.

[0054] Edge computing unit 104 receives toilet voice information transmitted by voice recognition alarm unit 103, and processes the toilet voice information using a preset second voice recognition model to obtain a second confidence level for the target keyword. When the second confidence level is greater than a preset second confidence threshold, edge computing unit 104 sends alarm information to management control center 101.

[0055] After receiving the alarm information, the management control center 101 broadcasts an emergency notice through the broadcast unit 105 and simultaneously sends the alarm information to the remote terminal 106 of the management personnel, so that the management personnel can respond to the alarm information in time.

[0056] This embodiment employs the aforementioned system, where the voice acquisition unit 102 and the voice recognition alarm unit 103 work collaboratively. A first voice recognition model is used to recognize toilet voice information in real time. When a target keyword is identified and the first confidence level is greater than a preset threshold, an alarm message is quickly sent to the management control center 101, providing timely warnings for emergencies such as bullying and calls for help, effectively ensuring user safety. When the first confidence level is insufficient, the edge computing unit 104 takes over processing, using a second voice recognition model for re-identification. If the second confidence level exceeds a preset threshold, an alarm is sent. This dual recognition mechanism significantly improves alarm accuracy, reduces false alarms and missed alarms, and optimizes system reliability.

[0057] The first speech recognition model is trained from the second model using knowledge distillation technology, balancing recognition efficiency and accuracy while achieving a lightweight model. This satisfies the resource limitations of the front-end devices while ensuring recognition performance. The entire system has a reasonable architecture with clear division of labor among units. The management and control center 101 enables centralized management and coordination, improving the level of intelligence in toilet management and making the toilet environment safer, more hygienic, and more efficient.

[0058] The dual verification mechanism of "coarse screening + fine inspection" designed in this embodiment not only avoids the response delay caused by the excessive complexity of a single model, but also solves the problem of misjudgment of lightweight models in complex scenarios. This enables the system to effectively control alarm delay while ensuring a high recognition accuracy, thus guaranteeing the timeliness of alarms.

[0059] In this embodiment, the management control center 101 can be implemented using one or more high-performance servers or workstations deployed in the security monitoring room or property management center. This hardware device needs to be equipped with a large-capacity hard drive for storing alarm logs, recorded evidence, and system data, and possess a powerful processor to smoothly run the backend management software. To achieve intuitive monitoring and rapid response, it will be connected to one or more large-size displays to show the status of each restroom in real time in the form of an electronic map or list, and will display a prominent audio-visual alarm interface when an alarm is received. Furthermore, the center needs to have a stable and reliable network interface (such as wired Ethernet) to ensure smooth communication with all front-end units, and may integrate an alarm push module to send alarm information to patrolling security personnel via SMS, App notifications, etc.

[0060] The voice acquisition unit 102 can be specifically implemented in hardware as a dedicated acoustic sensor device with an integrated microphone array. The core of this device consists of multiple (e.g., 3 to 8) high-sensitivity MEMS (Micro-Electro-Mechanical Systems) microphones arranged in a specific geometry (such as a ring or linear array) on a printed circuit board (PCB). The entire microphone array is encapsulated in a rugged, waterproof, and vandal-resistant enclosure suitable for public environments (e.g., achieving an IP65 protection rating) to withstand the damp and frequently used environment of public restrooms. To simplify wiring and power supply, the unit can employ PoE (Power over Ethernet) technology, simultaneously completing data transmission and power supply via a single network cable, allowing for convenient installation on the ceiling or high on the wall of the restroom.

[0061] In terms of hardware implementation, the voice recognition alarm unit 103 can be a compact embedded computing device, often integrated with the voice acquisition unit within the same physical housing, or installed nearby as a separate small box. Its core is a low-power microcontroller (MCU) or a domain-specific system-on-a-chip (SoC) with AI acceleration capabilities, such as a chip based on the ARM Cortex-M series with a DSP instruction set or an NPU (neural network processing unit). This hardware is equipped with sufficient flash memory to store the knowledge-distilled lightweight student model (the first voice recognition model) and firmware, as well as a certain amount of RAM to process real-time audio stream data. Its design goal is to achieve extremely low operating power consumption while meeting the requirements of rapid front-end recognition.

[0062] The edge computing unit 104 is a more powerful hardware device that can be deployed in the building's low-voltage electrical room or server room. One edge computing unit can serve multiple smart toilet management systems within its coverage area. Its hardware form can be an industrial-grade computer (IPC), an edge server, or a dedicated edge AI computing box. It is equipped with a high-performance multi-core CPU and a dedicated GPU (Graphics Processing Unit), NPU, or other AI accelerator card for AI computation, possessing powerful parallel computing capabilities. This unit has larger memory and storage space, sufficient to run complex and accurate teacher models (secondary speech recognition models) and perform computationally intensive tasks such as sound source localization and multi-dimensional acoustic feature analysis, thereby providing accurate secondary analysis for ambiguous events that the front-end unit cannot determine.

[0063] On the other hand, this embodiment hereby declares that in order to avoid the privacy risks that may arise from collecting voice information in high-privacy environments such as toilets and to meet the compliance requirements of relevant laws and regulations, this embodiment follows the principles of data minimization and privacy by design.

[0064] Specifically, in some embodiments, after receiving the raw audio stream from the voice acquisition unit 102, the voice recognition alarm unit 103 does not directly store or transmit it. Instead, it first uses acoustic feature extraction algorithms such as Mel-frequency cepstral coefficients (MFCC) and linear predictive coding (LPC) to convert the brief audio segment into a set of irreversible digital feature vectors in real time. All subsequent recognition and judgment, whether using the first or second voice recognition model, are based on these desensitized feature vectors, rather than the original voice waveform. After processing, the original audio data cache is cleared, overwritten, or encrypted for protection.

[0065] In this way, this embodiment only obtains the acoustic pattern information needed to determine the event, without involving the specific content of the voice, thereby maximizing the protection of the user's personal privacy while achieving security warnings.

[0066] In some embodiments, the voice acquisition unit 102 includes a microphone array for acquiring multi-directional audio signals within the target toilet and transmitting the raw multi-channel audio signals to the edge computing unit 104.

[0067] The microphone array can consist of multiple high-sensitivity microphones arranged in a specific geometric layout, such as linear, circular, or three-dimensional. In this embodiment, the microphone array is specifically installed in multiple key locations within the toilet (such as corridors and cubicles) to collect audio signals from all directions and angles. Each microphone can independently capture sound, thereby obtaining rich spatial audio information. The microphone array can receive sound signals from different directions and convert them into electrical signals, i.e., raw multi-channel audio signals, which are then transmitted to an edge computing unit for further processing.

[0068] By having multiple microphones in the array work together, the system can achieve functions such as sound source localization, tracking, and audio signal enhancement, providing high-quality audio data support for subsequent speech recognition and event detection.

[0069] In this embodiment, as Figure 2 As shown, the edge computing unit 104 is also used to perform the following steps: Step 201: Based on the multi-channel audio signals collected by the microphone array, construct a sound source direction energy distribution map. The sound source direction energy distribution map is used to characterize the intensity distribution of sound energy at different directional angles, as well as the concentration index of the dominant sound source direction and the sound source direction.

[0070] This step aims to utilize spatial acoustic information to determine the distribution of sound sources within the toilet, distinguishing between normal conversation and potential abnormal gatherings or conflicts. The edge computing unit 104 first receives multi-channel audio signals synchronously acquired from a pre-set microphone array within the toilet (e.g., a circular or linear array of three or more microphones). It then processes these signals using sound source localization (SSL) algorithms, such as Controlled Response Power-Phase Transform (SRP-PHAT) or Multiple Signal Classification (MUSIC) algorithms. Specifically, the computing unit virtually scans all possible sound source directions (e.g., from 0° to 360°) and estimates the sound energy intensity in each direction by calculating the correlation or energy focusing of the signals in each direction.

[0071] The output of this process is a "sound source direction energy distribution map"—a graph with direction angle as the x-axis and sound energy as the y-axis. Based on this sound source direction energy distribution map, the edge computing unit automatically identifies the angle with the highest energy peak as the "dominant sound source direction," and further calculates the "aggregation index" by analyzing the number, width, and dispersion of energy peaks. For example, a single, sharp energy peak yields a high aggregation index, indicating that there is only one definite sound source (such as a person speaking normally); while multiple dispersed energy peaks or a broad, flat energy peak will yield a low aggregation index, suggesting the existence of multiple sound sources or sound sources moving rapidly, which may be the acoustic characteristics of multiple people confronting, pulling, or surrounding each other.

[0072] Step 202: Perform time-domain analysis on the total energy of the original multi-channel audio signal to calculate the instantaneous fluctuation rate of energy and the frequency of energy burst events, so as to quantify the intensity of the sound.

[0073] The purpose of this step is to quantify the intensity and suddenness of the sound, capturing the energy change patterns unique to events such as cries for help, struggles, or impacts. The edge computing unit first synthesizes the multi-channel audio signals into a single signal, or calculates the average energy of all channels. Then, it calculates the short-time energy within each extremely short time window (e.g., 20 milliseconds), thus obtaining an energy curve that varies over time.

[0074] The instantaneous volatility of energy is calculated by taking the difference or derivative of the energy values ​​at adjacent time points on the energy curve. A high volatility indicates that the volume has changed drastically in an instant, such as from quiet to a sudden scream or loud noise.

[0075] The frequency of energy bursts is measured by setting a dynamic or fixed energy threshold (e.g., more than three times the average background noise energy) and then counting the number of times the energy curve crosses that threshold upwards within a specific time period (e.g., 1 minute). Frequent energy bursts are a strong indicator of intermittent intense behavior in the environment, such as repeated banging sounds, intermittent shouting, or arguments.

[0076] Step 203: Perform frequency domain analysis on the original multi-channel audio signal to calculate the proportion of high-frequency energy and the clarity of the speech harmonic structure.

[0077] The clarity of the harmonic structure of speech is an indicator of whether a single, regular human voice characteristic exists in a sound signal. If the clarity of the harmonic structure decreases, it indicates that multiple people are speaking at the same time or that the sound is noisy and chaotic.

[0078] This step aims to extract emotional and environmental complexity information from the frequency components of the sound. High-pitched screams and chaotic noises have significant characteristics in the frequency domain. The edge computing unit performs a Fast Fourier Transform (FFT) or similar time-frequency analysis on the original audio signal, transforming it from the time domain to the frequency domain to obtain a spectrum. The "high-frequency energy percentage" is obtained by calculating the sum of the energy above a specific frequency in the spectrum (e.g., 1kHz, where tense emotions and screams in human voices are mainly distributed), and then dividing it by the total energy of the signal. A significantly increased high-frequency energy percentage is often directly related to high-pitched shouts in a state of fear or tension.

[0079] The clarity of speech harmonic structure is measured by analyzing the peak clarity and regularity of the fundamental frequency and its integer multiples of harmonics in the spectrum. When a single person speaks clearly, the harmonic structure is very regular and clear; however, when multiple people speak at the same time, there is an argument, or there is a lot of background noise interference, the harmonics from different sound sources will overlap and mask each other, causing the harmonic structure to become blurry and chaotic, and the clarity will be significantly reduced.

[0080] Specifically, the intelligibility of speech harmonic structure can be quantified by calculating the Harmonics-to-Noise Ratio (HNR). First, the fundamental frequency (F0) of the audio frame is estimated using the autocorrelation function method or cepstral method. Then, the energy (harmonic energy) located at integer multiples of the fundamental frequency in the signal spectrum is compared with the energy (noise energy) in the remaining frequency bands. A high HNR value (e.g., greater than 15 dB) indicates that the signal has a regular and clear harmonic structure, corresponding to a single, clear human voice; conversely, a low HNR value (e.g., less than 5 dB) means that the harmonic components are severely submerged by noise or that there is harmonic interference from multiple sound sources, indicating a high probability of a noisy and chaotic environment with multiple people speaking simultaneously. Through this quantification method, the system can convert a qualitative acoustic concept into a precise numerical feature that can be used for subsequent probabilistic model calculations. Step 204: Based on the concentration index of the sound source direction, the instantaneous fluctuation rate of energy, the frequency of energy bursts, the proportion of high-frequency energy, and the clarity of the speech harmonic structure, the probability of bullying events is calculated.

[0081] This step fuses the multiple independent acoustic features extracted in the previous steps to make a comprehensive and reliable judgment. The edge computing unit first normalizes or standardizes the five key features obtained from the calculation—the concentration index of the sound source direction (spatial distribution feature), the instantaneous fluctuation rate of energy (temporal variation feature), the frequency of energy bursts (temporal intensity feature), the proportion of high-frequency energy (frequency emotion feature), and the clarity of the speech harmonic structure (frequency complexity feature)—to eliminate dimensional differences.

[0082] Next, the standardized values ​​of these five key features are multiplied by a preset weighting coefficient that represents their importance—for example, the energy fluctuation rate, which represents intense conflict, has the highest weight, while the sound source concentration, which represents environmental conditions, has a relatively low weight.

[0083] Finally, the standardized values ​​of all features are multiplied by their corresponding weight coefficients, and the resulting sum is the probability of a bullying event. This probability value intuitively and logically reflects the overall similarity between the current acoustic environment and a typical bullying scenario.

[0084] In some embodiments, step 204 may generate a result including the following steps: Step 2041: Based on the instantaneous volatility of energy, the frequency of sudden energy events, and the proportion of high-frequency energy, a basic risk score is calculated.

[0085] The edge computing unit first performs a basic risk score calculation, which aims to quantify the direct danger signals inherent in the sound itself, such as intense conflict or loud screaming. To achieve this, the edge computing unit weights and fuses three features that directly reflect the intensity and tension of the sound: the instantaneous fluctuation rate of energy, the frequency of energy bursts, and the proportion of high-frequency energy.

[0086] Specifically, the edge computing unit normalizes each feature value, mapping it to a uniform range of 0 to 1. Then, it performs a weighted summation based on preset weights (for example, the proportion of high-frequency energy and instantaneous volatility may be given higher weights because they are more directly related to screams and sudden movements) to generate a comprehensive basic risk score. The higher the score, the closer the sound event itself is to a danger signal.

[0087] Step 2042: Calculate the scene confidence score based on the convergence index of the sound source direction and the clarity of the speech harmonic structure.

[0088] The edge computing unit calculates a scene confidence score, which assesses whether the current acoustic environment matches typical bullying scene characteristics such as multiple people gathering, confrontation, or chaos. In this step, the edge computing unit focuses on analyzing two descriptive environmental indicators: the clustering index of the sound source direction and the intelligibility of the speech harmonic structure. Since a lower clustering index indicates a more dispersed sound source (possibly a group of people surrounding the sound source), and lower harmonic intelligibility indicates a more chaotic and noisy sound (possibly a group of people arguing), the unit performs an inverse mapping of these two indicators, converting them into scores between 0 and 1, with higher values ​​representing higher scene conformity.

[0089] Then, the two converted scores are weighted and averaged to obtain a comprehensive scenario confidence score, which is used to determine how likely the current environment is to support the occurrence of bullying events.

[0090] Step 2043: Calculate the probability of bullying events based on the basic risk score and the scenario confidence score.

[0091] The edge computing unit calculates the final probability of a bullying event based on the basic risk score and scenario confidence score obtained in the first two steps. This decision fusion process is not a simple addition, but rather uses the scenario confidence score as a correction factor or adjustment weight.

[0092] In practice, when the scene confidence score is high (i.e., the acoustic environment is highly suspected to be a bullying scene), it will significantly amplify the contribution of the basic risk score to the final probability, so that even if the basic risk score is not extremely high, it can still trigger a high-probability alarm. Conversely, if the scene confidence score is very low (for example, the sound is clear and comes from a single direction), it will correspondingly suppress the influence of the basic risk score, thereby effectively filtering out instantaneous high-risk misjudgments caused by non-bullying events such as a sudden cough by a single person or the dropping of objects.

[0093] Through this dynamic adjustment mechanism, the final calculated probability of bullying events takes into account both the intensity of the event itself and the corroboration of the environmental background, thus making the judgment more accurate and reliable.

[0094] In some embodiments, the probability of a bullying event can be calculated using the following formula:

[0095]

[0096]

[0097] in: P bullying ( t ): Probability of a bullying event, representing the probability at a given time point.t To determine the final probability value of a bullying incident occurring.

[0098] α: Global sensitivity coefficient, used to amplify or reduce the overall risk assessment results.

[0099] R base ( t The baseline risk score at time t represents the "basic aggressiveness" score inherent in the sound itself, without considering environmental factors.

[0100] C scene ( t ): The scenario confidence score at time t represents the "dangerous scenario correction coefficient" of the current toilet environment, which is used to amplify or adjust the basic risk.

[0101] tanh (·): Hyperbolic tangent function, an S-shaped curve function, used as an activation function in this embodiment to smoothly map calculation results of arbitrary size to the probability interval (0, 1). Risk perception does not grow infinitely linearly. When the level of danger reaches a certain point, the general judgment tends towards "100% certainty of danger" rather than "200% danger." Therefore, when the internal calculated value is very small, the probability is also small; when the calculated value exceeds a certain range, the probability will rapidly saturate, approaching 1. tanh The design of (·) not only conforms to the mathematical definition of probability, but also simulates the "saturation effect" of human perception.

[0102] S v ( t ): The standardized value of energy fluctuation rate. The larger the value, the more drastic the change in sound loudness.

[0103] S b ( t ): A standardized value for the frequency of sudden energy bursts. The larger the value, the more frequent the sudden loud noises.

[0104] S f ( t ): A standardized value for the proportion of high-frequency energy; the larger the value, the sharper and more piercing the sound.

[0105] w v : Preset weighting coefficient for energy volatility.

[0106] w b : Preset weighting coefficient for burst frequency.

[0107] w f : Preset weighting coefficient for high frequency ratio.

[0108] Among them, R base The formula for calculating S uses the square root of the sum of squares (Euclidean norm), which is a manifestation of "OR" logic. As long as S... v or S b If any one component is large, its square value will increase dramatically, resulting in a large final "resultant force" value. This avoids the problem of "one indicator being high but being averaged out by another low indicator" that can occur with simple addition. It treats the two energy characteristics as orthogonal risk dimensions and calculates the magnitude of their resultant vector, making the physical meaning clear.

[0109] High frequency ratio S f Designed as an "emotional booster." When the voice is not shrill ( S f ≈0), this factor is 1, and it does not change the energy risk. When the sound is very sharp ( S f ≈1), this factor significantly amplifies energy risk. This is equivalent to saying, "This sound is not only loud and chaotic, but also jarring and terrifying, so its danger is doubled!" This is a clever application of acoustic psychology.

[0110] S c (t): Standardized value of spatial clustering. The larger the value, the more concentrated the direction of the sound source (there may be more people).

[0111] H d (t): The standardized value of speech harmonic disorder. The larger the value, the more noisy and disordered the sound (multiple people speaking at the same time). Speech harmonic disorder is calculated based on the clarity of speech harmonic structure. Harmonic disorder = 1 - clarity of speech harmonic structure.

[0112] T h The preset disorder threshold is used to determine whether the scene is "to the point where vigilance is needed".

[0113] k The degree of confusion is used to determine the steepness of the "soft switch," that is, to determine the ambiguity or clarity of the boundary.

[0114] w c : Preset weighting coefficient for spatial clustering.

[0115] β: Base confidence offset, ensuring that the scene correction coefficient is not zero.

[0116] exist C scene In the calculation formula, this embodiment observes that human perception of environmental "chaos" is usually not linear. The change from "absolute silence" to "some human voices" may not seem significant to humans. However, when the noise becomes so chaotic that it's impossible to understand anyone's speech, a "qualitative change" occurs – the perception that "the situation is out of control."

[0117] Using the Sigmoid function 1 / (1+e^-x) is an effective tool for simulating this "qualitative change." By setting a threshold... T h This defines a critical point between "order" and "chaos." Only when the degree of chaos... H d Once this critical point is crossed, its amplification effect on risk becomes significant. This allows the model to tolerate a certain degree of normal noise, making high-risk judgments only for truly "out-of-control" chaotic scenarios, greatly improving the model's intelligence and anti-interference capabilities, which is the fifth core rationality of the design.

[0118] This embodiment employs the aforementioned series of formulas, going beyond a simple listing and linear combination of acoustic features. Instead, it delves into the causal relationships, synergistic effects, and contextual dependencies behind various acoustic phenomena in bullying incidents. Through a series of carefully selected nonlinear mathematical structures with clear physical and psychological significance, it solidifies these deep logics into a computable, interpretable, and highly robust model. It is a meticulously constructed "building" based on acoustic principles and engineering wisdom, rather than a collection of randomly piled "bricks."

[0119] Step 205: When the probability of a bullying event exceeds a preset probability threshold, send an alarm message to the management control center.

[0120] This step is the final decision-making and execution stage, designed to ensure the reliability of alarm decisions and avoid false alarms. The edge computing unit continuously calculates and updates the probability of bullying events.

[0121] The edge computing unit presets a probability threshold (e.g., 0.85), which is scientifically set through a trade-off between precision and recall (ROC curve analysis) on a test dataset to achieve optimal alarm performance. When the newly calculated "bullying event probability" exceeds this preset threshold, the system determines that the likelihood of a bullying event occurring is extremely high, and the decision condition is met. At this point, the edge computing unit immediately generates a structured alarm message, which may include the time and location of the event (specific toilet stall), the judgment criteria (such as a high probability value), etc., and then sends this alarm message to the management control center via a network connection. Upon receiving the information, the management center can trigger audible and visual alarms, push alerts to security personnel's mobile terminals, or link video surveillance (if applicable and compliant) and other subsequent emergency response measures.

[0122] In some embodiments, the background noise levels and acoustic characteristics of different toilets vary significantly at different times of day (e.g., during peak hours there is a lot of noise and frequent flushing, while in the early morning it is very quiet with only low-frequency noise from the ventilation equipment). If a fixed confidence threshold is used, it is difficult to simultaneously achieve high sensitivity and low false alarm rate. In noisy environments, a fixed low threshold is prone to false alarms; while in quiet environments, a fixed high threshold may miss weak distress signals.

[0123] To address the aforementioned issues, this embodiment also incorporates an environmental acoustic baseline learning module on the edge computing unit. This module periodically performs statistical analysis on the background speech information acquired by the speech acquisition unit to generate a dynamic acoustic baseline characterizing the normal acoustic environment of the target toilet at different times.

[0124] The environmental acoustic baseline learning module is designed as a self-learning unit that runs continuously in the background, functioning as a software program running on the hardware of the edge computing unit 104. Its core function is to generate an "acoustic fingerprint map" that changes over time for each target toilet.

[0125] In practice, the environmental acoustic baseline learning module periodically (e.g., every 10 minutes) acquires an audio sample from the voice acquisition unit. When acquiring a sample, it first confirms that the system is not currently in an alarm state to ensure that the acquired audio is "normal" background sound information, rather than an abnormal event itself. For the acquired audio samples, the environmental acoustic baseline learning module performs a series of acoustic feature analyses to extract key parameters, such as: average sound energy level, spectral centroid (reflecting the overall frequency response of the sound), energy variance (reflecting volume stability), and the energy proportion of specific frequency bands (e.g., the sound of flushing water or a hand dryer).

[0126] By performing statistical and cluster analysis on a large amount of data collected over a 24-hour period (or a longer period, such as a week), the environmental acoustic baseline learning module can generate a dynamic acoustic baseline. This dynamic acoustic baseline is not a single value, but a time-series database that records in detail the normal acoustic environmental characteristics of the target toilet during different typical time periods (such as "weekday morning rush hour: 7:00-9:00", "afternoon off-peak: 14:00-16:00", "nighttime quiet period: 0:00-5:00", etc.). For example, the dynamic acoustic baseline will clearly show that during the morning rush hour, the normal average sound energy level is around -20dB, with a large energy variance; while during the nighttime quiet period, the average sound energy level may be as low as -50dB, with an extremely small energy variance.

[0127] In some embodiments, to ensure the accuracy and robustness of the dynamic acoustic baseline, the learning cycle of the environmental acoustic baseline learning module is set to a dynamically adjusted, relatively long time scale, such as data acquisition and model update once per hour, in order to smooth out short-term noise fluctuations.

[0128] More importantly, this module incorporates a built-in mechanism to eliminate interference from abnormal events. Before collecting background audio samples for baseline updates, the system first performs a pre-screening: it uses the alarm judgment logic (including keyword recognition and bullying event probability calculation) provided earlier to quickly evaluate the background audio sample. Only when the sample is determined not to contain alarm event features (i.e., all confidence and probabilities are below the alarm threshold) will it be considered clean background noise and included in the baseline learning dataset. If potential abnormal signals such as brief screams are detected in the sample, the sample will be discarded and excluded from baseline updates.

[0129] This embodiment, through its filtering-then-learning strategy, effectively prevents abnormal events from contaminating the acoustic baseline of the normal environment, thus ensuring the long-term stability of the baseline model.

[0130] In some embodiments, the voice recognition alarm unit 103 is further configured to invoke the corresponding dynamic acoustic baseline according to the current real-time time and adaptively adjust the first confidence threshold; wherein, when the ambient background noise represented by the dynamic acoustic baseline is low, the first confidence threshold is reduced to improve detection sensitivity.

[0131] The voice recognition alarm unit 103 adds a preprocessing step before performing real-time voice recognition. First, it obtains the current system real-time time. Then, it queries the environmental acoustic baseline learning module and retrieves the dynamic acoustic baseline data that matches the current time period. After obtaining the dynamic acoustic baseline, the unit dynamically sets the first confidence threshold used for this recognition based on a preset adaptive adjustment strategy.

[0132] The core logic of this adaptive adjustment strategy is that the detection sensitivity should be inversely proportional to the background noise level. The specific implementation is as follows: When the invoked dynamic acoustic baseline indicates a period of extremely low background noise (e.g., the quiet of night), any weak human voice signal is considered highly significant. A victim might emit a weak cry for help due to exhaustion or coercion, with very low voice energy, resulting in a low confidence score (e.g., 0.65) from the speech recognition model. Using a fixed, high threshold (e.g., 0.8) in this situation would lead to missed detections. Therefore, the first confidence threshold should be automatically lowered (e.g., from the default 0.8 to 0.6) to significantly improve the detection sensitivity for such weak, low-confidence signals, ensuring that requests for help are not missed.

[0133] When the invoked dynamic acoustic baseline indicates that the current period is characterized by noisy and variable background noise (e.g., peak hours), the environment is filled with various noises that can interfere with identification. In this situation, to avoid misidentifying blurry background voice fragments or noise as keywords and causing frequent false alarms, a high initial confidence threshold must be maintained, and even slightly increased if necessary. This requires that an initial alarm be triggered only when the identified target keyword is very clear and specific, with a sufficiently high confidence score, thus ensuring the system's robustness and low false alarm rate in complex environments.

[0134] This embodiment introduces an environmental acoustic baseline learning module and a threshold adaptive adjustment mechanism. This system can intelligently learn and adapt to the unique, time-varying environmental characteristics of each toilet, realizing a shift from a static threshold judgment that is "one-size-fits-all" to a dynamic and intelligent decision-making process that is "tailored to local conditions and time." This greatly improves the system's comprehensive detection performance in various real-world scenarios, ensuring sensitive response at critical moments while effectively suppressing false alarms caused by environmental interference.

[0135] In some embodiments, the voice recognition alarm unit 103 is further configured to: extract the energy change characteristics of the toilet voice information in the time domain; determine whether the energy change characteristics match the preset voice energy reference characteristics; if so, transmit the toilet voice information to the edge computing unit.

[0136] Specifically, firstly, the voice recognition alarm unit 103 needs to quickly identify segments in the audio stream that may contain human speech. The voice recognition alarm unit continuously performs real-time temporal analysis on the raw audio signal received from the voice acquisition unit. It does not immediately perform complex speech recognition, but instead first calculates a series of underlying acoustic features that can depict the dynamic contours of the sound. The core calculation is short-time energy, which involves continuously calculating the energy level of the audio signal within a sliding time window (e.g., 20-30 milliseconds), thereby generating an energy envelope curve that varies over time. Based on this curve, key feature parameters are further extracted, such as the "onset time" of the energy peak (i.e., the time required for the energy to rapidly jump from the background level to the peak), the "duration" of the high-energy segment, and the "fluctuation pattern" or "modulation frequency" of the energy envelope (reflecting whether the sound has periodic fluctuations similar to syllables). These features together constitute an energy change feature that can describe the "transient" and "steady-state" behavior of the sound.

[0137] The extracted energy change features are then compared with a pre-defined "human speech energy model" to filter out obvious non-speech noise. This "pre-defined speech energy reference feature" is not a model of specific words, but a set of rules or thresholds that define the dynamic range of energy in typical human speech, shouting, or screaming. For example, the reference feature stipulates that: the energy "onset time" of an effective speech signal should be within a certain millisecond range (too slow might be a device startup sound, too fast might be current pulse noise); its "duration" should conform to the pronunciation length of a word or short phrase; and its energy envelope should have a certain "fluctuation pattern" to distinguish it from mechanical noises such as flushing water or fan noise, which have relatively stable or monotonous energy changes. The speech recognition alarm unit compares the feature vectors extracted in real time with these pre-defined rules one by one. As long as all features fall within the reasonable range of the reference model, it initially determines that the audio segment has the potential to be "human-like".

[0138] Finally, the voice recognition alarm unit 103 packages the complete original toilet voice information that has just passed the energy feature screening (usually including a short context before and after the matching point to ensure information integrity) and actively sends it to the more powerful edge computing unit via the internal network connection. The core significance of this is that it achieves optimized allocation of computing resources: for most background noise that does not have "human voice-like" characteristics, the processing flow is efficiently terminated at the low-power front-end alarm unit, avoiding meaningless data transmission and subsequent complex analysis; only those high-value audio segments that have passed the initial "sound recognition" screening are allowed to proceed to the back-end for in-depth processing, thus ensuring that the entire system maintains high sensitivity while also possessing extremely high operating efficiency and low power consumption characteristics.

[0139] In some embodiments, the voice recognition alarm unit 103 is further configured to: analyze the real-time collected toilet voice information, identify whether it contains both human voice and non-human voice noise, including flushing sounds or door opening and closing sounds; if so, separate the toilet voice information into human voice stream information and noise stream information in real time; and perform recognition processing on the human voice stream information.

[0140] Specifically, firstly, the voice recognition alarm unit 103 performs analysis by running multiple dedicated detectors based on acoustic features in parallel. It activates a voice activity detection algorithm, which works by calculating the fundamental frequency (pitch) and zero-crossing rate (ZCR) of the audio signal: a regular fundamental frequency and harmonic structure are typical characteristics of human voices, while a higher ZCR is usually associated with voiceless consonants. Simultaneously, a flushing noise detection algorithm is run, which is based on spectral analysis and specifically seeks out persistent, flat-distributed, broadband noise features covering low to high frequencies. There is also an impact noise detection algorithm, which monitors the instantaneous energy of the time-domain signal, looking for waveforms where the energy value rises sharply and then decays rapidly within a very short time; this is a typical indicator of impact sounds such as door opening and closing. If the flushing noise detection algorithm or the impact noise detection algorithm also detects non-human voice noise while the voice activity detection algorithm detects human voice, then the current scene is determined to be a mixed environment where human and non-human voice noise coexist.

[0141] Once the presence of mixed human voice and non-human voice noise is confirmed, signal separation can be achieved using spectral subtraction. First, the spectral estimation data of the current background noise needs to be obtained. This estimation data can come from a pure noise segment immediately preceding the human voice, or a noise spectral model dynamically generated based on the noise characteristics identified by a flushing sound detector. Next, the real-time acquired mixed audio is subjected to a short-time Fourier transform to obtain its spectrum. Then, in the spectral domain, the estimated noise energy spectrum is subtracted from the energy spectrum of the mixed signal. The core logic is to assume that the energy of the human voice and noise is simply superimposed. After the subtraction, the remaining spectrum mainly contains the energy information of the human voice. Then, this "purified" human voice spectrum is converted back to a time-domain waveform using an inverse Fourier transform, thus obtaining a "human voice stream information" with significantly suppressed noise; the subtracted portion constitutes the "noise stream information."

[0142] Finally, a template-matching-based Dynamic Time Warping (DTW) algorithm is used to identify keywords in the purified human speech stream. The system pre-records and stores standard pronunciation templates for target keywords (such as "help") and converts them into acoustic feature sequences (such as MFCC Mel-frequency cepstral coefficients). During runtime, the real-time separated human speech stream information is also converted into the same type of feature sequence. Then, the DTW algorithm calculates the "similarity distance" between the real-time speech sequence and the pre-stored keyword template sequence. It finds the optimal matching path between the two sequences by non-linearly warping the time axis, thus effectively solving the problem of inconsistent speech speed in actual speech. If the calculated similarity distance is less than a pre-set threshold, the system determines that the target keyword has been successfully matched and triggers the subsequent recognition processing.

[0143] In some embodiments, the edge computing unit 104 is also connected to the voice acquisition unit 102 and is further used to update and train the second voice recognition model using the voice samples of the toilet acquired by the voice acquisition unit 102.

[0144] The connection between the edge computing unit 104 and the voice acquisition unit 102 enables continuous optimization of the second speech recognition model. The edge computing unit 104 acquires voice samples from the toilet through the voice acquisition unit 102. These samples contain diverse real-world voice data, such as different accents, speech rates, volumes, and background information. The edge computing unit 104 uses these rich voice samples to update and train the second speech recognition model, enabling the model to continuously learn and adapt to new voice features, thereby improving the accuracy and robustness of recognition. This online update training mechanism allows the system to promptly capture and recognize new bullying or distress cries, enhancing the system's adaptability to complex voice environments. Simultaneously, by updating and training the model on the edge computing unit 104, reliance on cloud computing resources is reduced, data transmission latency is lowered, and the system's response speed and operating efficiency are improved.

[0145] Specifically, the edge computing unit 104 first preprocesses these speech samples, including noise reduction and enhancement operations, to improve speech quality. Then, the processed speech samples, along with corresponding labels (such as whether they contain target keywords), are input into the second speech recognition model. The model calculates the prediction result through forward propagation and compares it with the true labels to calculate the loss function. Next, the backpropagation algorithm is used to adjust the model's parameters, optimizing the model's weights and biases to minimize the prediction error. In this process, the edge computing unit 104 can utilize limited computing resources to iteratively train the model multiple times, gradually improving its recognition accuracy and robustness. Simultaneously, because the training process is performed on an edge device, it reduces dependence on cloud resources, lowers data transmission latency, and enables the model to adapt to new speech features and environmental changes more quickly.

[0146] In some embodiments, when the difference between the first confidence level and the second confidence level is greater than a set difference threshold, the edge computing unit 104 is further configured to send a model update instruction to the voice recognition alarm unit 103, and the voice recognition alarm unit 103 is further configured to update the first voice recognition model according to the model update instruction.

[0147] Specifically, the edge computing unit 104 first detects the difference between the first confidence level and the second confidence level, determining that the first speech recognition model may have inaccurate recognition or needs optimization. Therefore, the edge computing unit 104 sends a model update instruction to the speech recognition alarm unit 103. This instruction includes optimization parameters and adjustment strategies for the first speech recognition model, based on the edge computing unit 104's analysis of speech samples and the recognition results of the second speech recognition model.

[0148] After receiving the instruction, the voice recognition alarm unit 103 updates the first voice recognition model using these parameters. The update process may include adjusting the model's weights, optimizing the algorithm's threshold settings, etc., to improve the model's recognition accuracy and adaptability. This mechanism ensures that the system can continuously self-optimize, adapting to different voice environments and recognition needs, thereby improving the overall system reliability and performance.

[0149] In some embodiments, such as Figure 3 As shown, the system may also include one or more of the following: broadcast unit 105, remote terminal 106, and remote communication unit 107.

[0150] The broadcast unit 105 is connected to the voice recognition alarm unit 103 and the management control center 101 respectively, and is used to play background music or receive broadcast instructions from the voice recognition alarm unit 103 to issue emergency notifications.

[0151] Specifically, the broadcast unit can be one or more waterproof speakers installed on the ceiling or walls of the target toilet, connected to a local audio controller via audio cable or wirelessly. This audio controller then communicates with the voice recognition alarm unit 103 and the management control center 101, acting as an intelligent audio playback terminal.

[0152] Under normal circumstances, the broadcast unit 105 can play soothing background music to create a comfortable atmosphere and enhance the user experience. In case of an emergency, such as when the voice recognition alarm unit 103 identifies the target keyword and the confidence level meets the requirements, the broadcast unit 105 will immediately switch to emergency notification mode to promptly convey emergency information to all people in the restroom and guide them to make the correct response.

[0153] For example, in the event of bullying in a school restroom, the broadcast unit 105 can quickly play a pre-set warning message to notify other students to stay away from the danger zone and to inform the perpetrator to immediately stop the inappropriate behavior. Simultaneously, the broadcast unit 105 can also link with the management control center 101 to send emergency notifications to administrators, enabling them to quickly take intervention measures and effectively protect user safety.

[0154] For example, in a shopping mall restroom, the broadcast unit 105 routinely plays soft background music to create a comfortable atmosphere. When the voice recognition alarm unit 103 detects someone shouting "Help!" and the confidence level is met, the broadcast unit 105 immediately switches to playing a preset emergency notification, such as "Please note that there is an emergency on site. Please provide assistance and maintain order." At the same time, it sends a message to the management control center 101, simultaneously notifying management personnel to respond quickly and ensure the safety of people on site.

[0155] In some embodiments, the remote terminal 106 is connected to the management control center 101 and is used to provide the management personnel with the operating status of the target toilet, including disinfection status, odor concentration and alarm records; the remote communication unit 107 is installed in the target toilet and is connected to the management control center 101 for receiving notifications and instructions issued by the management personnel to the target toilet through the remote terminal 106.

[0156] By adopting the above design, the comprehensive management capabilities and human-computer interaction of the intelligent toilet management system can be further enhanced, constructing a more comprehensive and proactive management closed loop. This embodiment upgrades the system from a passive alarm tool to a proactive, remotely interactive comprehensive management platform, thereby solving the technical problems of information silos, limited response methods, and low on-site management efficiency in traditional management.

[0157] The remote terminal 106 is a user interface for management personnel. Its hardware can be client software installed on a computer, a web browser interface, or a dedicated application (APP) developed for smartphones and tablets. This terminal connects to the management and control center via a secure network, retrieving and integrating data from the central database in real time. On its interface, management personnel can not only view alarm records generated by the system's core functions but also intuitively see the overall operational status collected after linkage with other intelligent sensors in the toilet (such as ultraviolet disinfection lamp controllers and ammonia / hydrogen sulfide concentration sensors).

[0158] Specifically, the interface clearly displays the real-time "disinfection status" (such as "disinfecting in progress" or "disinfection completed"), quantitative values ​​of "odor concentration," and trend curves for each target toilet in the form of charts or lists. This information integration enables managers to grasp the safety, hygiene, and environmental conditions of toilets in a one-stop, comprehensive manner, providing strong data support for refined management and predictive maintenance.

[0159] The remote communication unit 107 is an audio-visual interactive hardware device deployed in the target toilet. It typically integrates a speaker, microphone, and indicator light, is installed in a prominent position at the toilet entrance or in a public area, and is designed to be waterproof and moisture-proof.

[0160] The remote communication unit 107 establishes a communication link with the management and control center via a network (such as VoIP protocol). Its core function is to serve as a remote "voice" outlet for management personnel. When management personnel need to disseminate information to people in the restrooms, they can operate through the remote terminal. For example, after the system triggers a fire alarm or security alarm, management personnel can use the microphone of the remote terminal to directly issue emergency evacuation instructions to the remote communication unit in the alarmed restroom; or in daily management, they can periodically broadcast reminders for civilized restroom use, issue lost and found notices, etc. This remote speaking function greatly expands the intervention methods of management personnel, enabling them to remotely command and reassure in emergencies, or efficiently convey information in daily work, significantly improving management efficiency and emergency response capabilities.

[0161] The remote communication unit 107 can be a wall-mounted or ceiling-mounted waterproof IP communication device, a smart speaker, etc. These devices are installed in the toilet and have waterproof, moisture-proof, and noise-resistant characteristics. They can clearly receive notifications and instructions issued by the management personnel through the remote terminal 106 and play them to the users in the toilet to ensure effective communication of information.

[0162] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An intelligent toilet management system with alarm function, characterized in that, include: Management and control center; A voice acquisition unit is installed inside the target toilet to acquire the toilet's voice information. A voice recognition alarm unit is connected to the voice acquisition unit and the management and control center respectively. It is used to recognize the voice information of the toilet through a preset first voice recognition model. When a target keyword is recognized and the first confidence level of the target keyword is greater than the preset first confidence threshold, it is determined that a bullying event has occurred in the target toilet and an alarm message is sent to the management and control center. The target keyword includes calling for help or bullying. An edge computing unit, connected to both the voice recognition alarm unit and the management control center, is used to receive the toilet voice information transmitted by the voice recognition alarm unit when the first confidence level is less than or equal to the first confidence level threshold, and to process the toilet voice information using a preset second voice recognition model to obtain the second confidence level of the target keyword. When the second confidence level is greater than the preset second confidence level threshold, it is determined that a bullying event has occurred in the target toilet, and an alarm message is sent to the management control center. The first speech recognition model is a student model obtained based on the second speech recognition model through knowledge distillation technology.

2. The system according to claim 1, characterized in that, The voice acquisition unit includes a microphone array, which is used to acquire multi-directional audio signals within the target toilet and transmit the raw multi-channel audio signals to the edge computing unit; the edge computing unit is also used for: Based on the multi-channel audio signals collected by the microphone array, a sound source direction energy distribution map is constructed. The sound source direction energy distribution map is used to characterize the intensity distribution of sound energy at different directional angles, as well as the concentration index of the dominant sound source direction and the sound source direction. The total energy of the original multi-channel audio signal is analyzed in the time domain to calculate the instantaneous fluctuation rate of the energy and the frequency of energy burst events, so as to quantify the intensity of the sound. Frequency domain analysis is performed on the original multi-channel audio signal to calculate the proportion of high-frequency energy and the clarity of the speech harmonic structure. If the clarity of the harmonic structure decreases, it indicates that multiple people are speaking at the same time or that the sound is noisy and chaotic. The probability of a bullying event is calculated based on the concentration index of the sound source direction, the instantaneous fluctuation rate of the energy, the frequency of the energy bursts, the proportion of high-frequency energy, and the clarity of the speech harmonic structure. When the probability of the bullying event exceeds a preset probability threshold, an alarm message is sent to the management and control center.

3. The system according to claim 2, characterized in that, The edge computing unit is also specifically used for: A basic risk score is calculated based on the instantaneous volatility of the energy, the frequency of the energy-related sudden events, and the proportion of high-frequency energy. Based on the convergence index of the sound source direction and the clarity of the speech harmonic structure, the scene confidence score is calculated. The probability of the bullying event is calculated based on the basic risk score and the scenario confidence score.

4. The system according to claim 1, characterized in that, The edge computing unit is also equipped with an environmental acoustic baseline learning module, which is used to periodically perform statistical analysis on the background voice information acquired by the voice acquisition unit to generate a dynamic acoustic baseline that characterizes the normal acoustic environment of the target toilet at different times. The voice recognition alarm unit is further configured to invoke the corresponding dynamic acoustic baseline according to the current real-time time and adaptively adjust the first confidence threshold; wherein, when the environmental background noise represented by the dynamic acoustic baseline is low, the first confidence threshold is reduced to improve detection sensitivity.

5. The system according to claim 1, characterized in that, The voice recognition alarm unit is also used for: Extract the energy change characteristics of the toilet voice information in the time domain; Determine whether the energy change characteristics match the preset speech energy reference characteristics; If so, the toilet voice information is transmitted to the edge computing unit.

6. The system according to claim 1, characterized in that, The voice recognition alarm unit is also used for: The real-time collected toilet voice information is analyzed to identify whether it contains both human voices and non-human voice noise, including flushing sounds or door opening and closing sounds. If so, the toilet voice information will be separated into human voice stream information and noise stream information in real time; The human voice audio stream information is then recognized and processed.

7. The system according to any one of claims 1-6, characterized in that, The edge computing unit is also connected to the voice acquisition unit and is used to update and train the second voice recognition model using the voice samples of the target toilet obtained by the voice acquisition unit.

8. The system according to claim 7, characterized in that, When the difference between the first confidence level and the second confidence level is greater than a set difference threshold, the edge computing unit is further configured to send a model update instruction to the speech recognition alarm unit, and the speech recognition alarm unit is further configured to update the first speech recognition model according to the model update instruction.

9. The system according to claim 1, characterized in that, It also includes remote terminals and remote communication units; The remote terminal is connected to the management and control center and is used to provide the management personnel with the operating status of the target toilet, including disinfection status, odor concentration and alarm records. The remote communication unit is installed inside the target toilet and is connected to the management and control center for receiving notifications and instructions issued by the management personnel to the target toilet through the remote terminal.

10. The system according to claim 1, characterized in that, It also includes a broadcasting unit, which is connected to the voice recognition alarm unit and the management control center respectively, and is used to play background music or receive broadcast instructions from the voice recognition alarm unit to issue emergency notifications.