Alarm system based on voice recognition
Through the combination of embedded installation and ResCNN-BiGRU model, the problem of campus alarms being easily destroyed and low speech recognition accuracy is solved, hidden installation and efficient speech recognition are achieved, and real-time and accuracy of campus security monitoring are improved.
Patent Information
- Application Number
- CN202510449348.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing alarms are easily deliberately avoided and destroyed in campus bullying scenarios, resulting in the failure of the alarm system, and the accuracy of speech recognition and poor real-time performance.
An alarm system based on speech recognition is designed, using an embedded installation structure, combined with the ResCNN network and BiGRU model, to realize hidden installation and anti-tearing protection, and voice signal processing is carried out through pre-emphasis processing, Hamming window filtering and bidirectional gated loop unit to improve recognition accuracy and real-timeness.
It realizes hidden installation and anti-tearing protection of the alarm system, improves the accuracy and real-timeness of voice recognition, and ensures the effectiveness of campus safety monitoring.
Smart Images

Figure CN120299159A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of alarm security, and particularly relates to an alarm system based on voice recognition. Background Art
[0002] School bullying is a special type of aggressive behavior that occurs in a school setting, where one or more students intentionally inflict harm on another or multiple students, causing damage to the physical, psychological, property, etc. of the victim.
[0003] The anti-bullying alarm meets the requirements of current campus security improvement, cooperates with cameras to achieve full coverage of the campus without dead angles, and forms a "campus sky net". Especially for behavioral bullying in the toilet and dormitory scenarios, by using sound detection and voice recognition technologies, the traditional manual alarm mode is upgraded to a voice alarm mode.
[0004] For example, a voice recognition alarm system and its recognition method for a bathroom with the application number CN202310901557.1, the system includes: a voice collection module; a voice storage module; a voice recognition module, the voice recognition module interacts with the voice storage module, and the voice recognition module retrieves sound data from the voice storage module and uses an end-to-end voice model for recognition and analysis; an alarm module, the alarm module interacts with the voice recognition module, and after the alarm module receives the information analyzed by the voice recognition module, it determines whether to activate the alarm. It can judge and alarm when bullying occurs, and at the same time can accurately locate the possible location of bullying and notify relevant personnel to go and check, fundamentally avoiding the occurrence of bullying.
[0005] Currently, there are certain deficiencies in the installation of alarms. Existing alarms are generally directly installed in the environment, making the alarm system not have a hidden effect. As a result, when bullying occurs, the alarm system will be deliberately avoided, making the time to deal with school bullying complicated. At the same time, the alarm is not protected, making the alarm vulnerable to damage and illegal removal, resulting in the alarm system losing its original function. Therefore, it is urgent to design an alarm system based on voice recognition to solve the above problems. Summary of the Invention
[0006] The purpose of the present invention is to provide an alarm system based on voice recognition to solve the above deficiencies in the prior art.
[0007] To achieve the above purpose, the present invention provides the following technical solutions:
[0008] An alarm system based on voice recognition, an alarm system based on voice recognition, includes an embedded body, the embedded body is embedded and installed in a groove inside the wall, a face shell is inserted inside the embedded body, a voice recognizer is arranged above the inside of the face shell, and the voice recognizer is at least used for collecting sounds in the environment;
[0009] A main control chip board distributed around the voice recognizer is arranged inside the face shell; the main control chip board is at least used for processing the sounds collected by the voice recognizer;
[0010] A camera module is arranged inside the face shell, the camera module is at least used for real-time video recording when a bullying incident occurs, a closing mechanism for blocking the camera module is arranged on the inner side of the face shell, and the closing mechanism is at least used for blocking the camera module;
[0011] Anti-disassembly structures are arranged on both inner walls of the embedded body, and the face shell is inserted inside the embedded body through the anti-disassembly structures;
[0012] A driving mechanism is arranged inside the embedded body, and the driving mechanism is at least used for opening the anti-disassembly structure to realize the disassembly of the face shell from inside the embedded body.
[0013] Preferably, a sound transmission mesh hole corresponding to the position of the voice recognizer is arranged above the front of the face shell;
[0014] A camera window is arranged in the middle of the face shell.
[0015] Preferably, a tongue plate protruding outwards is arranged on the outer surface of the face shell;
[0016] The outer wall of the embedded body coincides with the outer wall of the tongue plate.
[0017] Preferably, a backup battery is fixedly arranged below the inside of the face shell, and the backup battery is at least used for supplying power when the alarm system is powered off.
[0018] Preferably, the main control chip board includes a voice signal processor and an information exchange module, the voice signal processor is at least used for processing the collected language, and the information exchange module is at least used for remote transmission of data.
[0019] Preferably, the closing mechanism includes two relatively sliding closing plates, guide blocks are fixedly connected to the top and bottom of the closing plates, a lining plate is fixedly arranged on the back of the closing plates, a first rack and a second rack connected to the lining plate are respectively arranged below the two closing plates, a driving gear is meshed in the middle of the first rack and the second rack, and a driving motor is arranged on one side of the driving gear and is fixed on the inner wall of the face shell through a support rod.
[0020] Preferably, the imaging module includes a camera, and connection seats are arranged on both sides of the camera, and the ends of the connection seats are fixed on the inner wall of the front shell.
[0021] Preferably, the anti-disassembly structure includes a plurality of mounting lugs fixed on the inner wall of the embedded body, and a clamping groove is arranged on the outer wall of the top of the mounting lug;
[0022] A plurality of elastic support rods are fixedly arranged on the back of the front shell, and an embedded block inserted into the clamping groove is arranged below the end of the elastic support rod, and a plug rod is fixedly arranged on the outer wall of the top of the elastic support rod;
[0023] A support arm is arranged on the inner wall of one side of the front shell, and a conductive column is fixed at the end of the support arm. A power receiving sleeve is fixedly arranged inside the embedded body, and the conductive column is inserted into the power receiving sleeve.
[0024] Preferably, the driving mechanism includes a plurality of restraint sleeves fixed on the inner wall of the embedded body, and a synchronous frame is slidably inserted into the restraint sleeve, and a plurality of limiting rods are arranged on the outer walls of both sides of the synchronous frame;
[0025] A driving motor two is arranged on the inner wall of the bottom of the embedded body, and a swing rod is connected to the output shaft of the driving motor two, and the top end of the swing rod is in sliding contact with the lower part of the synchronous frame.
[0026] An alarm method based on voice recognition, including the described alarm system based on voice recognition, includes the following steps:
[0027] Step 1, using a voice signal processor to perform pre-emphasis processing on the voice signal;
[0028] The pre-emphasis processing uses a pre-emphasis coefficient to perform transfer function processing on the voice signal to enhance the high-frequency components in the signal;
[0029] The formula for performing pre-emphasis processing on the voice signal is:
[0030] H(z) = 1 - αz -1
[0031] In the formula, H(z) is the transfer function, z is the voice signal to be processed, and α is the pre-emphasis coefficient;
[0032] Step 2, dividing the voice signal after pre-emphasis processing into several frames, weighting the signal using a Hamming window, and performing filtering processing on the signal through FFT transformation and Gammatone filter bank to smooth the spectrum and eliminate the influence of harmonics, and then performing logarithmic operation and discrete cosine transformation on the processed spectrum to obtain the preprocessed signal;
[0033] The Hamming window is expressed as:
[0034]
[0035] In the formula, N represents the total number of frames into which the pre-emphasized signal is divided, and n represents the number of frames of the signal being currently processed;
[0036] Step 3: Determine the ResCNN network structure and introduce a residual connection layer to optimize the network parameters, enabling the ResCNN network to more effectively extract the features of the speech signal;
[0037] The ResCNN network structure includes a convolutional layer, a pooling layer, and an activation function;
[0038] The residual connection layer introduced in the ResCNN network enhances the model's expressive ability and training efficiency by increasing the number and positions of skip connections;
[0039] The residual layer is expressed by the formula:
[0040] Y = F(X, W i ) + X
[0041] In the formula, X represents the input of the residual block, Y represents the output of the ResCNN network, W i is the weight matrix, and F is the residual function;
[0042] Step 4: Based on the ResCNN network, determine the output value of the GRU model at the current moment, and simultaneously define and calculate the values of the update gate and reset gate of the GRU to control the flow of information;
[0043] The output value of the GRU model at the current moment is obtained by calculating the element-wise multiplication of the weight matrix of the candidate hidden state to capture the temporal dependence relationship in the speech signal;
[0044] Determine the output value of the GRU model, which is expressed by the formula:
[0045]
[0046] In the formula, represents the current moment, P represents the weight matrix of the candidate hidden state, r t is the reset gate of the GRU, * represents element-wise multiplication, and x t represents the input at the current time step of the input sequence;
[0047] Step 5: According to the structure and calculation process of the GRU model, determine the output result of the GRU model, and the output result of this GRU model reflects the recognition situation of the keywords in the speech signal;
[0048] Determine the output result of the GRU model, which is expressed by the formula:
[0049]
[0050] Among them, h t is the final output result;
[0051] Step 6, determine the output layer of the BiGRU network. The output layer of this BiGRU network combines the information of the forward GRU and the backward GRU to obtain the final recognition result of the BiGRU network for the voice signal, and trigger the corresponding alarm mechanism according to this result.
[0052] In the above technical solution, an alarm system based on voice recognition provided by the present invention:
[0053] The embedded body adopted is installed inside the wall, and the front shell is inserted into the inside of the embedded body to realize the hidden installation of the alarm system, making its structure similar to that of the photoelectric switch. At the same time, the camera module is blocked by the sealing mechanism to achieve a better hidden effect of the alarm system;
[0054] The anti-disassembly structure and drive structure adopted can realize the fixed installation between the front shell and the embedded body, prevent unauthorized personnel from arbitrarily disassembling or damaging the alarm system, and play a good protective role for the alarm system. At the same time, the operator can remotely start the drive mechanism to realize the removal of the front shell on the embedded body, and realize the installation and maintenance of the alarm system;
[0055] The present invention uses the ResCNN network, enabling the network to skip certain layers during the training process and directly connect the input of the previous layer to the output of the subsequent layer. This innovative connection method significantly improves the problems of gradient disappearance and gradient explosion that often occur in the training of deep networks. By optimizing gradient propagation, the present invention successfully accelerates the training process of the network and improves the training efficiency;
[0056] The present invention realizes the efficient encoding and understanding of the input sequence through the proposed BiGRU algorithm. This algorithm uses bidirectional gated recurrent units to capture the information of the previous and subsequent statements of each keyword, combines context modeling and long-term dependence modeling, and further enhances the algorithm's ability to model the context-related information of sequence data. This design enables the algorithm to more accurately understand the semantics of keywords and improves the accuracy and efficiency of information processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0058] Figure 1Schematic diagram of an angular structure provided for an embodiment of an alarm system based on voice recognition according to the present invention.
[0059] Figure 2 Schematic diagram of another angular structure provided for an embodiment of an alarm system based on voice recognition according to the present invention.
[0060] Figure 3 Schematic diagram of the open state structure of the closing mechanism provided for an embodiment of an alarm system based on voice recognition according to the present invention.
[0061] Figure 4 Schematic diagram of the state with the embedded body removed provided for an embodiment of an alarm system based on voice recognition according to the present invention.
[0062] Figure 5 Schematic diagram of the closing mechanism provided for an embodiment of an alarm system based on voice recognition according to the present invention Figure 1 。
[0063] Figure 6 Schematic diagram of the closing mechanism provided for an embodiment of an alarm system based on voice recognition according to the present invention Figure 2 。
[0064] Figure 7 Schematic diagram of the driving mechanism provided for an embodiment of an alarm system based on voice recognition according to the present invention.
[0065] Figure 8 Flow chart provided for an embodiment of an alarm system based on voice recognition according to the present invention.
[0066] Explanation of reference numerals:
[0067] 1. Embedded body; 2. Front shell; 21. Camera window; 22. Tongue plate; 3. Voice recognizer; 4. Main control chip board; 41. Voice signal processor; 42. Information exchange module; 5. Closing mechanism; 51. Closing plate; 52. Guide block; 53. Liner; 54. Rack one; 55. Rack two; 56. Driving gear; 57. Driving motor one; 6. Camera module; 61. Camera; 62. Connecting seat; 7. Anti-disassembly structure; 71. Mounting ear seat; 72. Card slot; 73. Elastic support rod; 74. Embedded block; 75. Plug rod; 76. Support arm; 77. Conductive column; 78. Power receiving sleeve; 8. Driving mechanism; 81. Driving motor two; 82. Swing rod; 83. Restraining sleeve; 84. Synchronization frame; 85. Limiting rod; 9. Backup battery. Detailed implementation manners
[0068] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.
[0069] AsFigure 1-7 As shown in Figure 1-7 , an alarm system based on voice recognition provided by an embodiment of the present invention includes an embedding body 1, which is embedded and installed in a groove inside a wall. A face shell 2 is inserted into the embedding body 1. Above the inside of the face shell 2, there is a voice recognizer 3, and the voice recognizer 3 is at least used for collecting sounds in the environment. Inside the face shell 2, there is a main control chip board 4 distributed around the voice recognizer 3. The main control chip board 4 is at least used for processing the sounds collected by the voice recognizer 3. Inside the face shell 2, there is a camera module 6, and the camera module 6 is at least used for real-time video recording when a bullying incident occurs. On the inner side of the face shell 2, there is a closing mechanism 5 for blocking the camera module 6, and the closing mechanism 5 is at least used for blocking the camera module 6. On both inner walls of the embedding body 1, there are anti-disassembly structures 7, and the face shell 2 is inserted into the embedding body 1 through the anti-disassembly structures 7. Inside the embedding body 1, there is a driving mechanism 8, and the driving mechanism 8 is at least used for opening the anti-disassembly structures 7 to realize the disassembly of the face shell 2 from the inside of the embedding body 1.
[0070] An alarm system based on voice recognition, in this embodiment, includes an embedding body 1, and the embedding body 1 is embedded and installed in a groove inside a wall;
[0071] Specifically, on the outer surface of the face shell 2, there is a tongue plate 22 protruding outward;
[0072] Specifically, the outer wall of the embedding body 1 coincides with the outer wall of the tongue plate 22. After the face shell 2 is buckled inside the embedding body 1, the tongue plate 22 is clamped on the front of the embedding body 1, and the two are closely closed.
[0073] In this embodiment, a face shell 2 is inserted into the embedding body 1;
[0074] Specifically, above the front of the face shell 2, there are sound-permeable mesh holes corresponding to the position of the voice recognizer 3, which is convenient for the sounds in the environment to enter the inside of the alarm system to achieve better collection of voices;
[0075] Specifically, in the middle of the face shell 2, there is a camera window 21. After the camera window 21 is blocked, the camera module 6 cannot be seen. When a bullying occurs, the closing mechanism 5 automatically opens, and the camera module 6 can be used for video recording.
[0076] In this embodiment, above the inside of the face shell 2, there is a voice recognizer 3, and the voice recognizer 3 is at least used for collecting sounds in the environment;
[0077] In this embodiment, inside the face shell 2, there is a main control chip board 4 distributed around the voice recognizer 3. The main control chip board 4 is at least used for processing the sounds collected by the voice recognizer 3, and the processing of the voice is specifically introduced in the following operation steps;
[0078] Specifically, the main control chip board 4 includes a voice signal processor 41 and an information exchange module 42. The voice signal processor 41 is at least used to process the collected language, and the information exchange module 42 is at least used for remote transmission of data.
[0079] In this embodiment, a camera module 6 is arranged inside the front shell 2. The camera module 6 is at least used for real-time video recording when a bullying incident occurs;
[0080] Specifically, the camera module 6 includes a camera 61. Connecting seats 62 are arranged on both sides of the camera 61, and the ends of the connecting seats 62 are fixed on the inner wall of the front shell 2.
[0081] In this embodiment, a closing mechanism 5 for blocking the camera module 6 is arranged inside the front shell 2. The closing mechanism 5 is at least used for blocking the camera module 6;
[0082] Specifically, the closing mechanism 5 includes two relatively sliding closing plates 51. Guide blocks 52 are fixedly connected to the top and bottom of the closing plates 51. A lining plate 53 is fixedly arranged on the back of the closing plates 51. A first rack 54 and a second rack 55 connected to the lining plate 53 are respectively arranged below the two closing plates 51. A driving gear 56 is meshed in the middle of the first rack 54 and the second rack 55. And a driving motor 57 is arranged on one side of the driving gear 56. The driving motor 57 is fixed on the inner wall of the front shell 2 through a support rod. Starting the driving motor 57 can drive the driving gear 56 to rotate. By using the driving gear 56, the first rack 54 and the second rack 55 can be driven to move towards or away from each other, so as to realize the opening and closing of the closing plates 51.
[0083] In this embodiment, anti-disassembly structures 7 are arranged on the inner walls on both sides of the embedded body 1. The front shell 2 is inserted into the embedded body 1 through the anti-disassembly structures 7;
[0084] Specifically, the anti-disassembly structure 7 includes a plurality of mounting ear seats 71 fixed on the inner wall of the embedded body 1. A card slot 72 is arranged on the outer wall of the top of the mounting ear seat 71;
[0085] Specifically, a plurality of elastic support rods 73 are fixedly arranged on the back of the front shell 2. And an insertion block 74 inserted into the card slot 72 is arranged below the end of the elastic support rod 73. A plug rod 75 is fixedly arranged on the outer wall of the top of the elastic support rod 73. When installing the front shell 2, the front shell 2 is inserted into the embedded body 1. At this time, the elastic support rod 73 drives the insertion block 74 to be clamped inside the card slot 72, so as to realize the fixation of the front shell 2 inside the embedded body 1 by using the insertion block 74 and the card slot 72. At the same time, the plug rod 75 is inserted into the limiting rod 85 at the end of the synchronous frame 84;
[0086] Specifically, a support arm 76 is provided on one inner wall of the front shell 2, and a conductive column 77 is fixed to the end of the support arm 76. A power receiving sleeve 78 is fixedly arranged inside the embedding body 1. The conductive column 77 is inserted into the power receiving sleeve 78. After the front shell 2 is installed inside the embedding body 1, the conductive column 77 is exactly inserted into the power receiving sleeve 78 to realize the connection of the anti-disassembly circuit. If the power supply of the two is cut off, a corresponding alarm will send out an alarm message.
[0087] In this embodiment, a driving mechanism 8 is arranged inside the embedding body 1. The driving mechanism 8 is at least used to open the anti-disassembly structure 7 to realize the disassembly of the front shell 2 from inside the embedding body 1.
[0088] Specifically, the driving mechanism 8 includes a plurality of restraint sleeves 83 fixed on the inner wall of the embedding body 1, and a synchronous frame 84 is slidably inserted into the restraint sleeves 83. A plurality of limiting rods 85 are arranged on the outer walls on both sides of the synchronous frame 84.
[0089] Specifically, a second driving motor 81 is arranged on the bottom inner wall of the embedding body 1, and a swing rod 82 is connected to the output shaft of the second driving motor 81. The top of the swing rod 82 is in sliding contact with the lower part of the synchronous frame 84. When the front shell 2 is removed from inside the embedding body 1, the operator can remotely control the second driving motor 81 to make the swing rod 82 swing upward, which can push the synchronous frame 84 upward. The synchronous frame 84 drives the limiting rods 85 to rise. At this time, the limiting rods 85 drive one end of the elastic support rod 73 to lift upward, so that the embedding block 74 is pulled out from the inside of the card slot 72. At this time, the front shell 2 can be taken out from inside the embedding body 1.
[0090] In this embodiment, a backup battery 9 is fixedly arranged below the inside of the front shell 2. The backup battery 9 is at least used for power supply when the alarm system loses power.
[0091] As Figure 8 shown, a voice recognition-based alarm method, including the above-mentioned voice recognition-based alarm system, includes the following steps:
[0092] The convolutional neural network (CNN) has demonstrated powerful feature extraction capabilities in speech recognition tasks. Its unique convolutional layer and pooling layer structures enable the CNN to effectively extract key features in the time domain and frequency domain from the original audio signal. These features not only retain the basic attributes of the audio signal but also obtain a more advanced abstract representation through convolutional operations, providing strong support for subsequent classification or recognition tasks.
[0093] In the practical application of speech recognition, recurrent neural networks such as LSTM and GRU play an important role. By introducing memory units, such networks can process sequential data with temporal dependencies. In speech recognition, speech signals often have continuity and temporality, and LSTM and GRU can capture the temporal information and context relevance in these signals, thereby improving the accuracy of speech recognition.
[0094] Specifically, by introducing a gating mechanism, LSTM can selectively retain or forget information at different time steps, thereby effectively modeling the long-term dependencies in speech signals. GRU simplifies the structure of LSTM, reducing the computational complexity while retaining its ability to capture temporal information. Both of these networks perform excellently in speech feature modeling and context modeling, providing strong support for speech recognition systems.
[0095] Deep neural networks (DNNs) play a key role in acoustic modeling. By constructing a multi-layer neural network structure, DNNs can learn complex mapping relationships from audio features to labels or probability distributions. In phoneme classification tasks, DNNs can accurately identify different phonemes in audio signals, providing accurate inputs for speech recognition. At the same time, in the training of acoustic models, DNNs continuously optimize network parameters to improve the model's representation ability for audio signals, thereby further enhancing the performance of speech recognition.
[0096] The present invention proposes an alarm method based on speech recognition, aiming to solve problems such as low accuracy and poor real-time performance in speech keyword recognition in existing campus alarm technologies. This method combines the ResCNN network and the BiGRU model in deep learning to achieve accurate recognition and real-time alarm of speech signals. It extracts features and models sequences from the input speech data, learns the representations and patterns of speech keywords, and detects and recognizes keywords. Once a keyword appears, the system can trigger the corresponding alarm mechanism to alert relevant personnel or systems for further processing or response. The following is a detailed description of the specific steps of this method:
[0097] Step 1: Use the speech signal processor 41 to perform pre-emphasis processing on the speech signal:
[0098] H(z) = 1 - αz -1
[0099] In the formula, H(z) is the transfer function, z is the speech signal to be processed, and α is the pre-emphasis coefficient.
[0100] Pre-emphasis processing is a common preprocessing step in speech signal processing. Its purpose is to enhance the energy of high-frequency components to improve the signal-to-noise ratio of the signal. Pre-emphasis processing is achieved through a transfer function, where α is the pre-emphasis coefficient. By selecting an appropriate pre-emphasis coefficient, the spectral characteristics of the speech signal can be effectively improved, laying a foundation for subsequent feature extraction and keyword recognition.
[0101] Step 2: Divide the signal into several frames:
[0102] To perform time-frequency analysis of the speech signal, the continuous speech signal needs to be divided into several frames. In the present invention, the Hamming window weighting method is used to frame the signal to increase the continuity between adjacent frames. The Hamming window is expressed as:
[0103]
[0104] In the formula, N represents the total number of frames into which the pre-emphasized signal is divided, and n represents the current frame number of the signal being processed.
[0105] The Hamming window is a window function with smooth transition characteristics. The signal weighted by it can retain more useful information. After framing, the FFT transform is used to convert the signal from the time domain to the frequency domain, and then the Gammatone filter bank is used to filter the signal to eliminate harmonic effects and smooth the spectrum. Finally, by performing logarithmic operations and discrete cosine transforms on the processed spectrum, the correlation between noise and feature components is further removed to obtain the preprocessed signal y(t).
[0106] Step 3: Determine the ResCNN network:
[0107] The ResCNN network combines the ideas of deep convolutional neural networks (DCNN) and residual connections, aiming to improve the model's expressive ability and training efficiency. In the present invention, the ResCNN network consists of multiple convolutional layers, pooling layers, and activation functions. By stacking 8 convolutional layers, the deep features of the speech signal can be gradually extracted. Each convolutional layer uses a 3×3 convolutional kernel and sets different numbers of convolutional kernels to adapt to different levels of feature representation. The numbers of convolutional kernels are 32, 43, 128, and 256 respectively. At the same time, a residual connection layer is introduced into the DCNN network. By means of skip connections, the low-level features are directly transmitted to the high-level, which helps to alleviate the problem of gradient disappearance and accelerate the model training process. Introducing a residual connection layer into the DCNN network makes the parameter optimization of the network easier. Among them, the residual layer can be expressed by the formula:
[0108] Y = F(X, W i ) + X
[0109] In the formula, X represents the input of the residual block, Y represents the output of the ResCNN network, and Wi is the weight matrix, and F is the residual function.
[0110] Step 4: Determine the output value of the GRU model at the current moment:
[0111]
[0112] In the formula, represents the current moment, P represents the weight matrix of the candidate hidden state, and r t is the reset gate of the GRU, * represents element-wise multiplication, and x t represents the input at the current time step of the input sequence.
[0113] The GRU model is a recurrent neural network suitable for sequence modeling. It effectively captures long-term dependencies in sequence data by introducing a gating mechanism to control the flow of information. In the present invention, by calculating the update gate and reset gate of the GRU, the update and reset of the hidden state at the current moment can be achieved. The update gate determines the amount of information passed from the previous moment to the current moment, while the reset gate controls the influence degree of the input information at the current moment on the hidden state. This gating mechanism enables the GRU to flexibly process speech sequences of different lengths and extract discriminative feature representations. Among them, the formulas for the update gate and reset gate of the GRU are:
[0114] z t = σ(W z · [h t-1 , x t )
[0115] r t = σ(W r · [h t-1 , x t )
[0116] In the formula, z t is the update gate of the GRU, r t is the reset gate of the GRU, W z is the weight matrix of the update gate, W r is the weight matrix of the reset gate, σ is the Sigmoid function, and z t and r t represent the influence degree from h t-1 to h t in one time step.
[0117] Step 5: Determine the output result of the GRU model:
[0118]
[0119] Among them, h t is the final output result.
[0120] By calculating the final output result of the GRU model, the recognition result of the keyword in the voice signal can be obtained. This output result is jointly determined by the GRU networks in the forward and backward directions, and can fully utilize the bidirectional dependency relationship in the voice signal. By performing threshold judgment or classification processing on the output result, it can be determined whether the keyword appears.
[0121] Step 6: Determine the output layer of the BiGRU network:
[0122]
[0123] In the formula, represents the information obtained by the forward GRU for the i-th voice frame, represents the information obtained by the backward GRU for the i-th voice frame, is the final result obtained by this voice frame through the BiGRU.
[0124] The BiGRU network combines the GRU networks in the forward and backward directions, and can comprehensively capture the timing information in the voice signal. By splicing the output results of the forward and backward GRUs, the final output of the BiGRU network can be obtained. This output layer not only contains the deep feature representation of the voice signal, but also reflects the timing position information of the keyword. By decoding or post-processing the output layer, accurate recognition and positioning of the keyword can be achieved.
[0125] After determining the output layer of the BiGRU network, the present invention also uses the corresponding evaluation index to evaluate the effect of the model. The evaluation index is:
[0126]
[0127] In the formula, F1 is the evaluation index, Precision is the precision rate, and Recall is the recall rate.
[0128] The evaluation index includes the precision rate and the recall rate, etc., and is used to measure the performance of the model in the voice keyword recognition task. By adjusting the model parameters and structure, the performance of the model can be further optimized, and the accuracy and real-time performance of keyword recognition can be improved.
[0129] In summary, the present invention combines the ResCNN network and the BiGRU model to implement an efficient and accurate voice keyword recognition campus alarm method. This method can real-time monitor the voice signal and recognize the keyword therein. Once the keyword appears, the alarm mechanism can be triggered to remind relevant personnel or systems to perform further processing or response. This method has broad application prospects and practical value in the field of campus security monitoring.
[0130] (1) The present invention ingeniously introduces the ResCNN technology to perform delicate operations on the preprocessed speech frames. The core feature of the ResCNN network lies in that it allows the network to skip certain layers during the training process and directly connect the input of the previous layer to the output of the subsequent layer. This unique connection method not only optimizes the network structure but also significantly improves the network performance.
[0131] In practical applications, deep networks often face the problems of vanishing gradients and exploding gradients, which seriously restrict the training effect and speed of the network. However, through the layer skipping mechanism of the ResCNN network, the present invention successfully solves this problem. It enables the gradients to propagate more smoothly in the network, thus avoiding the phenomena of vanishing gradients and exploding gradients. This not only improves the training efficiency of the network but also enables the network to converge to the optimal solution more quickly.
[0132] Therefore, by introducing the ResCNN technology, the present invention realizes the efficient processing of speech frames, accelerates the training process of the network, and lays a solid foundation for subsequent speech recognition and speech analysis tasks.
[0133] (2) The present invention proposes the BiGRU algorithm. This algorithm uses bidirectional gated recurrent units (BiGRU), providing a new idea for the processing and analysis of sequential data.
[0134] Bidirectional gated recurrent units can capture both forward and backward information in sequential data, thus more comprehensively understanding the position and role of each keyword in the context. This feature enables the algorithm to more accurately capture the information of the sentences before and after the keyword, providing rich context basis for subsequent semantic understanding.
[0135] Therefore, through the application of the BiGRU algorithm, the present invention realizes the efficient encoding and understanding of sequential data. It can not only obtain global features but also more deeply understand the semantics of information keywords, providing strong support for subsequent speech analysis and processing tasks.
[0136] Only some exemplary embodiments of the present invention are described by way of illustration above. Undoubtedly, for those of ordinary skill in the art, without departing from the spirit and scope of the present invention, the described embodiments can be modified in various different ways. Therefore, the above drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
Claims
1. An alarm system based on speech recognition, comprising an embedding body (1), and the embedding body (1) is embedded and installed in a groove inside a wall body, characterized in that, A face shell (2) is inserted inside the embedding body (1). Above the inside of the face shell (2), a voice recognizer (3) is provided, and the voice recognizer (3) is at least used for collecting sounds in the environment; Inside the face shell (2), a main control chip board (4) is provided which is distributed around the voice recognizer (3); the main control chip board (4) is at least used for processing the sounds collected by the voice recognizer (3); Inside the face shell (2), a camera module (6) is provided, and the camera module (6) is at least used for real-time video recording when a bullying incident occurs. Inside the face shell (2), a closing mechanism (5) for blocking the camera module (6) is provided, and the closing mechanism (5) is at least used for blocking the camera module (6); Anti-disassembly structures (7) are provided on both inner walls of the embedding body (1), and the face shell (2) is inserted inside the embedding body (1) through the anti-disassembly structures (7); A driving mechanism (8) is provided inside the embedding body (1), and the driving mechanism (8) is at least used for opening the anti-disassembly structures (7) to realize the disassembly of the face shell (2) from inside the embedding body (1).
2. The alarm system based on speech recognition according to claim 1, characterized in that, Above the front of the face shell (2), sound-permeable mesh holes corresponding to the position of the voice recognizer (3) are provided; A camera window (21) is provided in the middle of the face shell (2).
3. An alarm system based on speech recognition according to claim 1, wherein On the outer surface of the face shell (2), a tongue plate (22) protruding outwards is provided; The outer wall of the embedding body (1) coincides with the outer wall of the tongue plate (22).
4. The alarm system based on voice recognition according to claim 1, characterized in that, Below the inside of the face shell (2), a backup battery (9) is fixedly provided, and the backup battery (9) is at least used for power supply when the alarm system is powered off.
5. An alarm system based on speech recognition according to claim 1, characterized in that, The main control chip board (4) includes a voice signal processor (41) and an information exchange module (42). The voice signal processor (41) is at least used for processing the collected language, and the information exchange module (42) is at least used for remote transmission of data.
6. The alarm system based on speech recognition according to claim 1, characterized in that, The closing mechanism (5) includes two relatively sliding closing plates (51). At the top and bottom of the closing plates (51), guide blocks (52) are fixedly connected. On the back of the closing plates (51), a lining plate (53) is fixedly provided. Below the two closing plates (51), a first rack (54) and a second rack (55) connected to the lining plate (53) are respectively provided. In the middle of the first rack (54) and the second rack (55), a driving gear (56) is meshed, and on one side of the driving gear (56), a first driving motor (57) is provided. The first driving motor (57) is fixed on the inner wall of the face shell (2) through a support rod.
7. An alarm system based on speech recognition according to claim 1, characterized in that, The camera module (6) includes a camera (61). On both sides of the camera (61), connecting seats (62) are provided, and the ends of the connecting seats (62) are fixed on the inner wall of the face shell (2).
8. The alarm system based on voice recognition according to claim 1, characterized in that, The anti-disassembly structure (7) includes a plurality of mounting ear seats (71) fixed on the inner wall of the embedding body (1). On the outer wall of the top of the mounting ear seats (71), a card slot (72) is provided; A plurality of elastic support rods (73) are fixedly arranged on the back surface of the front shell (2), and an embedding block (74) inserted into the internal of the clamping groove (72) is arranged below the end of the elastic support rod (73), and a plug rod (75) is fixedly arranged on the outer wall of the top of the elastic support rod (73); A support arm (76) is arranged on the inner wall of one side of the front shell (2), and a conductive column (77) is fixed at the end of the support arm (76). A power receiving sleeve (78) is fixedly arranged inside the embedding body (1), and the conductive column (77) is inserted into the internal of the power receiving sleeve (78).
9. An alarm system based on speech recognition according to claim 1, characterized in that, The driving mechanism (8) includes a plurality of restraint sleeves (83) fixed on the inner wall of the embedding body (1), and a synchronous frame (84) is slidably inserted into the internal of the restraint sleeve (83). A plurality of limiting rods (85) are arranged on the outer walls of both sides of the synchronous frame (84); A driving motor two (81) is arranged on the inner wall of the bottom of the embedding body (1), and a swing rod (82) is connected to the output shaft of the driving motor two (81). The top end of the swing rod (82) is in sliding contact with the lower part of the synchronous frame (84).
10. An alarm method based on speech recognition, comprising an alarm system based on speech recognition according to any one of claims 1-9, characterized in that, It includes the following steps: Step 1, using a voice signal processor (41) to perform pre-emphasis processing on the voice signal; The pre-emphasis processing uses a pre-emphasis coefficient to perform transfer function processing on the voice signal to enhance the high-frequency components in the signal; The formula for performing pre-emphasis processing on the voice signal is: H(z) = 1 - αz -1 In the formula, H(z) is the transfer function, z is the voice signal to be processed, and α is the pre-emphasis coefficient; Step 2, dividing the voice signal after pre-emphasis processing into several frames, weighting the signal using a Hamming window, and performing filtering processing on the signal through FFT transformation and Gammatone filter bank to smooth the spectrum and eliminate the influence of harmonics. Subsequently, performing logarithmic operation and discrete cosine transformation on the processed spectrum to obtain the pre-processed signal; The Hamming window is expressed as: In the formula, N represents the total number of frames into which the signal after pre-emphasis is divided, and n represents the number of frames of the signal currently being processed; Step 3, determining the ResCNN network structure and introducing a residual connection layer to optimize the network parameters, so that the ResCNN network can more effectively extract the features of the voice signal; The ResCNN network structure includes a convolutional layer, a pooling layer, and an activation function; The residual connection layer introduced in the ResCNN network improves the expression ability and training efficiency of the model by increasing the number and position of skip connections; The residual layer is expressed by the formula: Y = F(X, W i ) + X Wherein, X represents the input of the residual block, Y represents the output of the ResCNN network, and W i is the weight matrix, and F is the residual function; Step 4, based on the ResCNN network, determining the output value of the GRU model at the current moment, and simultaneously defining and calculating the values of the update gate and reset gate of the GRU to control the flow of information; The output value of the GRU model at the current moment is obtained by calculating the weight matrix of the candidate hidden state and performing element-wise multiplication to capture the temporal dependence relationship in the voice signal; Determining the output value of the GRU model at the current moment is expressed by the formula: In the formula, represents the current moment, P represents the weight matrix of the candidate hidden state, and r t is the reset gate of the GRU, * represents element-wise multiplication, and x t represents the input at the current time step of the input sequence; Step 5, according to the structure and calculation process of the GRU model, determining the output result of the GRU model, and the output result of the GRU model reflects the recognition situation of the keywords in the voice signal; Determining the output result of the GRU model is expressed by the formula: Among them, h t is the final output result; Step 6, determine the output layer of the BiGRU network. The output layer of this BiGRU network combines the information of the forward GRU and the backward GRU to obtain the final recognition result of the BiGRU network for the voice signal, and triggers the corresponding alarm mechanism according to this result.
Citation Information
Patent Citations
Toilet voice recognition alarm system and recognition method thereof
CN116863933A