A method for identifying baby crying sounds based on neural network
Through the neural network-based infant cry recognition method, combined with long-term feature fusion and model comprehensive voting technology, the problem of inaccurate identification in the existing technology under complex background noise is solved, and efficient recognition is achieved on devices with weak computing power, improving the accuracy and efficiency of recognition.
Patent Information
- Application Number
- CN202111419882.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-26
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-11-26
AI Technical Summary
The existing infant cry recognition method is difficult to achieve accurate and rapid recognition under complex background noise, and the existing deep learning models require high processor computing power and are difficult to run on devices with weak computing power.
The infant crying recognition method based on neural network is used to obtain different types of infant crying audio and non-infant crying audio for preprocessing and fbank feature extraction. Combining long-term feature fusion and model comprehensive voting technology, fixed-point models can be trained that can run on devices with weak computing power, and the recognition accuracy can be improved through silent detection and online learning.
It realizes accurate recognition of baby crying under complex background noise, reduces the requirements for processor computing power, can operate effectively on devices with weak computing power, and improves the accuracy and efficiency of recognition.
Smart Images

Figure CN114155833B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of sound detection, and in particular to a method for recognizing baby crying sounds based on a neural network. Background Art
[0002] Guardians of infants always hope to better understand the mental state of their children. Since infants have weak language skills, they mainly use crying to express their discomfort, such as feelings such as fear and loneliness, and physical feelings such as hunger and pain. Therefore, if the mental state of infants can be understood in time through the crying of infants, it is possible to better care for and comfort infants and enable them to grow up healthily. However, since guardians are not always with infants, such as when they need to work, do housework, prepare food, etc., they often coax infants to sleep or put them in a safe area to play. However, the crying of infants may not be heard by guardians due to background noise and other reasons. In view of this, if there is a way to accurately and quickly identify the crying of infants, guardians can immediately understand the needs of infants through baby monitors or similar simple sound transmission devices.
[0003] Traditional baby cry recognition is mainly based on the sound quality, frequency and other properties of different sounds. After filtering, the sound is classified or recognized. It is difficult to distinguish accurately and quickly in the case of complex background noise, and it is easy to be affected by the environment and miss false alarms. In view of this, some researchers have adopted deep learning methods to train the recognition system to intelligently analyze the characteristics of the sound, and detect baby cries through deep learning classification or audio feature extraction. For example, a patent application with publication number CN112185364A discloses a method for detecting baby cries. The sound of baby cries is obtained by training based on a deep learning model, but its model is too complex and has high requirements on the processor computing power. It is difficult to load it onto the low-cost baby monitor device with weak computing power; some detection methods have the disadvantage that the sound features selected for training are too simple. Although they can ensure a faster speed, they cannot guarantee accuracy.
[0004] Therefore, it is necessary to improve and optimize the existing baby crying recognition methods. Summary of the invention
[0005] The purpose of the present invention is to overcome the above-mentioned deficiencies in the prior art and to provide a method for identifying baby crying based on a neural network. The method can train the system to accurately identify baby crying based on a simple neural network.
[0006] The technical solution adopted by the present invention to solve the above problem is: a method for recognizing baby crying based on neural network, characterized in that the steps are as follows:
[0007] Step 1: Obtain different types of baby crying audio and different types of non-baby crying audio;
[0008] Step 2: Set the baby crying audio as the positive sample and the non-baby crying audio as the negative sample, and pre-process the positive sample and negative sample audio data;
[0009] Step 3: Extract fbank features from the preprocessed positive and negative sample audios;
[0010] Step 4: For each type of audio, during the fbank feature extraction process, the time domain dimension and frequency domain dimension of the feature are randomly set to zero to enhance the audio;
[0011] Step 5: Fusion of long- and short-time features: Extract feature data from audio clips of different lengths for each type of audio, and fuse the feature data corresponding to each length to obtain a training model;
[0012] Step 6: Place all training models into the neural network for training, and use the model comprehensive voting technology to perform weighted averaging of different training models to obtain the final prediction model;
[0013] Step 7: Convert the prediction model from a floating-point model to a fixed-point model for easy embedded transplantation;
[0014] Step 8: Use the microphone to collect ambient audio data and determine whether the ambient audio data is silent. If it is not silent, extract features from it. Then use the fixed-point model obtained in step 7 to determine whether there is a baby crying. If there is a baby crying, use the fixed-point model to determine the type of baby crying.
[0015] Step 9: Use different methods to comfort babies according to their crying patterns.
[0016] Preferably, in the step 2, preprocessing refers to unifying the positive sample and negative sample audio data into: wav format, mono, 16KHz sampling, 16-bit expression.
[0017] Preferably, in step three, the steps of the fbank feature extraction method are: pre-emphasis processing, frame processing, window filtering and short-time Fourier transform.
[0018] Preferably, in step five, three audio clips of 100ms, 50ms and 25ms in length are taken for each type of audio, and fbank feature extraction operations are performed on them respectively. Each audio clip of the same length can obtain 128-dimensional feature data. Then, the 128-dimensional feature data of the three lengths are fused to obtain 384-dimensional features and put into the neural network for training to obtain a training model. This operation can improve the prediction accuracy of the test set.
[0019] Preferably, in step six, when the training model is placed in a neural network for training, the training model is made to learn a baby's cry online. If the features of the collected baby's cry are similar to the previous features, the confidence of this cry category is improved, wherein whether the crying features are similar is represented by the cosine distance of the feature vector of the fully connected layer.
[0020] Preferably, in step eight, the mute judgment method is: averaging 200 frames of audio in the frequency domain, taking the average value + 1.5 as a threshold, and if none of the 200 frames of audio is greater than the threshold, it is judged to be a mute state.
[0021] Compared with the prior art, the present invention has the following advantages and effects:
[0022] 1. Long- and short-time fusion: In the feature extraction process, the audio frame length is divided into 100ms, 50ms and 25ms, and the fbank feature extraction operation is performed on them respectively. Each audio of different lengths obtains 128-dimensional feature data. Then the 128-dimensional features are fused to obtain 384-dimensional features and put into the neural network for training. This operation can improve the prediction accuracy of the test set;
[0023] 2. Silence detection: The current mature silence detection is a piece of code cut out from WebRTC and related to silence detection (VAD). However, due to its relatively large code complexity, it is necessary to load a lot of static libraries and dynamic libraries. The present invention designs a simple silence detection method, and the effect can meet the requirements through testing. The method is to calculate the average value of 200 frames of audio in the frequency domain, and use this average value + 1.5 as the threshold. If none of the 200 frames of audio is greater than this threshold, it is determined to be in silence. This method can greatly reduce the consumption of CPU resources while ensuring the detection effect;
[0024] 3. Model comprehensive voting: At a certain moment, different lengths of audio are intercepted and different models are trained through feature extraction. The final crying category is obtained through weighted average of different models;
[0025] 4. Online learning of a baby’s cry: If the current cry feature is very similar to the previously recognized cry feature, the confidence of this recognition is improved. Here, whether the cry feature is similar is represented by the cosine distance of the feature vector of the fully connected layer. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the specific implementation methods of the present invention or the solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.
[0027] Figure 1 4 is a flow chart of a method for recognizing a baby's cry according to an embodiment of the present invention. DETAILED DESCRIPTION
[0028] The present invention will be further described in detail below with reference to the accompanying drawings and by way of examples. The following examples are intended to explain the present invention but the present invention is not limited to the following examples.
[0029] Example
[0030] See also Figure 1 .
[0031] This embodiment discloses a method for recognizing baby crying sounds based on a neural network, and the steps are as follows:
[0032] Step 1: You can obtain different types of baby crying audio and different types of non-baby crying audio from the freesound website;
[0033] Step 2: Set the baby crying audio as the positive sample and the non-baby crying audio as the negative sample, and pre-process the positive and negative sample audio data. The pre-processing here means unifying the positive and negative sample audio data into: wav format, mono, 16KHz sampling, 16-bit expression;
[0034] Step 3: Perform fbank feature extraction on the preprocessed positive sample and negative sample audio. The steps of the fbank feature extraction method are: pre-emphasis processing, frame processing, window filtering and short-time Fourier transform;
[0035] Step 4: For each type of audio, during the fbank feature extraction process, the time domain dimension and frequency domain dimension of the feature are randomly set to zero to enhance the audio;
[0036] Step 5: Fusion of long- and short-time features: For each type of audio, take audio clips of 100ms, 50ms, and 25ms in length, perform fbank feature extraction on them respectively, and obtain 128-dimensional feature data for each audio clip of the same length. Then, the 128-dimensional feature data of the three lengths are fused to obtain 384-dimensional features and put into the neural network for training to obtain the training model. This operation can improve the prediction accuracy of the test set;
[0037] Step 6: All training models are placed in the neural network for training, and the final prediction model is obtained by weighted averaging different training models through model comprehensive voting technology; when the training model is placed in the neural network for training, the training model is made to learn a baby's cry online. If the characteristics of the collected baby's cry are similar to the previous characteristics, the confidence of this cry category is improved. Whether the crying characteristics are similar is represented by the cosine distance of the feature vector of the fully connected layer;
[0038] Step 7: Convert the prediction model from a floating-point model to a fixed-point model for easy embedded transplantation;
[0039] Step 8: Collect environmental audio data using a microphone, and determine whether the environmental audio data is silent. If it is not silent, extract features from it, and then use the fixed-point model obtained in step 7 to determine whether there is a baby crying. If there is a baby crying, use the fixed-point model to determine the type of baby crying. The silent judgment method is: average the 200 frames of audio in the frequency domain, and use the average value + 1.5 as a threshold. If no frame of audio in the 200 frames of audio is greater than the threshold, it is determined to be silent.
[0040] Step 9: Use different methods to comfort babies according to their crying patterns.
[0041] Although the present invention has been disclosed as above by way of embodiments, it is not intended to limit the protection scope of the present invention. Any changes and modifications made by any technician familiar with the technology without departing from the concept and scope of the present invention should fall within the protection scope of the present invention.
Claims
1. A method for recognizing baby crying based on neural network, characterized in that: Here are the steps: Step 1: Obtain different types of baby crying audio and different types of non-baby crying audio; Step 2: Set the baby crying audio as the positive sample and the non-baby crying audio as the negative sample, and pre-process the positive sample and negative sample audio data; Step 3: Extract fbank features from the preprocessed positive and negative sample audios; Step 4: For each type of audio, during the fbank feature extraction process, the time domain dimension and frequency domain dimension of the feature are randomly set to zero to enhance the audio; Step 5: Fusion of long- and short-time features: For each type of audio, take audio clips of 100ms, 50ms, and 25ms in length, perform fbank feature extraction on them respectively, and each audio clip of the same length can obtain 128-dimensional feature data. Then, the 128-dimensional feature data of the three lengths are fused to obtain 384-dimensional features and put into the neural network for training to obtain a training model; Step 6: Place all training models into the neural network for training, and use the model comprehensive voting technology to perform weighted averaging of different training models to obtain the final prediction model; Step 7: Convert the prediction model from a floating-point model to a fixed-point model for easy embedded transplantation; Step 8: Use the microphone to collect ambient audio data and determine whether the ambient audio data is silent. If it is not silent, extract features from it. Then use the fixed-point model obtained in step 7 to determine whether there is a baby crying. If there is a baby crying, use the fixed-point model to determine the type of baby crying. Step 9: Use different methods to comfort babies according to their crying patterns.
2. The method for identifying infant crying sound based on neural network according to claim 1, characterized in that: In the step 2, preprocessing refers to unifying the positive sample and negative sample audio data into: wav format, mono, 16KHz sampling, 16-bit expression.
3. The method for identifying infant crying sound based on neural network according to claim 1, characterized in that: In the step three, the steps of the fbank feature extraction method are: pre-emphasis processing, frame processing, window filtering and short-time Fourier transform.
4. The method for identifying infant crying sound based on neural network according to claim 1, characterized in that: In step five, for each type of audio, three audio clips of 100ms, 50ms and 25ms are taken, and fbank feature extraction operations are performed on them respectively. Each audio clip of the same duration can obtain 128-dimensional feature data. Then, the 128-dimensional feature data of the three durations are fused to obtain 384-dimensional features and put into the neural network for training to obtain a training model. This operation can improve the prediction accuracy of the test set.
5. The method for identifying infant crying sound based on neural network according to claim 1, characterized in that: In step six, when the training model is placed in the neural network for training, the training model is made to learn a baby's cry online. If the features of the collected baby's cry are similar to the previous features, the confidence of this cry category is improved. Whether the crying features are similar is represented by the cosine distance of the feature vector of the fully connected layer.
6. The method for identifying infant crying sound based on neural network according to claim 1, characterized in that: In step eight, the mute judgment method is: find the average value of 200 frames of audio in the frequency domain, and use the average value + 1.5 as a threshold. If no frame of audio in the 200 frames of audio is greater than the threshold, it is determined to be a mute state.
Citation Information
Patent Citations
Baby crying sound translation method based on sound feature recognition
CN109065034A
Baby crying detection method and device
CN112185364A