Crying sound detection model training method, device, electronic device and storage medium

By obtaining the training data set containing audio level and segment level labels, the cry detection model is jointly trained in the cry detection model, which solves the problems of insufficient data sets and noise instability in the prior art, and improves the accuracy and accuracy of cry detection.

CN114550729BActive Publication Date: 2025-05-16EEASY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210075516.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-22
Publication Date
2025-05-16
Estimated Expiration
2042-01-22

AI Technical Summary

Technical Problem

The cry detection model in existing intelligent monitoring devices has low detection accuracy and accuracy due to insufficient data sets and unstable environmental noise.

Method used

By obtaining a training data set containing audio-level and segment-level tags, the pre-constructed cry detection model is trained in joint dual-label. Using a structure combined with the backbone network and the dual-label network, features are extracted and masked to improve the generalization ability of the model.

Benefits of technology

While taking into account the insufficient data set, the training effect of the cry detection model is improved, and the accuracy and accuracy of the detection are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114550729B_ABST
    Figure CN114550729B_ABST
Patent Text Reader

Abstract

The present invention is applicable to the field of audio processing technology, and provides a training method, device, electronic device and storage medium for a crying detection model. The method comprises: obtaining a training data set, wherein the training data set comprises an audio-level label, a segment-level label and an audio feature corresponding to each training sample, and using the training data set to train a pre-built crying detection model to obtain a trained crying detection model, thereby jointly training the crying detection model with dual labels, thereby improving the training effect of the crying detection model while taking into account the insufficient data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of audio processing, and in particular relates to a training method, device, electronic equipment and storage medium for a crying detection model. Background Art

[0002] For ordinary families, baby care requires a lot of manpower, material resources and financial resources. The emergence of smart baby care equipment can effectively reduce the burden on caregivers.

[0003] The existing sound detection models of smart monitoring devices have low accuracy and precision in the training process because there are few open source crying data sets and the number of sounds is limited, making it difficult to collect them. In addition, the uncertainty and instability of noise in the actual environment lead to low accuracy and precision in baby crying detection. Summary of the invention

[0004] The purpose of the present invention is to provide a crying detection model training method, device, electronic device and storage medium, aiming to solve the problem that the sound detection model trained based on a small amount of data set in the prior art has insufficient detection performance.

[0005] In one aspect, the present invention provides a method for training a crying detection model, the method comprising the following steps:

[0006] Obtain a training data set, wherein the training data set includes an audio-level label, a segment-level label, and an audio feature corresponding to each training sample;

[0007] The pre-built crying detection model is trained using the training data set to obtain a trained crying detection model.

[0008] Optionally, the step of obtaining a training data set includes:

[0009] Generate a double label for each of the training samples, and extract audio features of each of the training samples, wherein the double label includes a label at the audio level and a label at the segment level.

[0010] Optionally, the step of generating a double label for each training sample includes:

[0011] Generate the audio level label according to whether the training sample contains crying sound;

[0012] According to the preset speech framing parameters, the preset input feature dimension downsampling multiples and the preset length of the training sample, the audio segment corresponding to each segment-level label is obtained, and the segment-level label is generated based on the comparison result of the number of audio sampling points corresponding to the crying in each of the audio segments with the preset sampling point number threshold.

[0013] Optionally, before the step of extracting the audio features of each of the training samples, the step further includes:

[0014] Performing data augmentation on each of the training samples;

[0015] The step of extracting the audio features of each of the training samples further includes:

[0016] Extract audio features of each training sample after data augmentation.

[0017] Optionally, the audio feature is an Fbank feature.

[0018] Optionally, the crying detection model includes a backbone network and a dual-label network connected in sequence, and the dual-label network includes a segment-level network and an audio-level network connected in sequence.

[0019] Optionally, the segment-level network includes a one-dimensional convolutional layer and an activation layer connected in sequence, and the audio-level network includes a pooling layer.

[0020] Optionally, the step of training the pre-built crying detection model includes:

[0021] The feature map extracted by the backbone network is subjected to mask processing, and the crying detection model is trained based on the feature map after the mask processing.

[0022] Optionally, the step of performing mask processing on the feature map extracted by the backbone network includes:

[0023] Determine whether to perform mask processing on the feature map extracted from the current layer according to a preset first probability value;

[0024] If so, obtaining a plurality of channels to be subjected to mask processing;

[0025] The plurality of channels are subjected to mask processing according to a preset second probability value and a preset mask area size.

[0026] Optionally, the pooling layer is a temporal pooling layer.

[0027] Optionally, the pooling function used by the temporal pooling layer is as follows:

[0028] pool=sum(pred_seg^2) / sum(pred_seg), where pool represents the pooling result, sum represents the summation operation, pred_seg represents the input of the temporal pooling layer, and ^2 represents square.

[0029] Optionally, the step of using the training data set to train a pre-built crying detection model includes:

[0030] The crying detection model is jointly trained according to the loss of the segment level network and the loss of the audio level network.

[0031] Optionally, the loss function used by the crying detection model is as follows:

[0032] loss=loss1+0.1*loss2, where loss represents the total loss of the crying detection model, loss1 represents the loss of the audio level network, and loss2 represents the loss of the segment level network.

[0033] In another aspect, the present invention provides a training device for a crying detection model, the device comprising:

[0034] A training set acquisition unit, configured to acquire a training data set, wherein the training data set includes an audio level label, a segment level label, and an audio feature corresponding to each training sample; and

[0035] The model training unit is used to train the pre-built crying detection model using the training data set to obtain a trained crying detection model.

[0036] On the other hand, the present invention further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.

[0037] On the other hand, the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0038] The present invention obtains a training data set, which includes audio-level labels, segment-level labels and audio features corresponding to each training sample, and uses the training data set to train a pre-built crying detection model to obtain a trained crying detection model, thereby jointly training the crying detection model with dual labels, thereby improving the training effect of the crying detection model while taking into account the insufficient data set. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1A is a flow chart of the implementation of the training method of the crying detection model provided in the first embodiment of the present invention;

[0040] Figure 1B This is an example diagram of a double label provided in the first embodiment of the present invention;

[0041] Figure 1C is a structural example diagram of a crying detection model provided in Embodiment 1 of the present invention;

[0042] Figure 2 is a structural schematic diagram of a crying detection model training device provided in Embodiment 2 of the present invention; and

[0043] Figure 3 It is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION

[0044] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0045] The specific implementation of the present invention is described in detail below in conjunction with specific embodiments:

[0046] Embodiment 1:

[0047] Figure 1A The implementation process of the training method of the crying detection model provided in the first embodiment of the present invention is shown. For the convenience of explanation, only the part related to the embodiment of the present invention is shown, which is described in detail as follows:

[0048] In step S101, a training data set is obtained, where the training data set includes an audio-level label, a segment-level label, and an audio feature corresponding to each training sample.

[0049] The embodiments of the present invention are applicable to electronic devices, which may be mobile phones, speakers, tablet computers, wearable devices, vehicle-mounted devices, surveillance cameras, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPC), netbooks, personal digital assistants (PDA) and other terminal devices. The embodiments of the present application do not impose any restrictions on the specific types of electronic devices.

[0050] In an embodiment of the present invention, when obtaining a training data set, the original audio data set may be obtained first, and each original audio in the original audio data set may be preprocessed to obtain a plurality of training samples of preset length. The above-mentioned original audio data set may be obtained based on existing open source data, or may be obtained through actual recording. Since the open source data of crying is limited, the total duration is short, and the actual scene recording is also difficult, the crying audio obtained in the above manner is usually limited. Considering that the crying detection model only accepts fixed-length input during the deployment process, in order to ensure the data consistency of training and application, the audio input for the training of the crying detection model also needs to adopt fixed-length input. Therefore, the above-mentioned preprocessing may include audio sampling processing and fixed-length processing. Specifically, the original audio may be sampled according to a preset audio sampling rate, and the above-mentioned original audio may be slidably intercepted or padded using a preset length data window to obtain a plurality of training samples of preset length, that is, the audio with a duration of more than 10s is truncated, and the audio with a duration of less than 10s is padded. As an example, among the multiple original audios contained in the original audio data set, the length of each audio segment varies from 5s to 30min, the audio sampling rate is 16k, and each training sample after preprocessing is an audio file in wav format with a length of 10s and a sampling rate of 16k. Among them, for audio with a duration of less than 10s, the values ​​of the sampling points of the part less than 10s are padded with 0.

[0051] After obtaining the above-mentioned multiple training samples, a label for each training sample can be generated. Considering that it is difficult to obtain frame-level labels and it is difficult to avoid the negative impact of manual labeling errors on detection performance, and the audio-level labels require a large amount of audio to achieve the effect at the frame level, optionally, a double label for each training sample is generated, and the audio features of each training sample are extracted to obtain a training data set based on the double labels and audio features of each training sample. Among them, the double labels include audio-level labels and segment-level labels, and the audio features can be Fbank features. Of course, the above-mentioned audio features can also be features other than the above, such as MFCC (Mel-frequency cepstral coefficients) features. As an example, if the speech frame length is 40ms, the frame shift is 20ms, and the preset audio length is 10s, then a segment of audio has 501 frames. If the number of Mel filters is 128, the feature dimension of the input feature map of the crying detection model is [batch, 1, 501, 128]. Among them, batch is a hyperparameter, that is, the number of samples selected for one training, 1 represents the number of channels of the input feature map, 501 represents the time feature dimension of the input feature map, and 128 represents the frequency feature dimension of the input feature map.

[0052] When generating dual labels for each training sample, optionally, generate audio-level labels based on whether the training sample contains crying sounds; obtain the audio segment corresponding to each segment-level label based on the preset speech framing parameters, the preset input feature dimension downsampling multiples, and the preset length of the training sample, and generate segment-level labels based on the comparison results of the number of audio sampling points corresponding to crying sounds in each audio segment and the preset sampling point number threshold, so as to fully consider the time for breathing during crying, thereby achieving dual label generation while improving the accuracy and generation efficiency of segment-level labels, thereby improving the training effect of the subsequent crying detection model. Among them, the speech framing parameters may include frame length and frame shift. Among them, the above-mentioned downsampling multiple of the input feature dimension is determined according to the following crying detection model. Specifically, the downsampling multiple of the input feature dimension can be determined according to the ratio of the time dimension of the input feature of the backbone network of the following crying detection model to the time dimension of the output feature, that is, the downsampling multiple of the input feature dimension = the time dimension of the input feature of the backbone network / the time dimension of the output feature of the backbone network. More specifically, the downsampling multiple of the input feature dimension can be determined according to the number of pooling layers of the backbone network and the step size of each pooling layer. The downsampling multiple of the input feature dimension is specifically the product of the step sizes of all pooling layers.

[0053] For example, the duration of the training sample is 10s, the crying sound occurs at 0-4.58s, 5.11-5.68s, 7.24s-10s, and the audio sampling rate is 16k, then:

[0054] (a) For audio level labels:

[0055] If the training sample does not contain crying, the audio level label is 0. If the training sample contains crying, the audio level label is 1. Since the training sample contains crying, its audio level label is 1. Figure 1B As shown in the figure above;

[0056] (b) For segment-level labels:

[0057] b-1) The initialized segment-level labels are: 0, 81760, 115840 | 73279, 90879, 15999. The crying sound is divided into three segments. The starting time of each segment is 0s, 5.11s, and 7.24s, and the corresponding audio sampling points are 0, 81760, and 115840. "|" represents the separator. Similarly, the end time of each segment is 4.58s, 5.68s, and 10s, and the corresponding sampling points are 73279, 90879, and 15999.

[0058] b-2) Generation of segment-level labels: If the speech frame length is 40ms and the frame shift is 20ms, the crying division needs to take into account the baby's breathing time when marking the initial segment-level labels. Therefore, the threshold of the number of sampling points is 1600. The above preset input feature dimension downsampling multiple is 16. The audio segment has 501 frames in total and 31 labels in the time dimension, that is, each label corresponds to 5120 sampling points on average. Therefore, when the number of sampling points corresponding to crying in the 5120 sampling points corresponding to each label meets the above sampling point number threshold of 1600, the corresponding label is 1, otherwise the corresponding label is 0. The basic labels of the obtained segments are [1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,0,0,0,0,1,1,1,1,1,1,1,1,1,1]. For the convenience of display, the segment-level labels are repeatedly sampled into the original audio length, such as Figure 1B As shown in the figure below.

[0059] It is explained here that the speech frame length of an audio signal is usually between 20ms and 50ms, and the segment-level labels described in this embodiment are non-frame-level labels. On the one hand, the generation of the segment-level labels saves the time of manual labeling, and compared with the frame-level labels, the segment-level labels are more conducive to the convergence of the crying detection model; on the other hand, considering that the crying of infants is generally continuous, the crying detection model described in this embodiment does not require overly sophisticated frame-level labels.

[0060] Optionally, before extracting the audio features of each training sample, data augmentation is performed on each training sample, and then the audio features of each training sample after data augmentation are extracted, and a training set is obtained based on the double labels and audio features of each training sample after data augmentation to further improve the training effect of the subsequent crying detection model. Considering that the background noise of the crying data of the original audio may be low, the above-mentioned data augmentation may include noise addition. Of course, data augmentation methods other than noise addition may also be used according to the actual scenario, such as adjusting the speed, adjusting the frequency, adjusting the volume, etc.

[0061] In step S102, the pre-built crying detection model is trained using the training data set to obtain a trained crying detection model.

[0062] In an embodiment of the present invention, the crying detection model includes a backbone network and a dual-label network connected in sequence, and the dual-label network includes a segment-level network and an audio-level network connected in sequence. The backbone network is used for feature extraction, the segment-level network is used to predict the posterior probability of the segment level based on the feature extraction result, and the audio-level network is used to predict the posterior probability of the audio level based on the crying detection result at the segment level. The above-mentioned backbone network can select one of a variety of network structures, such as VGG, ResNet, MobileNet, etc.

[0063] Optionally, the backbone network includes a first convolution block, a first maximum pooling layer, a second convolution block and a second maximum pooling layer connected in sequence, and the first convolution block and the second convolution block are both composed of a two-dimensional convolution layer, a BN layer and a RELU activation layer connected in sequence, so as to realize the construction of the backbone network based on VGG.

[0064] Optionally, the segment-level network includes a one-dimensional convolutional layer and an activation layer connected in sequence, and the audio-level network includes a pooling layer to realize the construction of a dual-label network. The activation layer can be a sigmoid layer, and the pooling layer can be a maximum pooling layer or an average pooling layer. Based on the multiple judgment phenomenon of the maximum pooling layer and the missed judgment phenomenon of the average pooling layer, the pooling layer is optionally a temporal pooling layer to improve the detection performance of the crying detection model.

[0065] Optionally, the pooling function used by the temporal pooling layer is as follows:

[0066] pool = sum(pred_seg^2) / sum(pred_seg), where pool represents the pooling result, sum represents the summation operation, pred_seg represents the input of the temporal pooling layer, and ^2 represents square. That is, as long as there is crying in a certain audio segment, the output label of the audio segment corresponds to 1, thereby avoiding the application limitations of other pooling functions by fully considering the crying detection scenario.

[0067] When a pre-built cry detection model is trained using a training dataset, the cry detection model is optionally jointly trained based on the loss of a segment-level network and the loss of an audio-level network, thereby increasing the convergence speed of the cry detection model and improving the performance of the cry detection model.

[0068] Optionally, the loss function used by the crying detection model is as follows:

[0069] loss=loss1+0.1*loss2, where loss represents the total loss of the crying detection model, loss1 represents the loss of the audio-level network, and loss2 represents the loss of the segment-level network. Since the segment-level labels are only for the detection network to better predict the audio-level labels, loss2 can be understood as adding a regularization term to the original loss.

[0070] Figure 1C An example diagram of the network structure of a crying detection model provided in an embodiment of the present invention, the crying detection model includes a backbone network backbone, a segment-level network and an audio-level network connected in sequence, the backbone network backbone includes a first convolution block block1, a first maximum pooling layer maxpool1, a second convolution block block2, and a second maximum pooling layer maxpool2 connected in sequence, the segment-level network includes a one-dimensional convolution layer Conv1d and a sigmoid layer connected in sequence, and the audio-level network includes a time series pooling layer. Among them, the step sizes of the first maximum pooling layer and the second maximum pooling layer are both set to 4, and the time dimension is downsampled to 1 / 16 of the input dimension. Assuming that the input feature is [batch, 1, 501, 128], the output feature after the backbone is [batch, 256, 31, 1]. In the time dimension, the corresponding input time feature dimension 501 is downsampled to 31 dimensions after several layers of pooling functions. The output features of the backbone are reshaped to obtain a feature map of [batch, 256, 31]. After passing through the segment-level network, the segment-level output is [batch, 2, 31], which is recorded as the above pred_seg. That is, the original audio is divided into 31 small segments, each of which corresponds to a 1x2-dimensional output, corresponding to the probability values ​​of the output label 0 and label 1 (where label 0 means no crying, and label 1 means crying). The output of the segment-level network is input into the audio-level network, and the audio-level posterior probability is obtained through the temporal pooling layer.

[0071] When training a pre-built crying detection model, the dropout mechanism can be used to improve the generalization ability of the crying detection model. However, dropout randomly discards some neurons with a certain probability, and the neurons are independent, that is, most of the neurons discarded by dropout are single neurons, and the audio signal has temporal continuity. The feature map has spatial continuity through convolution. In this way, the effect of dropout is not rational for the convolution layer. Therefore, optionally, the feature map extracted by the backbone network is masked, and the crying detection model is trained based on the feature map after masking to avoid the adverse effects of the above-mentioned temporal continuity and spatial continuity on the generalization ability.

[0072] When performing mask processing on the feature map extracted by the backbone network, it is optional to determine whether to perform mask processing on the feature map extracted by the current layer according to the preset first probability value. If so, multiple channels to be masked are obtained, and mask processing is performed on the multiple channels according to the preset second probability value and the preset mask area size, thereby realizing mask processing on high-dimensional feature maps. Among them, the multiple channels to be masked can be randomly selected, and the number of channels N to be masked depends on the situation. If the amount of training data is relatively large and the covered scenes are relatively wide, then N can be relatively small. In general, N takes 1 / 3-1 / 2 of all channels; the above-mentioned first probability value and second probability value are usually set by the user, for example, the second probability value is usually set to 0.1; the size of the above-mentioned mask area can be determined according to the actual size of any feature map.

[0073] As an example, the dimension of the feature map of a certain layer is [batch, 128, 62, 16]. The first probability value p1 is used to determine whether to perform feature map masking. Otherwise, the feature map remains unchanged. If so, N channels are randomly selected for masking for the 128 channels. After that, for the N channels selected above, the feature map size is 62x16. The mask size [H, W] can be selected by the second probability value p2 to perform masking step by step. For a feature map size of 62x16, the time feature dimension H can be selected from 20 to 40, and the frequency feature dimension W can be selected from 6 to 12.

[0074] In an embodiment of the present invention, a training data set is obtained, which includes audio-level labels, segment-level labels and audio features corresponding to each training sample. The training data set is used to train a pre-built crying detection model to obtain a trained crying detection model, thereby jointly training the crying detection model with dual labels, thereby improving the training effect of the crying detection model while taking into account the insufficient data set.

[0075] Embodiment 2:

[0076] Figure 2 The structure of the training device of the crying detection model provided in the second embodiment of the present invention is shown. For the convenience of description, only the part related to the embodiment of the present invention is shown, including:

[0077] A training set acquisition unit 21 is used to acquire a training data set, wherein the training data set includes an audio level label, a segment level label and an audio feature corresponding to each training sample; and

[0078] The model training unit 22 is used to train the pre-built crying detection model using the training data set to obtain a trained crying detection model.

[0079] Optionally, the training set acquisition unit includes:

[0080] The label and feature acquisition unit is used to generate a double label for each training sample and extract the audio features of each training sample. The double label includes an audio level label and a segment level label.

[0081] Optionally, the label and feature acquisition unit includes:

[0082] A first label generating unit, configured to generate an audio level label according to whether the training sample contains crying sound; and

[0083] The second label generation unit is used to obtain the audio segment corresponding to each segment-level label according to the preset speech frame parameters, the preset input feature dimension downsampling multiples and the preset length of the training sample, and generate a segment-level label based on the comparison result of the number of audio sampling points corresponding to the crying in each audio segment and the preset sampling point number threshold.

[0084] Optionally, the training set acquisition unit further includes:

[0085] A data augmentation unit, used to perform data augmentation on each training sample;

[0086] The label and feature acquisition unit also includes:

[0087] The feature extraction subunit is used to extract the audio features of each training sample after data augmentation.

[0088] Optionally, the audio feature is an Fbank feature.

[0089] Optionally, the crying detection model includes a backbone network and a dual-label network connected in sequence, and the dual-label network includes a segment-level network and an audio-level network connected in sequence.

[0090] Optionally, the segment-level network includes a one-dimensional convolutional layer and an activation layer connected in sequence, and the audio-level network includes a pooling layer.

[0091] Optionally, the model training unit includes:

[0092] The mask processing and training unit is used to perform mask processing on the feature map extracted by the backbone network, and train the crying detection model based on the feature map after mask processing.

[0093] Optionally, the mask processing and training unit includes:

[0094] A judging unit, used for judging whether to perform mask processing on the feature map extracted from the current layer according to a preset first probability value;

[0095] a channel acquisition unit, configured to acquire a plurality of channels to be subjected to mask processing if mask processing is performed on the feature map extracted from the current layer; and

[0096] The mask processing subunit is used to perform mask processing on multiple channels according to a preset second probability value and a preset mask area size.

[0097] Optionally, the pooling layer is a temporal pooling layer.

[0098] Optionally, the model training unit includes:

[0099] A joint training unit for jointly training the cry detection model based on the loss of the segment-level network and the loss of the audio-level network.

[0100] In the embodiment of the present invention, each unit of the training device of the crying detection model can be implemented by a corresponding hardware or software unit, and each unit can be an independent software or hardware unit, or can be integrated into a software or hardware unit, which is not intended to limit the present invention. The specific implementation of each unit of the training device of the crying detection model can refer to the description of the aforementioned method embodiment, which will not be repeated here.

[0101] Embodiment three:

[0102] Figure 3 The structure of an electronic device provided by the third embodiment of the present invention is shown. For the convenience of description, only the part related to the embodiment of the present invention is shown.

[0103] The electronic device 3 of the embodiment of the present invention includes a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30. When the processor 30 executes the computer program 32, the steps in the above-mentioned method embodiments are implemented, for example: Figure 1A Alternatively, when the processor 30 executes the computer program 32, the functions of each unit in the above-mentioned device embodiments are realized, for example Figure 3 The functions of the units 21 to 22 are shown.

[0104] In an embodiment of the present invention, a training data set is obtained, which includes audio-level labels, segment-level labels and audio features corresponding to each training sample. The training data set is used to train a pre-built crying detection model to obtain a trained crying detection model, thereby jointly training the crying detection model with dual labels, thereby improving the training effect of the crying detection model while taking into account the insufficient data set.

[0105] Embodiment 4:

[0106] In an embodiment of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above method embodiment are implemented, for example, Figure 1A Alternatively, when the computer program is executed by a processor, the functions of each unit in the above-mentioned device embodiments are realized, for example Figure 2 The functions of the units 21 to 22 are shown.

[0107] In an embodiment of the present invention, a training data set is obtained, which includes audio-level labels, segment-level labels and audio features corresponding to each training sample. The training data set is used to train a pre-built crying detection model to obtain a trained crying detection model, thereby jointly training the crying detection model with dual labels, thereby improving the training effect of the crying detection model while taking into account the insufficient data set.

[0108] The computer-readable storage medium of the embodiment of the present invention may include any entity or device or recording medium capable of carrying computer program code, for example, ROM / RAM, magnetic disk, optical disk, flash memory and other memories.

[0109] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for training a crying detection model, characterized in that: The method comprises the following steps: Acquire a training data set, wherein the training data set includes an audio level label, a segment level label, and an audio feature corresponding to each training sample, wherein acquiring the training data set includes: generating a double label for each training sample, and extracting the audio feature of each training sample, wherein the double label includes the audio level label and the segment level label; Using the training data set to train the pre-built crying detection model to obtain a trained crying detection model, the crying detection model includes a backbone network and a dual-label network connected in sequence, the dual-label network includes a segment-level network and an audio-level network connected in sequence, the segment-level network includes a one-dimensional convolutional layer and an activation layer connected in sequence, and the audio-level network includes a pooling layer; The step of generating a double label for each training sample includes: Generate the audio level label according to whether the training sample contains crying sound; According to the preset speech framing parameters, the preset input feature dimension downsampling multiples and the preset length of the training sample, the audio segment corresponding to each segment-level label is obtained, and the segment-level label is generated based on the comparison result of the number of audio sampling points corresponding to the crying in each of the audio segments with the preset sampling point number threshold.

2. The method according to claim 1, characterized in that Before the step of extracting the audio features of each of the training samples, the method further includes: Performing data augmentation on each of the training samples; The step of extracting the audio features of each of the training samples further includes: The audio features of each training sample after data augmentation are extracted, and the audio features are Fbank features.

3. The method according to claim 1, characterized in that The steps to train the pre-built cry detection model include: Performing mask processing on the feature graph extracted by the backbone network, and training the crying detection model based on the feature graph after the mask processing; The step of performing mask processing on the feature map extracted by the backbone network comprises: Determine whether to perform mask processing on the feature map extracted from the current layer according to a preset first probability value; If so, obtaining a plurality of channels to be subjected to mask processing; The plurality of channels are subjected to mask processing according to a preset second probability value and a preset mask area size.

4. The method according to claim 1, characterized in that The pooling layer is a temporal pooling layer, and the pooling function used by the temporal pooling layer is as follows: pool=sum(pred_seg^2) / sum(pred_seg), where pool represents the pooling result, sum represents the summation operation, pred_seg represents the input of the temporal pooling layer, and ^2 represents square.

5. The method according to claim 1, characterized in that The step of using the training data set to train the pre-built crying detection model includes: The crying detection model is jointly trained according to the loss of the segment-level network and the loss of the audio-level network. The loss function adopted by the crying detection model is as follows: loss=loss1+0.1*loss2, where loss represents the total loss of the crying detection model, loss1 represents the loss of the audio level network, and loss2 represents the loss of the segment level network.

6. A crying detection model training device, characterized in that: The device comprises: A training set acquisition unit, configured to acquire a training data set, wherein the training data set includes an audio level label, a segment level label, and an audio feature corresponding to each training sample, wherein acquiring the training data set includes: generating a double label for each training sample, and extracting the audio feature of each training sample, wherein the double label includes the audio level label and the segment level label; and a model training unit, configured to train a pre-built crying detection model using the training data set to obtain a trained crying detection model, wherein the crying detection model comprises a backbone network and a dual-label network connected in sequence, wherein the dual-label network comprises a segment-level network and an audio-level network connected in sequence, wherein the segment-level network comprises a one-dimensional convolutional layer and an activation layer connected in sequence, and wherein the audio-level network comprises a pooling layer; Wherein, when generating the double label of each training sample, the training set acquisition unit includes: Generate the audio level label according to whether the training sample contains crying sound; According to the preset speech framing parameters, the preset input feature dimension downsampling multiples and the preset length of the training sample, the audio segment corresponding to each segment-level label is obtained, and the segment-level label is generated based on the comparison result of the number of audio sampling points corresponding to the crying in each of the audio segments with the preset sampling point number threshold.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Infant cry recognition method and system

    CN111862991A

  • Endpoint detection method and system based on joint deep neural network

    CN112735482A