Method and System for Training Voice Recognition Model and Computer Readable Medium

By judging the relationship between event tones and preset parameters in the sound recognition model training, multiple training tones files are sampled and generated, which solves the problems of high sampling repetition and change in feature distribution in the existing methods, and a small deep learning model is established to improve the recognition effect.

CN114360579BActive Publication Date: 2025-06-13WISTRON CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011344557.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-10-14
Filing Date
2020-11-26
Publication Date
2025-06-13
Estimated Expiration
2040-11-26

AI Technical Summary

Technical Problem

When the existing sound recognition model training methods deal with event tones of different lengths, the sampling method has the problem of high repetition or change in feature distribution, resulting in a decrease in recognition accuracy.

Method used

By judging the relationship between event tones and preset parameters, the sampling parameters are determined, and the event tones are sampled to generate multiple training tones files. The length of each training tones is associated with the first parameter, and the time difference between each two training tones is associated with the second parameter, avoiding high repetition and change in feature distribution.

Benefits of technology

A small deep learning model was established, which reduced the training complexity and initial R&D costs, improved the quality of training sound files, avoided the impact of event sound lengths on training results, and thus improved the identification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114360579B_ABST
    Figure CN114360579B_ABST
Patent Text Reader

Abstract

A method for training a voice recognition model includes determining the relationship between an event sound and a first parameter, and in response to the relationship, determining a second parameter. Using the first parameter and the second parameter, sampling the event sound to generate a plurality of training sound files, and inputting at least a part of the plurality of training sound files to train the voice recognition model, wherein the length of each training sound file is associated with the first parameter, the time difference between every two training sound files is associated with the second parameter, and the voice recognition model is used to determine the type of sound.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for training an identification model, and particularly to a method for training a voice identification model. Background Art

[0002] There are many different types of sounds in the living environment or the working environment, and different sounds can represent the occurrence of different events. Therefore, through voice identification technology, it is possible to assist in judging the situation of the living environment or the working environment, such as judging the occurrence of abnormal events. In the Detection and Classification of Acoustic Scenes and Events (DCASE) competition in 2017, the method for obtaining and processing training audio files proposed by the first-place winner was to decompose an event audio into detailed features and increase the temporal correlation. However, it belongs to a large deep learning model, and the architecture specifications used are relatively high, resulting in relatively high costs.

[0003] In addition, for the known sampling method of training audio files where the event audio is longer than the sampling length, sampling is performed with a fixed-length sampling displacement. Therefore, the longer the event audio is sampled, the more training audio files will be obtained, resulting in too high a repetition rate of the training audio files, causing the identification ability of the trained model to focus on specific sounds; for the known sampling method of training audio files where the event audio is shorter than the sampling length, it is performed by replicating the event audio to supplement the sampling length. In this way, the obtained training audio files will contain multiple identical and consecutive event audio, which instead changes the original feature distribution and may lead to a decrease in the subsequent identification accuracy. Summary of the Invention

[0004] In view of the above, the present invention provides a method and a system for training a voice identification model.

[0005] According to an embodiment of the present invention, a method for training a voice identification model includes determining the relationship between an event audio and a first parameter, and in response to this relationship, determining a second parameter. By using the first parameter and the second parameter, sampling the event audio to generate a plurality of training audio files, and inputting at least a part of the plurality of training audio files to train a voice identification model, wherein the length of each training audio file is related to the first parameter, the time difference between every two training audio files is related to the second parameter, and the voice identification model is used to judge the type of sound.

[0006] A voice recognition model training system according to an embodiment of the present invention includes an audio capture device, a processing device, and a storage device, wherein the processing device is connected to the audio capture device and the storage device. The audio capture device is used to obtain event sounds. The processing device is connected to the audio capture device and is configured to: determine the relationship between the event sound and a first parameter, and in response to this relationship, determine a second parameter, sample the event sound by using the first parameter and the second parameter to generate a plurality of training sound files, and input at least a part of the plurality of training sound files to train a voice recognition model, wherein the length of each training sound file is associated with the first parameter, the time difference between every two training sound files is associated with the second parameter, and the voice recognition model is used to determine the voice type. The storage device is used to store the voice recognition model.

[0007] A computer-readable medium according to an embodiment of the present invention includes program code, and the program code is used to be run by a processor to execute: determine the voice type according to a voice recognition model, wherein the voice recognition model is trained by the voice recognition model training method described in the foregoing embodiment.

[0008] With the above structure, the voice recognition model training method and system disclosed in the present application can establish a small deep learning model as the voice recognition model. Compared with a large deep learning model, the small deep learning model has a lower training complexity and a lower initial R & D cost. Through a special preprocessing process of the training sound files, the voice recognition model and the computer-readable medium established by the voice recognition model training method and system disclosed in the present application can have good training sound file quality, avoid the influence of the length of the event sound on the training result, and thus have good recognition effectiveness.

[0009] The above description of the disclosure content and the following description of the embodiments are used to demonstrate and explain the spirit and principle of the present invention, and provide a further explanation of the claims of the present invention. Description of the Drawings

[0010] Figure 1 It is a functional block diagram of a voice recognition model training system shown according to an embodiment of the present invention.

[0011] Figure 2 It is a flowchart of a voice recognition model training method shown according to an embodiment of the present invention.

[0012] Figure 3 It is a flowchart of a voice recognition model training method shown according to another embodiment of the present invention.

[0013] Figure 4 It is a sampling schematic diagram of an event sound shown according to an embodiment of the present invention.

[0014] Figure 5It is a flowchart of a method for training a voice recognition model according to another embodiment of the present invention. Detailed implementation manners

[0015] The detailed features and advantages of the present invention will be described in detail in the following implementation manners. The content is sufficient for any person skilled in the relevant art to understand the technical content of the present invention and implement it accordingly. And according to the content disclosed in this specification, the claims and the drawings, any person skilled in the relevant art can easily understand the relevant purposes and advantages of the present invention. The following embodiments further illustrate the viewpoints of the present invention, but do not limit the scope of the present invention in any way.

[0016] The present invention provides a voice recognition model training system and method for establishing a voice recognition model associated with special types of event sounds. The special types of event sounds are, for example, baby cries, dog barks, screams, voices, horn sounds, alarm sounds, gunshots, glass breaking sounds, etc. Please refer to Figure 1 , Figure 1 It is a functional block diagram of a voice recognition model training system according to an embodiment of the present invention. As Figure 1 shown, the voice recognition model training system 1 includes an audio capture device 11, a processing device 13 and a storage device 15, wherein the processing device 13 is connected to the audio capture device 11 and the storage device 15 in a wired or wireless manner.

[0017] The audio capture device 11 is used to obtain the original audio file. For example, the audio capture device 11 includes a wired transmission port such as USB, micro USB, etc., or a wireless transmission port such as a Bluetooth transceiver, a WIFI transceiver, etc., and can receive the original audio file from other devices. Take another example, the audio capture device 11 includes a radio such as a microphone and can receive external sounds to generate the original audio file. In one embodiment, the audio capture device 11 transmits the original audio file to the processing device 13 as an event sound. In another embodiment, in addition to the audio input component (such as the above-mentioned transmission connection port or radio) of the audio capture device 11, it also includes a processor such as a central processing unit, a microcontroller, a programmable logic controller, etc., and can process the original audio file to generate an event sound.

[0018] The processing device 13 can be a processor or an electronic device including a processor. The processor is, for example, a central processing unit, a microcontroller, a programmable logic controller, etc. The processing device 13 can obtain the event sound from the audio capture device 11, preprocess the event sound to generate a plurality of training audio files, and then input at least a part of the plurality of training audio files to train the voice recognition model. Among them, the process of the preprocessing will be described later. In one embodiment, the processing device 13 can include a plurality of processors to respectively execute the preprocessing of the event sound and the training of the training audio files.

[0019] The storage device 15 may include one or more non-volatile memories, such as flash memory, read-only memory (ROM), magnetoresistive random access memory (MRAM), etc. The storage device 15 may store the voice recognition model established by the processing device 13. In one embodiment, the storage device 15 and the processing device 13 may be disposed in the same host. In another embodiment, the storage device 15 may be disposed in a cloud server, and the processing device 13 may upload the voice recognition model to the storage device 15 through a wireless network. Further, the cloud server may provide a voice recognition service using the voice recognition model. The cloud server may receive the voice input to be recognized, perform recognition using the voice recognition model, and push the recognition result to the user device (such as a mobile phone) that has downloaded the corresponding application from the cloud server in the form of a notification or warning.

[0020] Please refer to Figure 1 and Figure 2 where Figure 2 is a flowchart of a method for training a voice recognition model according to an embodiment of the present invention. As Figure 2 shown, the method for training a voice recognition model may include step S11: determining the relationship between the event sound and the first parameter, and determining the second parameter in response to the relationship; S12: sampling the event sound by means of the first parameter and the second parameter to generate a plurality of training sound files, wherein the length of each training sound file is associated with the first parameter, and the time difference between every two training sound files is associated with the second parameter; and S13: inputting at least a part of the plurality of training sound files to train the voice recognition model, and the voice recognition model is used to determine the voice type. Figure 2 The method for training a voice recognition model shown in Figure 1 is applicable to the voice recognition model training system 1 shown in Figure 2 Further, the voice recognition model training system 1 may obtain the event sound by means of the audio capture device 11, execute steps S11 to S13 by the processing device 13 to obtain the training sound files for training the voice recognition model from the event sound, and store the voice recognition model by means of the storage device 15. However, Figure 2 the method for training a voice recognition model shown in Figure 1 is not limited to being executed by the system architecture shown in

[0021] Please refer to Figure 1 , Figure 2 and Figure 3 where Figure 3 is a flowchart of a method for training a voice recognition model according to another embodiment of the present invention. As Figure 3As shown, the method for training a voice recognition model may include step S21: obtaining an original voice file; step S22: comparing the time length of the event voice in the original voice file with a preset sampling length; step S23: when the time length is greater than the preset sampling length, obtaining an estimated sampling displacement based on the time length, the preset sampling length, and the upper limit of the number of samples; step S24: determining whether the estimated sampling displacement is greater than or equal to a displacement threshold; when the determination result is yes, executing step S25: sampling the event voice according to the preset sampling length and the estimated sampling displacement to generate a plurality of training voice files; when the determination result is no, executing step S26: sampling the event voice according to the preset sampling length and the displacement threshold to generate a plurality of training voice files; step S27: storing at least a part of the plurality of training voice files in a training voice file set; and step S28: training a voice recognition model according to the training voice file set.

[0022] Specifically, Figure 3 The steps of the method for training a voice recognition model shown may be Figure 2 The detailed implementation of the steps of the method for training a voice recognition model shown. Further, Figure 2 The determination of the relationship between the event voice and the first parameter in step S11 of Figure 3 May be implemented by step S22 of Figure 2 Where the preset sampling length is used as the first parameter; Figure 3 The determination of the second parameter in response to the relationship between the event voice and the first parameter in step S11 of Figure 2 May include step S23, step S24 of Figure 2 And when the estimated parameter (estimated sampling displacement) is greater than or equal to the displacement threshold, the estimated parameter will be used as the second parameter. Additionally, when the estimated parameter is less than the displacement threshold, Figure 3 Step S12 of Figure 2 Also includes using the displacement threshold as the second parameter; Figure 4 Step S12 of

[0023] Figure 3 The method for training a voice recognition model shown may be applicable to Figure 1 The voice recognition model training system 1 shown. The following describes multiple implementation aspects in which the voice recognition model training system 1 executes the method for training a voice recognition model. However, the method for training a voice recognition model proposed by the present invention is not limited to being implemented by Figure 1 The system architecture shown.

[0024] In step S21, the audio capture device 11 obtains the original audio file. As mentioned above, the audio capture device 11 can receive the original audio file from other devices or capture external sounds to generate the original audio file. In S22, the processor of the audio capture device 11 can extract the event sound from the original audio file and then transmit it to the processing device 13 for subsequent judgment.

[0025] Furthermore, the original audio file from other devices may have start tags and end tags of event sounds previously annotated after analysis by other devices or manually marked by users. The audio capture device 11 can then extract the event sound based on the start tag and end tag. Alternatively, the internal memory of the audio capture device 11 can store various types of event sound templates (such as the sound waveforms corresponding to various types of event sounds, and the types can be, for example, baby crying, dog barking, screaming, speaking, alarm sound, horn sound, gunshot sound, glass breaking sound, etc.). The processor of the audio capture device 11 can determine that the sound segment in the original audio file that conforms to a certain event sound template is the event sound. Or, the audio capture device 11 can regard the entire original audio file as the event sound and transmit it to the processing device 13 for subsequent judgment.

[0026] In step S22, the processing device 13 determines the time length of the event sound received from the audio capture device 11 and compares the time length of the event sound with a preset sampling length. Specifically, since the subsequent training audio file will be transformed into the frequency domain by Fourier transform during training, it is better to set the preset sampling length to a power of 2. In addition, the preset sampling length can be set to different lengths depending on whether the event sound is a long sound or a short sound. Among them, long sounds can include types such as baby crying, dog barking, screaming, speaking, etc., and short sounds can include types such as alarm sound, horn sound, gunshot sound, glass breaking sound, etc. Furthermore, when the event sound is a long sound, the preset sampling length is set to have a first value, and when the event sound is a short sound, the preset sampling length is set to have a second value, where the first value is greater than the second value. For example, when the event sound is a long sound, the preset sampling length is set to 4, and when the event sound is a short sound, the preset sampling length is set to 1. The preset sampling length can be set by the user through the user interface of the sound recognition model training system 1, or the processing device 13 of the sound recognition model training system 1 can determine whether the event sound is a long sound or a short sound and then set the preset sampling length accordingly.

[0027] To further illustrate the manner in which the processing device 13 determines whether an event sound belongs to a long sound or a short sound. In one embodiment, the processing device 13 can determine whether an event sound belongs to a long sound or a short sound based on the time length of the event sound. For example, when the time length of the event sound is 2 seconds (inclusive) or less, the processing device 13 determines that it belongs to a short sound, and vice versa for a long sound. In another embodiment, in addition to the start tag and end tag of the previously marked event sound in the original sound file, there is also a tag indicating whether the event sound belongs to a long sound or a short sound. The processing device 13 can set the preset sampling length based on this tag. In yet another embodiment, the internal memory of the processing device 13 stores a look-up table indicating whether each type of event sound belongs to a long sound or a short sound, and in addition to the start tag and end tag of the previously marked event sound in the original file, there is also a tag indicating the type of the event sound. The processing device 13 can determine whether the event sound belongs to a long sound or a short sound based on this tag and the above look-up table, and then set the preset sampling length accordingly. In still another embodiment, the internal memory of the processing device 13 stores various types of event sound templates (such as the event sound templates stored in the audio capture device 11) and a look-up table indicating whether each type of event sound belongs to a long sound or a short sound. The processing device 13 can determine the type of the event sound in the original sound file based on the event sound template, and determine whether the event sound belongs to a long sound or a short sound based on the look-up table, and then set the preset sampling length accordingly.

[0028] In step S23, when the time length of the event sound is greater than the preset sampling length, the processing device 13 obtains an estimated sampling displacement based on the time length, the preset sampling length, and the sampling number limit. The sampling number limit can be adjusted according to the number of training sound files in the current training sound file set. For example, when the number of training sound files is sufficient (such as when more than 1000 original sound files are appropriately cut and amplified to more than 3000 training sound files), the sampling number limit is set to 3, and when the number of training sound files is insufficient, the sampling number limit is set to 5. Further, the obtaining of the estimated sampling displacement performed by the processing device 13 can include performing the following calculation formula:

[0029] where S L represents the estimated sampling displacement, E L represents the time length of the event sound, the W L represents the preset sampling length, and the N represents the sampling number limit.

[0030] In this embodiment, when the time length of the event sound is less than or equal to the preset sampling length, the processing device 13 can fill the difference between the time length of the event sound and the preset sampling length according to the background sound to generate a training sound file, and store it in the training sound file set. Further, when the time length of the event sound is equal to the preset sampling length, the difference between the time length of the event sound and the preset sampling length is zero, and the processing device 13 directly uses the event sound as the training sound file; when the time length of the event sound is less than the preset sampling length, the processing device 13 can add a segment of the background sound before or after or both before and after the segment of the event sound to make up the preset sampling length, and use the combined sound segment as the training sound file, and the background sound is, for example, the sound segments before and after the event sound in the original sound file. By the above method of generating the training sound file according to the event sound and the background sound, the problem of changing the original characteristic distribution of the event sound by copying the event sound to make up the length can be avoided.

[0031] In step S24, the processing device 13 determines whether the estimated sampling displacement is greater than or equal to the displacement threshold. Among them, this displacement threshold is related to the preset sampling length, especially one-fourth of the preset sampling length, but the present invention is not limited thereto. When the judgment result is yes, the processing device 13 executes step S25 to sample the event sound according to the preset sampling length and the estimated sampling displacement to generate a plurality of training sound files. Further, please refer to Figure 4 , Figure 4 which is a sampling schematic diagram of the event sound shown in an embodiment of the present invention. As Figure 4 shown, the processing device 13 will process the event sound with a time length of E L , using the preset sampling length W L as the time length of each sampling, and using the estimated sampling displacement S L as the time difference between every two samplings, that is, using the preset sampling length W L as the time length of the training sound file, and using the estimated sampling displacement S L as the time difference between every two training sound files. Specifically, the time difference between the two samplings / training sound files can refer to the time interval between the start time points of the first sampling / training sound file and the second sampling / training sound file, where the start time point can be replaced by the end time point or any time point between the start time point and the end time point, and the present invention is not limited thereto.

[0032] Please refer to Figure 1 and Figure 3, when the judgment result of step S24 is negative, the processing device 13 executes step S26, and samples the event sound according to a preset sampling length and displacement threshold to generate a plurality of training sound files. Further, when the processing device 13 determines that the estimated sampling displacement is less than the displacement threshold, it uses the preset sampling length as the time length of each sampling (training sound file), and uses the displacement threshold to replace the estimated sampling displacement as the time difference between two samplings (training sound files). Thereby, it is possible to avoid the problem that the identification model built is too concentrated on specific sound features due to the too short time interval between two samplings and the too high overlap rate of the training sound files.

[0033] In step S27, the processing device 13 stores at least a part of the plurality of training sound files generated in step S25 or step S26 into the training sound file set. The processing device 13 can store all the generated training sound files into the training sound file set. Alternatively, the processing device 13 can also perform a volume filtering step to screen the training sound files that can be stored in the training sound file set. Further, the processing device 13 can perform the following steps for each training sound file: dividing the training sound file into a plurality of segmentation results in a preset unit; respectively determining whether the volume of each of the plurality of segmentation results is greater than or equal to a preset volume; and when more than half of the volumes of the plurality of segmentation results are greater than or equal to the preset volume, storing the training sound file into the training sound file set, and when more than half of the volumes of the plurality of segmentation results are less than the preset volume, discarding the training sound file. Among them, the preset unit is, for example, 64 ms, and the preset volume is, for example, -30 dBFS, which is equivalent to the volume of white noise. Since the quality of the collected training sound files may vary, the quality of some training sound files is poor, resulting in insufficient signal strength and too small volume, which in turn affects model learning. By the above volume filtering step, this problem can be avoided. In addition, the processing device 13 can also perform the above volume filtering step on the event sound when obtaining the event sound, so as to discard the event sound with too small volume before performing sampling.

[0034] In step S28, the processing device 13 trains a sound identification model based on the training sound file set. Further, the processing device 13 can generate a deep learning model in a deep learning manner from the training sound files to be used as a sound identification model. For example, the framework selected for the deep learning model is Keras, the architecture is MobileNetV2, and it is accelerated with OpenVINO. MobileNetV2 is a lightweight neural network. Compared with other neural networks such as VGG and ResNet, it has fewer parameters, faster instruction cycles, and good accuracy. Figure 3Exemplarily presenting the method for training a voice recognition model includes steps S21 to S28. In another embodiment, before performing step S28 to train the voice recognition model, the voice recognition model training system 1 may execute steps S21 to S27 multiple times. Specifically, the voice recognition model established by the method proposed above belongs to a small deep learning model, and the multiple training audio files generated by steps S21 to S27 have no temporal association with each other. Therefore, compared with the training method of a large deep learning model that uses both features and temporal correlation as training parameters, it has a lower training complexity. By the voice recognition model training method proposed in this application and based on the concept of microservices (Microservice), a large deep learning model can be replaced by multiple small deep learning models respectively corresponding to different types of event sounds.

[0035] Specifically, the original audio file may include one or more event sounds. When the processing device 13 determines that the original audio file has multiple event sounds, it can compare each event sound with a preset sampling length and perform steps S23 to S27 to process each event sound into multiple training audio files and store them in the training audio file set, and then execute step S28 to generate a voice recognition model. The multiple event sounds included in the original audio file may belong to different types. The processing device 13 can determine the type of each event sound according to the aforementioned multiple types of event sound templates, and store the training audio files corresponding to different types of event sounds in different training audio file sets. Alternatively, in addition to the start label and end label of the previously marked event sound, the original audio file further has a type label for the processing device 13 to identify.

[0036] The present invention also proposes another method for training a voice recognition model. Please refer to Figure 1 、 Figure 3 and Figure 5 , where Figure 5 is a flowchart of the method for training a voice recognition model shown in another embodiment of the present invention. Figure 5 The steps of the voice recognition model training method shown are generally similar to those of the voice recognition model training method shown in Figure 3 , and can also be the detailed implementation of the voice recognition model training method shown in Figure 2 . And it can also be applied to the voice recognition model training system 1 shown in Figure 1 . The difference is that the voice recognition model training method shown in Figure 5 compares the time length of the event sound with the preset sampling length by judging the proportional relationship between the time length of the event sound and the preset sampling length, as shown in step S32, and has different processing methods for multiple proportional ranges.

[0037] When the processing device 13 determines that the time length of the event sound is more than 100% of the preset sampling length, it means that the time length of the event sound is greater than the preset sampling length. At this time, the processing device 13 will execute steps S33 to S38, which are the same as Figure 3 the steps S23 and subsequent steps S24 to S28 shown, to train the voice recognition model. The detailed implementation manners are all as described in the foregoing embodiments and will not be elaborated herein. When the processing device 13 determines that the time length of the event sound is between X% and 100% of the preset sampling length (that is, the time length of the event sound is less than or equal to the preset sampling length and greater than or equal to X% of the preset sampling length, where X% is a preset ratio and less than 100%), as shown in step S33', the processing device 13 can fill the difference between the time length of the event sound and the preset sampling length according to the background sound to generate a training sound file and store it in the training sound file set. Further, when the time length of the event sound is equal to the preset sampling length, the difference between the time length of the event sound and the preset sampling length is zero, and the processing device 13 directly uses the event sound as the training sound file; when the time length of the event sound is less than the preset sampling length and greater than or equal to X% of the preset sampling length, the processing device 13 can add a segment of the background sound before or after or both before and after the segment of the event sound to make up the preset sampling length, and use the combined sound segment as the training sound file, and the background sound is, for example, the sound segments before and after the event sound in the original sound file. By the above method of generating the training sound file based on the event sound and the background sound, the problem of changing the original feature distribution of the event sound by copying the event sound to make up the length in the prior art can be avoided.

[0038] Next, the processing device 13 can be as Figure 3Perform step S38 as shown in step S28 to train the voice recognition model. Alternatively, before performing step S38, the processing device 13 can drive the audio capture device 11 to perform step S31 multiple times to obtain another event sound from another original audio file or from the original audio file, so as to make the judgment in step S32. When the processing device 13 determines that the time length of the event sound is less than X% of the preset sampling length, as shown in step S33, the processing device 13 will discard this event sound and execute step S31 to request or receive another event sound from the audio capture device 11, so as to execute step S32 and subsequent steps again. Specifically, the preset ratio X% can be set to different values according to whether the event sound is a long sound or a short sound. Among them, long sounds can include types: baby crying, dog barking, screaming, speaking, etc., and short sounds can include types: alarm sound, horn sound, gunshot sound, glass breaking sound, etc. For example, when the event sound is a long sound, the preset ratio is set to 40% - 60%, and 50% is preferred; when the event sound is a short sound, the preset ratio is set to 15% - 35%, and 25% is preferred. Since the time length of the event sound belonging to the short sound is relatively short, if the same preset ratio as that of the event sound belonging to the long sound is adopted, many short event sounds will be discarded, so through the adjustment of the above preset ratio, the problem that the time for collecting training audio files becomes longer due to too many short event sounds being discarded can be avoided.

[0039] The preset ratio can be set by the user through the user interface of the voice recognition model training system 1, or the processing device 13 of the voice recognition model training system 1 can determine whether the event sound is a long sound or a short sound and then set the preset ratio accordingly. In one embodiment, the processing device 13 can determine whether the event sound is a long sound or a short sound according to the time length of the event sound. For example, when the time length of the event sound is 2 seconds (inclusive) or less, the processing device 13 determines that it is a short sound, otherwise it is a long sound. In another embodiment, in addition to the start label and end label of the previously marked event sound in the original audio file, there is also a label indicating whether the event sound is a long sound or a short sound, and the processing device 13 can set the preset ratio according to this label. In yet another embodiment, the internal memory of the processing device 13 stores a look-up table for each type of event sound indicating whether it is a long sound or a short sound, and in addition to the start label and end label of the previously marked event sound in the original file, there is also a label indicating the type of the event sound. The processing device 13 can determine whether the type of the event sound is a long sound or a short sound according to this label and the above look-up table, and then set the preset ratio accordingly. In still another embodiment, the internal memory of the processing device 13 stores various types of event sound templates and a look-up table for each type of event sound indicating whether it is a long sound or a short sound. The processing device 13 can determine the type of the event sound in the original audio file according to the event sound template, and determine whether the type of the event sound is a long sound or a short sound according to the look-up table, and then set the preset ratio accordingly.

[0040] The voice recognition model trained by the voice recognition model training method described in the above embodiments can be included in a non-transitory medium in the form of program code, such as computer-readable storage media like CD-ROMs, USB flash drives, memory cards, hard disks of cloud servers, etc. When the processor of a computer loads the program code from this computer-readable medium and executes it, the processor can determine the type of voice based on the voice recognition model. Further, the processor can input the voice to be recognized into the voice recognition model so that the voice recognition model can recognize the type of the voice to be recognized.

[0041] Please refer to Table 1 and Table 2. Table 1 presents the evaluation metrics of the recognition results of the voice recognition model for long event voices established by the voice recognition model training system and method of an embodiment of the present invention, and Table 2 presents the evaluation metrics of the recognition results of the voice recognition model for short event voices established by the voice recognition model training system and method of an embodiment of the present invention. Further, the voice recognition models corresponding to Table 1 and Table 2 are established using the voice recognition model training method described in the foregoing Figure 5 embodiments. Among them, the voice recognition model training method corresponding to Table 1 sets the preset sampling length to 4 seconds and the preset ratio to 50%, and the voice recognition model training method corresponding to Table 2 sets the preset sampling length to 1 second and the preset ratio to 25%. Table 1 and Table 2 respectively include multiple evaluation metrics. Among them, precision represents the ratio of correctly judging positive and negative cases when judged as positive; recall represents how many are correctly judged among the positive samples; and the f1 score is the harmonic mean of the former two, serving as a comprehensive indicator of the two. As shown in Table 1 and Table 2, it can be seen that the voice recognition models for short event voices and long event voices established by the voice recognition model training system and method of the present application all have good recognition effects.

[0042] Table 1

[0043] Type Precision Recall F1 Score Number of Test Samples Baby Cry 0.81 0.84 0.82 512 Dog Bark 0.88 0.89 0.88 663 Scream 0.97 0.98 0.98 487 Speech 0.79 0.81 0.8 1195 Others 0.79 0.66 0.69 878 Average / Sum 0.848 0.836 0.834 3735

[0044] Table 2

[0045] Type Precision Recall F1 Score Number of Test Samples Alarm 0.74 0.67 0.7 268 Horn 0.85 0.86 0.86 497 Breaking Glass 0.8 0.8 0.8 469 Gunshot 0.86 0.9 0.88 564 Average / Sum 0.8125 0.8075 0.81 1798

[0046] With the above structure, the voice recognition model training method and system disclosed in the present application can establish a small deep learning model as the voice recognition model. Compared with large deep learning models, small deep learning models have lower training complexity and lower initial R & D costs. Through a special pre-processing process for training voice files, the voice recognition model and computer-readable medium established by the voice recognition model training method and system disclosed in the present application can have good training voice file quality, avoid the influence of the length of event voices on the training results, and thus have good recognition effects.

[0047] Although the present invention has been disclosed above in the foregoing embodiments, it is not intended to limit the present invention. Any modifications and refinements made without departing from the spirit and scope of the present invention fall within the scope of patent protection of the present invention. For the scope of protection defined by the present invention, please refer to the appended claims.

Claims

1. A method for training a voice recognition model, comprising: Determining the relationship between an event sound and a first parameter, and determining a second parameter in response to the relationship; Sampling the event sound by means of the first parameter and the second parameter to generate a plurality of training sound files, wherein the length of each of the training sound files is associated with the first parameter, and the time difference between every two of the training sound files is associated with the second parameter; and Inputting at least a part of the training sound files to train the voice recognition model, which is used to determine the voice type.

2. The method for training a voice recognition model according to claim 1, wherein determining the relationship between the event sound and the first parameter, and determining the second parameter in response to the relationship comprises: Comparing the time length of the event sound with the first parameter; When the time length of the event sound is greater than the first parameter, obtaining a predicted parameter according to the time length of the event sound, the first parameter and the sampling number upper limit; Determining whether the predicted parameter is greater than or equal to a displacement threshold; and When the predicted parameter is greater than or equal to the displacement threshold, using the predicted parameter as the second parameter.

3. The method for training a voice recognition model according to claim 2, wherein determining the relationship between the event sound and the first parameter, and determining the second parameter in response to the relationship further comprises: When the predicted parameter is less than the displacement threshold, using the displacement threshold as the second parameter.

4. The method for training a voice recognition model according to claim 2, wherein the displacement threshold is one-fourth of the first parameter.

5. The method for training a voice recognition model according to claim 2, wherein obtaining the predicted parameter according to the time length of the event sound, the first parameter and the sampling number upper limit comprises executing a calculation formula to obtain the predicted parameter, and the calculation formula is: Where S L represents the estimated parameter, E L represents the time length of the event sound, the W L represents the first parameter, and the N represents the upper limit of the number of samples.

6. The method for training a voice recognition model according to claim 2 further comprises: When the event sound belongs to a long sound, setting the first parameter to have a first value; and When the time length of the event sound is less than a preset ratio of the first parameter, obtaining another event sound; Wherein the first value is a power of 2, and the preset ratio is between 40% and 60%.

7. The method for training a voice recognition model according to claim 2 further comprises: When the event sound belongs to a short sound, setting the first parameter to have a second value; and When the time length of the event sound is less than a preset ratio of the first parameter, obtaining another event sound; Wherein the second value is a power of 2, and the preset ratio is between 15% and 35%.

8. The method for training a voice recognition model according to claim 2 further comprises: When the time length of the event sound is between the first parameter and a preset ratio of the first parameter, filling the difference between the time length of the event sound and the first parameter according to the background sound to generate a training sound file, and inputting the training sound file to train the voice recognition model.

9. The method for training a voice recognition model according to claim 1, wherein inputting at least a part of the training sound files to train the voice recognition model comprises performing, for each of the training sound files: Dividing the training sound file into a plurality of division results in a preset unit; Determine respectively whether the volume of the segmentation result is greater than or equal to a preset volume; and When more than half of the volumes of the segmentation result are greater than or equal to the preset volume, input the training audio file to train the voice recognition model.

10. A voice recognition model training system, comprising: An audio capture device for obtaining an event sound; A processing device connected to the audio capture device for performing: Judging the relationship between the event sound and a first parameter, and determining a second parameter in response to the relationship; Sampling the event sound by means of the first parameter and the second parameter to generate a plurality of training audio files, wherein the length of each training audio file is associated with the first parameter, and the time difference between every two training audio files is associated with the second parameter; and Input at least a part of the training audio files to train a voice recognition model for judging the voice type; and A storage device connected to the processing device for storing the voice recognition model.

11. The voice recognition model training system according to claim 10, wherein the judging of the relationship between the event sound and the first parameter by the processing device and determining the second parameter in response to the relationship includes: Comparing the time length of the event sound with the first parameter; When the time length of the event sound is greater than the first parameter, obtaining a predicted parameter according to the time length of the event sound, the first parameter and the sampling number upper limit; Judging whether the predicted parameter is greater than or equal to a displacement threshold; and When the predicted parameter is greater than or equal to the displacement threshold, using the predicted parameter as the second parameter.

12. The voice recognition model training system according to claim 11, wherein the judging of the relationship between the event sound and the first parameter by the processing device and determining the second parameter in response to the relationship further includes when the predicted parameter is less than the displacement threshold, using the displacement threshold as the second parameter.

13. The voice recognition model training system according to claim 11, wherein the displacement threshold is one-fourth of the first parameter.

14. The voice recognition model training system according to claim 11, wherein the processing device obtains the predicted parameter by executing a calculation formula, and the calculation formula is: Where S L represents the estimated parameter, E L represents the time length of the event sound, the W L represents the first parameter, and the N represents the upper limit of the number of samples.

15. The voice recognition model training system according to claim 11, wherein the processing device further sets the first parameter to have a first value when the event sound is a long sound, and obtains another event sound when the time length of the event sound is less than a preset ratio of the first parameter, wherein the first value is a power of 2, and the preset ratio is between 40% and 60%.

16. The voice recognition model training system according to claim 11, wherein the processing device further sets the first parameter to have a second value when the event sound is a short sound, and obtains another event sound when the time length of the event sound is less than a preset ratio of the first parameter, wherein the second value is a power of 2, and the preset ratio is between 15% and 35%.

17. The voice recognition model training system as described in claim 11, wherein when the time length of the event sound is between the first parameter and a preset ratio of the first parameter, the processing device fills the difference between the time length of the event sound and the first parameter with background sound to generate a training sound file, and inputs the training sound file to train the voice recognition model.

18. The voice recognition model training system as described in claim 10, wherein the input of at least a part of the training sound file by the processing device to train the voice recognition model includes performing for each of the training sound files: Dividing the training sound file into a plurality of segmentation results in a preset unit; Respectively determining whether the volume of the segmentation results is greater than or equal to a preset volume; and When more than half of the volumes of the segmentation results are greater than or equal to the preset volume, inputting the training sound file to train the voice recognition model.

19. A computer-readable medium, comprising program code, wherein the program code is used to be run by a processor to execute: Judging the voice type according to the voice recognition model; Wherein the voice recognition model is trained by the voice recognition model training method as described in claim 1.

Citation Information

Patent Citations

  • Device and method for audio classification and audio processing

    CN104078050A