Methods, apparatus, media and electronic equipment for generating speech activity detection models
By using the inverse of a smooth approximation of the F1 score as the loss function, combined with the cross-entropy loss function, the parameters of the speech activity detection model are adjusted, thus solving the problem of inconsistent model objectives and improving detection performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING YOUZHUJU NETWORK TECH CO LTD
- Filing Date
- 2023-01-29
- Publication Date
- 2026-05-05
AI Technical Summary
In practical applications, existing speech activity detection models often fail to achieve ideal detection results due to the inconsistency between the cross-entropy loss function and the target.
The negative of the smooth approximation of the F1 score is used as the loss function, combined with the cross-entropy loss function. By adjusting the model parameters until the preset conditions are met, a speech activity detection model is generated.
This improved the performance of the speech activity detection model in practical applications, ensured the consistency between model training and application goals, and enhanced detection accuracy and recall.
Smart Images

Figure CN116312619B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of neural network technology, and more specifically, to a method, apparatus, medium, and electronic device for generating a speech activity detection model. Background Technology
[0002] Voice activity detection (VAD) models can detect speech in an audio clip.
[0003] In related technologies, voice activity detection (VAD) models are usually optimized using the cross-entropy loss function as the objective function. However, in practical applications of VAD models, the objective function of the VAD model and the cross-entropy loss function used for optimization are not strictly consistent, resulting in the VAD model not performing well in real-world applications. Summary of the Invention
[0004] This section is provided to briefly introduce the concepts, which will be described in detail in the Detailed Description section later. This section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0005] In a first aspect, this disclosure provides a method for generating a speech activity detection model, comprising: acquiring a speech sample dataset, wherein the speech sample dataset includes speech sample frames and sample labels corresponding to the speech sample frames;
[0006] The speech sample frame is input into the speech activity detection model to obtain the speech probability result corresponding to the speech sample frame output by the speech activity detection model.
[0007] The loss value is determined based on the preset loss function, the corresponding speech probability results and sample labels of the speech sample frames in the speech sample dataset, wherein the loss function includes a first loss function, which is the inverse of a smooth approximation of the F1 score;
[0008] Based on the loss value, the model parameters of the speech activity detection model are adjusted until the training parameters of the speech activity detection model meet the preset conditions to obtain the speech activity detection model.
[0009] Secondly, this disclosure provides a speech activity detection model generation apparatus, comprising:
[0010] An acquisition module is used to acquire a speech sample dataset, wherein the speech sample dataset includes speech sample frames and sample labels corresponding to the speech sample frames;
[0011] The input module is used to input the speech sample frame into the speech activity detection model to obtain the speech probability result corresponding to the speech sample frame output by the speech activity detection model.
[0012] The determination module is used to determine the loss value based on a preset loss function, the corresponding speech probability results and sample labels of the speech sample frames in the speech sample dataset, wherein the loss function includes a first loss function, which is the inverse of a smooth approximation of the F1 score;
[0013] An adjustment module is used to adjust the model parameters of the speech activity detection model according to the loss value until the training parameters of the speech activity detection model meet preset conditions to obtain the speech activity detection model.
[0014] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon that, when executed by a processing device, implements the steps of the method described in the first aspect.
[0015] Fourthly, this disclosure provides an electronic device, comprising:
[0016] A storage device having at least one computer program stored thereon;
[0017] At least one processing device is configured to execute the at least one computer program in the storage device to implement the steps of the method in the first aspect.
[0018] The above technical solution considers that the F1 score is an indicator for evaluating the practical application of a model. Given that the F1 score is not continuously differentiable and that a lower loss value for model training generally indicates better model performance, the constructed first loss function is the inverse of a smooth approximation of the F1 score. The smooth approximation of the F1 score satisfies the condition that a lower loss value during model training indicates better model performance. Therefore, using the inverse of the smooth approximation of the F1 score as the first loss function for training the speech activity detection model ensures that the loss function used in the speech activity detection model generation process is consistent with the target in the actual application of the model, thus improving the performance of the trained speech activity detection model in practical applications.
[0019] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0020] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings:
[0021] Figure 1 This is a flowchart illustrating a method for generating a speech activity detection model according to an exemplary embodiment.
[0022] Figure 2 This is a block diagram illustrating a speech activity detection model generation apparatus according to an exemplary embodiment.
[0023] Figure 3 This is a schematic diagram of the structure of an electronic device according to an exemplary embodiment. Detailed Implementation
[0024] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0025] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0026] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0027] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0028] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0029] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0030] All actions involving the acquisition of signals, information, or data in this disclosure are carried out in accordance with the relevant data protection laws and policies of the country where the location is situated, and with the authorization granted by the owner of the relevant device.
[0031] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0032] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0033] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0034] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0035] Meanwhile, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0036] As mentioned in the background section, voice activity detection (VAD) models are typically optimized using a cross-entropy loss function as the objective function, also known as the loss function. The optimization goal of the cross-entropy loss function is to make the model's identified results as consistent as possible with the labeled sample results. However, in practical applications, the objectives of VAD models can be quantified as precision and recall. Precision characterizes the total number of frames where the model identifies speech as a result, while recall characterizes the ratio of the number of frames identified as speech to the total number of frames identified as speech. Higher precision results in lower recall, and these two are inversely proportional. Therefore, to balance precision and recall, the geometric mean of precision and recall, known as the F1 score, is derived. In practical applications of VAD models, the goal is usually to maximize both precision and recall, or in other words, to maximize the F1 score. However, this objective does not have a strict consistency guarantee with the optimization objective of the cross-entropy loss function used in the model generation process, leading to less than ideal performance of VAD models in real-world applications.
[0037] In view of this, the present disclosure provides a method, apparatus, medium, and electronic device for generating a speech activity detection model, which improves the effectiveness of the speech activity detection model in practical applications.
[0038] The present disclosure will be explained below with reference to the accompanying drawings.
[0039] Figure 1 This is a flowchart illustrating a method for generating a speech activity detection model according to an exemplary embodiment. This method can be applied to electronic devices, such as mobile terminals (e.g., mobile phones, tablets) or fixed terminals (e.g., servers, desktop computers). (Refer to...) Figure 1 This includes the following steps:
[0040] Step S101: Obtain the speech sample dataset, which includes speech sample frames and corresponding sample labels for the speech sample frames.
[0041] The sample label corresponding to the speech sample frame is used to characterize whether the speech sample frame is speech. For example, label 0 can be used to characterize non-speech, and label 1 can be used to characterize speech.
[0042] The speech sample dataset can be sampled from pre-collected samples. The pre-collection method can refer to relevant technologies, which will not be elaborated here.
[0043] Step S102: Input the speech sample frame into the speech activity detection model to obtain the speech probability result corresponding to the speech sample frame output by the speech activity detection model.
[0044] Among them, the speech probability result corresponding to the speech sample frame can be used to characterize the probability that the speech sample frame is speech.
[0045] The model structure of the speech activity detection model can be referred to in relevant technologies, and will not be elaborated here in this embodiment.
[0046] Step S103: Determine the loss value based on the preset loss function, the corresponding speech probability results and sample labels of the speech sample frames in the speech sample dataset, wherein the loss function includes a first loss function, which is the inverse of a smooth approximation of the F1 score.
[0047] The smooth approximation of the F1 score can be characterized by the following equation (1):
[0048] (1);
[0049] in, For a smooth approximation of the F1 score, () is the summation function. Let be the sample label of the i-th speech sample frame in the speech sample dataset. Let be the probability that the i-th speech sample frame is predicted as speech, and e be a natural constant.
[0050] The negative of the smooth approximation of the F1 score is - .
[0051] Step S104: Adjust the model parameters of the speech activity detection model according to the loss value until the training parameters of the speech activity detection model meet the preset conditions to obtain the speech activity detection model.
[0052] It is worth noting that adjusting the model parameters of the speech activity detection model is an iterative update process. When the training parameters of the speech activity detection model do not meet the preset conditions, the speech sample frames are re-inputted into the speech activity detection model after adjusting the model parameters. Based on the speech probability results and sample labels of the speech activity detection model after adjusting the model parameters, the loss value is determined, and then the model parameters of the speech activity detection model are adjusted based on the loss value.
[0053] Among them, the parameters for evaluating the training of the speech activity detection model can be the number of times the model parameters of the speech activity detection model are adjusted. In this case, if the number of times the model parameters of the speech activity detection model are adjusted reaches a preset number, it can be determined that the parameters for evaluating the training of the speech activity detection model meet the preset conditions. The speech activity detection model after adjusting the model parameters is the trained speech activity detection model. It is worth noting that the number of times the model parameters of the speech activity detection model are adjusted reaches a preset number can indicate that the model has converged.
[0054] The parameters for evaluating the training of the speech activity detection model can be such that the difference between the loss value and the previously determined loss value is less than a first preset threshold. In this case, the parameters for evaluating the training of the speech activity detection model can be determined to meet the preset conditions when the difference between the loss value and the previously determined loss value is less than the first preset threshold. The model corresponding to the adjusted model parameters of the speech activity detection model based on the previous loss value is the trained speech activity detection model. It is worth noting that the difference between the loss value and the previously determined loss value being less than the first preset threshold can indicate that the model has converged.
[0055] The trained speech activity detection model can be used for speech detection in different scenarios.
[0056] The above approach considers that the F1 score is used as an indicator to evaluate the practical application of the model. Given that the F1 score is not continuously differentiable and that a lower loss value for model training generally indicates better model performance, the first loss function is constructed as the inverse of a smooth approximation of the F1 score. The smooth approximation of the F1 score is continuously differentiable, and its inverse satisfies the condition that a lower loss value during model training indicates better model performance. Therefore, using the inverse of the smooth approximation of the F1 score as the first loss function for training the speech activity detection model ensures that the loss function used during the generation process aligns with the model's objective in practical applications, thus improving the effectiveness of the trained speech activity detection model in real-world applications.
[0057] In some embodiments, the loss function further includes a second loss function, which is a cross-entropy loss function. In this case, step S103 can be implemented as follows: determine a first loss value based on the first loss function, the corresponding speech probability results and sample labels of the speech sample frames in the speech sample dataset; determine a second loss value based on the second loss function, the corresponding speech probability results and sample labels of the speech sample frames in the speech sample dataset; and weight the first loss value and the second loss value according to a weight relationship to obtain a loss value. The weight relationship is used to characterize that the sum of the first weight corresponding to the first loss value and the second weight corresponding to the second loss value is a preset value. The first weight is directly proportional to the number of times the model parameters of the speech activity detection model have been adjusted, and the second weight is inversely proportional to the number of times the model parameters of the speech activity detection model have been adjusted.
[0058] For example, the weighted average of the first and second loss values can be represented by the following equation (2):
[0059] (2);
[0060] Wherein, loss is the weighted sum of the first loss value and the second loss value, i.e., the loss value. As the first weight, The first loss value, As the second weight, This is the second loss value. It's worth noting that here... The first loss function, constructed from the inverse of the smooth approximation of the F1 score in the example above, can be calculated. It can be characterized by the following formula (3):
[0061] (3);
[0062] Among them, it can be determined by the above formula (3). In equation (3) above, the meaning of each physical parameter can be found in the explanation of equation (1) above. This embodiment will not repeat the explanation here.
[0063] The cross-entropy loss function can be characterized by the following equation (4):
[0064] (4);
[0065] in, The second loss value is given by N, where N is the number of speech sample frames in the speech sample dataset. Let be the sample label of the i-th speech sample frame in the speech sample dataset. The probability that the i-th speech sample frame is predicted as speech can be determined by the above equation (4).
[0066] The preset value can be 1, and the relationship between the first weight and the second weight can be represented by the following formula (5):
[0067] (5);
[0068] in, As the first weight, The first weight increases with the number of times the model parameters of the speech activity detection model have been adjusted; that is, the first weight is directly proportional to the number of times the model parameters of the speech activity detection model have been adjusted. For example, the first weight can increase to 1 as the number of times the model parameters of the speech activity detection model have been adjusted increases. The second weight decreases with the number of times the model parameters of the speech activity detection model have been adjusted; that is, the second weight is inversely proportional to the number of times the model parameters of the speech activity detection model have been adjusted. For example, the second weight can decrease to 0 as the number of times the model parameters of the speech activity detection model have been adjusted increases.
[0069] Furthermore, the second weight can be characterized by the following equation (6):
[0070] (6);
[0071] in, As the second weight, This represents the number of times the model parameters of the current speech activity detection model have been adjusted. The preset number of times is the total number of times the model parameters of the speech activity detection model are adjusted, i.e., the preset number mentioned above, where e is a natural constant.
[0072] The second weight can be determined by the above formula (6), and the first weight can be determined based on the weight relationship. Based on the first weight and the second weight, the first loss value and the second loss value are weighted to obtain the loss value.
[0073] It is worth noting that when using a large number of samples to construct a speech sample dataset for training a speech activity detection model, the speech sample dataset obtained each time is not uniformly sampled. The F1 score involves summing the samples in the speech sample dataset, which leads to a biased estimate of the F1 score. This results in the problem of training divergence caused by using the opposite of the smooth approximation of the F1 score as the objective function of the model. Therefore, to address the training divergence problem, a first loss function, constructed by combining the cross-entropy loss function and the inverse of the smooth approximation based on the F1 score, is used to train the speech activity detection model. The first weight is proportional to the number of times the model parameters have been adjusted, and the second weight is inversely proportional to the number of times the model parameters have been adjusted. The model is first trained to convergence using the cross-entropy loss function, and then fine-tuned using the first loss function to obtain a well-trained speech activity detection model. Furthermore, a second weight is set that automatically changes based on the preset total number of times the model parameters have been adjusted and the number of times the model parameters have been adjusted. This allows the model to automatically change the first and second weights during the loss value determination process, enabling the loss values corresponding to the first and second loss functions to change to different degrees at different training stages without manual intervention, i.e., without manually replacing the model's loss function at different training stages.
[0074] In the case where the model parameters of the speech activity detection model are adjusted using the loss value obtained by weighting the first loss value and the second loss value according to the weight relationship, the above-mentioned preset condition includes that the number of times the model parameters of the speech activity detection model have been adjusted has reached the total number of times.
[0075] In some embodiments, step S102 can be implemented as follows: inputting a speech sample frame into a candidate speech activity detection model to obtain the speech probability result corresponding to the speech sample frame output by the candidate speech activity detection model, wherein the candidate speech activity detection model is trained using the cross-entropy loss function as the objective function.
[0076] It is worth noting that in the aforementioned embodiments, a scheme is provided that combines the cross-entropy loss function and the inverse of the smooth approximation of the F1 score as the loss function of the model to collaboratively generate a speech activity detection model. It can be understood that different loss functions can also be used as the loss functions of the speech activity detection model at different stages to adjust the model parameters of the speech activity detection model, thereby obtaining a trained speech activity detection model.
[0077] For example, the first stage of obtaining a trained speech activity detection model is to use the cross-entropy loss function as the objective function of the speech activity detection model. Then, the second stage of obtaining a trained speech activity detection model is to use the inverse of the smooth approximation of the F1 score as the objective function of the candidate speech activity detection model.
[0078] The candidate speech activity detection model can be obtained as follows: inputting speech sample frames into the speech activity detection model, and outputting the speech probability results corresponding to the speech sample frames; using the cross-entropy loss function, the corresponding speech probability results and sample labels of the speech sample frames in the speech sample dataset are used to determine the loss value, and the model parameters of the speech activity detection model are adjusted based on the loss value until the training parameters of the speech activity detection model meet the second preset condition, thus obtaining the trained candidate speech activity detection model. The second preset condition may include the difference between the loss value and the previously determined loss value being less than the second preset threshold. It is worth noting that the difference between the loss value and the previously determined loss value being less than the second preset threshold can indicate that the model has converged.
[0079] By using the above method, the cross-entropy loss function is first used as the objective function of the speech activity detection model for training, thereby obtaining candidate speech activity detection models; then, the inverse of the smooth approximation of the F1 score is used as the objective function of the candidate speech activity detection model for training, thereby obtaining a trained speech activity detection model, which can solve the training divergence problem mentioned above.
[0080] Figure 2 This is a block diagram illustrating a speech activity detection model generation apparatus according to an exemplary embodiment. (Refer to...) Figure 2 The speech activity detection model generation device may include:
[0081] The acquisition module 201 is used to acquire a speech sample dataset, wherein the speech sample dataset includes speech sample frames and sample labels corresponding to the speech sample frames;
[0082] Input module 202 is used to input the speech sample frame into the speech activity detection model to obtain the speech probability result corresponding to the speech sample frame output by the speech activity detection model;
[0083] The determining module 203 is used to determine the loss value based on a preset loss function, the corresponding speech probability results and sample labels of the speech sample frames in the speech sample dataset, wherein the loss function includes a first loss function, which is the inverse of a smooth approximation of the F1 score;
[0084] The adjustment module 204 is used to adjust the model parameters of the speech activity detection model according to the loss value until the training parameters of the speech activity detection model meet the preset conditions to obtain the speech activity detection model.
[0085] In some embodiments, the loss function further includes a second loss function, wherein the second loss function is a cross-entropy loss function, and the determining module 203 includes:
[0086] The first determining submodule is used to determine a first loss value based on the first loss function, the corresponding speech probability results and sample labels of the speech sample frames in the speech sample dataset;
[0087] The second determining submodule is used to determine the second loss value based on the second loss function, the corresponding speech probability results and sample labels of the speech sample frames in the speech sample dataset;
[0088] The weighting submodule is used to weight the first loss value and the second loss value according to the weight relationship to obtain the loss value. The weight relationship is used to represent that the sum of the first weight corresponding to the first loss value and the second weight corresponding to the second loss value is a preset value. The first weight is directly proportional to the number of times the model parameters of the speech activity detection model have been adjusted, and the second weight is inversely proportional to the number of times the model parameters of the speech activity detection model have been adjusted.
[0089] In some embodiments, the second weight is characterized by the following formula:
[0090] ;
[0091] in, As the second weight, This represents the number of times the model parameters of the current speech activity detection model have been adjusted. The preset total number of times the model parameters of the speech activity detection model are adjusted, where e is a natural constant.
[0092] In some embodiments, the preset condition includes the number of times the model parameters of the voice activity detection model have been adjusted reaching the total number of times.
[0093] In some embodiments, the input module 202 is specifically used for:
[0094] The speech sample frame is input into the candidate speech activity detection model to obtain the speech probability result corresponding to the speech sample frame output by the candidate speech activity detection model. The candidate speech activity detection model is trained using the cross-entropy loss function as the objective function.
[0095] In some embodiments, the preset condition includes the difference between the loss value and the previously determined loss value being less than a first preset threshold.
[0096] The implementation methods of each module in the above-mentioned device can be referred to the above-mentioned related embodiments, and will not be repeated here.
[0097] This disclosure also provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the above-described method.
[0098] This disclosure also provides an electronic device, including:
[0099] A storage device having at least one computer program stored thereon;
[0100] At least one processing device is configured to execute the at least one computer program in the storage device to implement the steps of the above method.
[0101] The following is for reference. Figure 3 The diagram illustrates a structural schematic of an electronic device 300 suitable for implementing embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0102] like Figure 3 As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0103] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0104] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of embodiments of this disclosure.
[0105] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0106] In some implementations, electronic devices can communicate using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communications (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0107] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0108] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire a speech sample dataset, wherein the speech sample dataset includes speech sample frames and sample labels corresponding to the speech sample frames; input the speech sample frames into a speech activity detection model to obtain speech probability results output by the speech activity detection model corresponding to the speech sample frames; determine a loss value based on a preset loss function, the speech probability results corresponding to the speech sample frames in the speech sample dataset, and the sample labels, wherein the loss function includes a first loss function, which is the negative of a smooth approximation of the F1 score; and adjust the model parameters of the speech activity detection model based on the loss value until the parameters for evaluating the training of the speech activity detection model meet preset conditions to obtain a speech activity detection model.
[0109] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0110] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0111] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules are not, in some cases, intended to limit the functionality of the module itself.
[0112] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0113] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0114] According to one or more embodiments of this disclosure, Example 1 provides a method for generating a speech activity detection model, including:
[0115] Obtain a speech sample dataset, wherein the speech sample dataset includes speech sample frames and sample labels corresponding to the speech sample frames;
[0116] The speech sample frame is input into the speech activity detection model to obtain the speech probability result corresponding to the speech sample frame output by the speech activity detection model.
[0117] The loss value is determined based on the preset loss function, the corresponding speech probability results and sample labels of the speech sample frames in the speech sample dataset, wherein the loss function includes a first loss function, which is the inverse of a smooth approximation of the F1 score;
[0118] Based on the loss value, the model parameters of the speech activity detection model are adjusted until the training parameters of the speech activity detection model meet the preset conditions to obtain the speech activity detection model.
[0119] According to one or more embodiments of this disclosure, Example 2 provides the method of Example 1, wherein the loss function further includes a second loss function, the second loss function being a cross-entropy loss function, and the step of determining the loss value based on the preset loss function, the corresponding speech probability results of the speech sample frames in the speech sample dataset, and the sample labels includes:
[0120] The first loss value is determined based on the first loss function, the corresponding speech probability results and sample labels of the speech sample frames in the speech sample dataset;
[0121] The second loss value is determined based on the second loss function, the corresponding speech probability results and sample labels of the speech sample frames in the speech sample dataset;
[0122] The first loss value and the second loss value are weighted according to the weight relationship to obtain the loss value. The weight relationship is used to represent that the sum of the first weight corresponding to the first loss value and the second weight corresponding to the second loss value is a preset value. The first weight is directly proportional to the number of times the model parameters of the speech activity detection model have been adjusted, and the second weight is inversely proportional to the number of times the model parameters of the speech activity detection model have been adjusted.
[0123] According to one or more embodiments of this disclosure, Example 3 provides the method of Example 2, wherein the second weight is characterized by the following formula:
[0124] ;
[0125] in, As the second weight, This represents the number of times the model parameters of the current speech activity detection model have been adjusted. The preset total number of times the model parameters of the speech activity detection model are adjusted, where e is a natural constant.
[0126] According to one or more embodiments of this disclosure, Example 4 provides the method of Example 3, wherein the preset condition includes the number of times the model parameters of the speech activity detection model have been adjusted reaching the total number of times.
[0127] According to one or more embodiments of this disclosure, Example 5 provides the method of Example 1, wherein inputting the speech sample frame into a speech activity detection model to obtain a speech probability result corresponding to the speech sample frame output by the speech activity detection model includes:
[0128] The speech sample frame is input into the candidate speech activity detection model to obtain the speech probability result corresponding to the speech sample frame output by the candidate speech activity detection model. The candidate speech activity detection model is trained using the cross-entropy loss function as the objective function.
[0129] According to one or more embodiments of this disclosure, Example 6 provides the method of Example 5, wherein the preset condition includes that the difference between the loss value and the previously determined loss value is less than a first preset threshold.
[0130] According to one or more embodiments of this disclosure, Example 7 provides a speech activity detection model generation apparatus, comprising:
[0131] An acquisition module is used to acquire a speech sample dataset, wherein the speech sample dataset includes speech sample frames and sample labels corresponding to the speech sample frames;
[0132] The input module is used to input the speech sample frame into the speech activity detection model to obtain the speech probability result corresponding to the speech sample frame output by the speech activity detection model.
[0133] The determination module is used to determine the loss value based on a preset loss function, the corresponding speech probability results and sample labels of the speech sample frames in the speech sample dataset, wherein the loss function includes a first loss function, which is the inverse of a smooth approximation of the F1 score;
[0134] An adjustment module is used to adjust the model parameters of the speech activity detection model according to the loss value until the training parameters of the speech activity detection model meet preset conditions to obtain the speech activity detection model.
[0135] According to one or more embodiments of this disclosure, Example 8 provides the apparatus of Example 7, wherein the loss function further includes a second loss function, the second loss function being a cross-entropy loss function, and the determining module includes:
[0136] The first determining submodule is used to determine a first loss value based on the first loss function, the corresponding speech probability results and sample labels of the speech sample frames in the speech sample dataset;
[0137] The second determining submodule is used to determine the second loss value based on the second loss function, the corresponding speech probability results and sample labels of the speech sample frames in the speech sample dataset;
[0138] The weighting submodule is used to weight the first loss value and the second loss value according to the weight relationship to obtain the loss value. The weight relationship is used to represent that the sum of the first weight corresponding to the first loss value and the second weight corresponding to the second loss value is a preset value. The first weight is directly proportional to the number of times the model parameters of the speech activity detection model have been adjusted, and the second weight is inversely proportional to the number of times the model parameters of the speech activity detection model have been adjusted.
[0139] According to one or more embodiments of this disclosure, Example 9 provides the apparatus of Example 8, wherein the second weight is characterized by the following formula:
[0140] ;
[0141] in, As the second weight, This represents the number of times the model parameters of the current speech activity detection model have been adjusted. The preset total number of times the model parameters of the speech activity detection model are adjusted, where e is a natural constant.
[0142] According to one or more embodiments of this disclosure, Example 10 provides the apparatus of Example 9, wherein the preset condition includes the number of times the model parameters of the speech activity detection model have been adjusted reaching the total number of times.
[0143] According to one or more embodiments of this disclosure, Example 11 provides the apparatus of Example 7, wherein inputting the speech sample frame into a speech activity detection model to obtain a speech probability result corresponding to the speech sample frame output by the speech activity detection model includes:
[0144] The speech sample frame is input into the candidate speech activity detection model to obtain the speech probability result corresponding to the speech sample frame output by the candidate speech activity detection model. The candidate speech activity detection model is trained using the cross-entropy loss function as the objective function.
[0145] According to one or more embodiments of this disclosure, Example 12 provides the apparatus of Example 11, wherein the preset condition includes that the difference between the loss value and a previously determined loss value is less than a first preset threshold.
[0146] According to one or more embodiments of the present disclosure, Example 13 provides a computer-readable medium having a computer program stored thereon that, when executed by a processing device, implements the steps of the method described in any one of Examples 1-6.
[0147] According to one or more embodiments of this disclosure, Example 14 provides an electronic device comprising:
[0148] A storage device having at least one computer program stored thereon;
[0149] At least one processing device is configured to execute the at least one computer program in the storage device to implement the steps of the method described in any one of Examples 1-6.
[0150] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0151] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0152] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.
Claims
1. A method for generating a speech activity detection model, characterized in that, include: Obtain a speech sample dataset, wherein the speech sample dataset includes speech sample frames and sample labels corresponding to the speech sample frames; The speech sample frame is input into the speech activity detection model to obtain the speech probability result corresponding to the speech sample frame output by the speech activity detection model. The loss value is determined based on the preset loss function, the corresponding speech probability results and sample labels of the speech sample frames in the speech sample dataset, wherein the loss function includes a first loss function, which is the negative of the smooth approximation of the F1 score, and the smooth approximation of the F1 score satisfies continuous differentiability; Based on the loss value, the model parameters of the speech activity detection model are adjusted until the training parameters of the speech activity detection model meet the preset conditions to obtain the speech activity detection model.
2. The method according to claim 1, characterized in that, The loss function further includes a second loss function, which is a cross-entropy loss function. The step of determining the loss value based on the preset loss function, the corresponding speech probability results of the speech sample frames in the speech sample dataset, and the sample labels includes: The first loss value is determined based on the first loss function, the corresponding speech probability results and sample labels of the speech sample frames in the speech sample dataset; The second loss value is determined based on the second loss function, the corresponding speech probability results and sample labels of the speech sample frames in the speech sample dataset; The first loss value and the second loss value are weighted according to the weight relationship to obtain the loss value. The weight relationship is used to represent that the sum of the first weight corresponding to the first loss value and the second weight corresponding to the second loss value is a preset value. The first weight is directly proportional to the number of times the model parameters of the speech activity detection model have been adjusted, and the second weight is inversely proportional to the number of times the model parameters of the speech activity detection model have been adjusted.
3. The method according to claim 2, characterized in that, The second weight is represented by the following formula: ; in, As the second weight, This represents the number of times the model parameters of the current speech activity detection model have been adjusted. The preset total number of times the model parameters of the speech activity detection model are adjusted, where e is a natural constant.
4. The method according to claim 3, characterized in that, The preset conditions include the number of times the model parameters of the voice activity detection model have been adjusted reaching the total number of times.
5. The method according to claim 1, characterized in that, The step of inputting the speech sample frame into the speech activity detection model to obtain the speech probability result corresponding to the speech sample frame output by the speech activity detection model includes: The speech sample frame is input into the candidate speech activity detection model to obtain the speech probability result corresponding to the speech sample frame output by the candidate speech activity detection model. The candidate speech activity detection model is trained using the cross-entropy loss function as the objective function.
6. The method according to claim 5, characterized in that, The preset conditions include that the difference between the loss value and the previously determined loss value is less than a first preset threshold.
7. A speech activity detection model generation device, characterized in that, include: An acquisition module is used to acquire a speech sample dataset, wherein the speech sample dataset includes speech sample frames and sample labels corresponding to the speech sample frames; The input module is used to input the speech sample frame into the speech activity detection model to obtain the speech probability result corresponding to the speech sample frame output by the speech activity detection model. The determination module is used to determine the loss value based on a preset loss function, the corresponding speech probability results and sample labels of the speech sample frames in the speech sample dataset, wherein the loss function includes a first loss function, which is the negative of the smooth approximation of the F1 score, and the smooth approximation of the F1 score satisfies continuous differentiability; An adjustment module is used to adjust the model parameters of the speech activity detection model according to the loss value until the training parameters of the speech activity detection model meet preset conditions to obtain the speech activity detection model.
8. The apparatus according to claim 7, characterized in that, The loss function further includes a second loss function, which is a cross-entropy loss function. The determining module includes: The first determining submodule is used to determine a first loss value based on the first loss function, the corresponding speech probability results and sample labels of the speech sample frames in the speech sample dataset; The second determining submodule is used to determine the second loss value based on the second loss function, the corresponding speech probability results and sample labels of the speech sample frames in the speech sample dataset; The weighting submodule is used to weight the first loss value and the second loss value according to the weight relationship to obtain the loss value. The weight relationship is used to represent that the sum of the first weight corresponding to the first loss value and the second weight corresponding to the second loss value is a preset value. The first weight is directly proportional to the number of times the model parameters of the speech activity detection model have been adjusted, and the second weight is inversely proportional to the number of times the model parameters of the speech activity detection model have been adjusted.
9. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by the processing device, the program implements the steps of the method according to any one of claims 1-6.
10. An electronic device, characterized in that, include: A storage device having at least one computer program stored thereon; At least one processing device is configured to execute the at least one computer program in the storage device to implement the steps of the method according to any one of claims 1-6.
Citation Information
Patent Citations
Model generation method and device
CN109545192A
Speech recognition and model training method, device and computer readable storage medium
CN111261146A