Voice activity detection model training method, voice activity detection method and related device

By training a speech activity detection model with semantic integrity discrimination function, the problem of poor voice activity detection effect in the prior art is solved, especially in segmenting semantic complete speech segments, which achieves higher detection accuracy and subsequent processing effects.

CN120199284APending Publication Date: 2025-06-24XIAN XUNFEI SUPER BRAIN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510417647.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-06-24

Smart Images

  • Figure CN120199284A_ABST
    Figure CN120199284A_ABST
Patent Text Reader

Abstract

The invention discloses a voice activity detection model training method, a voice activity detection method and a related device, and relates to the technical field of audio processing, and the training method comprises the steps: carrying out the training of a first training audio marked with a frame-level audio class, and obtaining a first voice activity detection model with a semantic integrity discrimination function, the audio category of an audio frame of the first training audio is one of voice, non-voice at a semantic incomplete position and non-voice after semantic complete voice; a second training audio marked with a frame-level audio category is utilized, the first voice activity detection model is assisted, a second voice activity detection model capable of capturing semantic information in the voice is obtained through training, and the audio category of an audio frame of the second training audio is one of voice and non-voice. The voice activity detection model trained by the training method disclosed by the invention can capture semantic information of the audio, so that a reasonable category can be given for each audio frame of the audio by referring to the semantic information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio processing, and in particular, to a method for training a voice activity detection model, a voice activity detection method, and related devices. Background Art

[0002] With the rapid development of speech technology, applications such as speech recognition, speech synthesis, and speech communication have been widely used in multiple industries. In these applications, it is crucial to effectively distinguish speech segments from non-speech segments. For example, in a speech recognition system, processing non-speech parts may lead to misrecognition and reduce the accuracy of the system. Similarly, in a speech communication system (such as a telephone or VoIP system), transmitting non-speech data will increase bandwidth consumption and affect communication quality.

[0003] Voice activity detection (VAD) is a speech processing technology that can distinguish speech segments from non-speech segments. Current voice activity detection methods all make judgments based on the differences between speech and non-speech to achieve the distinction between speech segments and non-speech segments. However, the detection effect of the voice activity detection scheme that distinguishes speech segments from non-speech segments based on the differences between speech and non-speech is not good. Summary of the Invention

[0004] In view of this, the present application provides a method for training a voice activity detection model, a voice activity detection method, and related devices, and its technical solutions are as follows:

[0005] The first aspect of the present application provides a method for training a voice activity detection model, including:

[0006] Using a first training audio labeled with frame-level audio categories to train a first voice activity detection model with a semantic integrity discrimination function, where the audio category of any audio frame of the first training audio is one of speech, non-speech at the semantic incomplete part, and non-speech after the semantically complete speech;

[0007] Using a second training audio labeled with frame-level audio categories, supplemented by the first voice activity detection model, to train a second voice activity detection model capable of capturing semantic information in speech, and the trained second voice activity detection model is used as the target voice activity detection model, where the audio category of any audio frame of the second training audio is one of speech and non-speech.

[0008] In a possible implementation manner, the using a first training audio labeled with frame-level audio categories to train a first voice activity detection model with a semantic integrity discrimination function includes:

[0009] Using the first training audio labeled with frame-level audio categories and corresponding texts, in the speech activity detection task combined with semantic information, and simultaneously jointly with the speech recognition task, train to obtain a first speech activity detection model with a semantic integrity discrimination function.

[0010] In a possible implementation manner, the step of using the first training audio labeled with frame-level audio categories and corresponding texts, in the speech activity detection task combined with semantic information, and simultaneously jointly with the speech recognition task, to train and obtain a first speech activity detection model with a semantic integrity discrimination function includes:

[0011] Extract audio features from the first training audio to obtain first audio features;

[0012] Use the first speech activity detection model to encode the first audio features to obtain first audio encoded features;

[0013] Use the first speech activity detection model to obtain word-level speech text features corresponding to the first training audio according to the first audio encoded features;

[0014] Use the first speech activity detection model to process the first audio encoded features and the word-level speech text features corresponding to the first training audio into features containing the semantic information of the first training audio to obtain target features;

[0015] Taking the frame-level audio categories predicted by the first speech activity detection model according to the target features to be consistent with the frame-level audio categories labeled in the first training audio, and the texts predicted by the first speech activity detection model according to the word-level speech text features corresponding to the first training audio to be consistent with the texts labeled in the first training audio as the goal, update the parameters of the first speech activity detection model.

[0016] In a possible implementation manner, the step of using the first speech activity detection model to obtain word-level speech text features corresponding to the first training audio according to the first audio encoded features includes:

[0017] Use the connectionist temporal classification (CTC) module of the first speech activity detection model to obtain word-level first speech text features corresponding to the first training audio according to the first audio encoded features;

[0018] And / or, use the autoregressive decoder of the first speech activity detection model to obtain word-level second speech text features corresponding to the first training audio according to the first audio encoded features.

[0019] In a possible implementation, with the goal of making the frame-level audio categories predicted by the first voice activity detection model based on the target features tend to be consistent with the frame-level audio categories annotated for the first training audio, and making the text predicted by the first voice activity detection model based on the word-level speech text features corresponding to the first training audio tend to be consistent with the text annotated for the first training audio, the parameter update of the first voice activity detection model includes:

[0020] Using the first voice activity detection model, based on the target features, obtain the frame-level category prediction probabilities of the first training audio, where the category prediction probability of each audio frame of the first training audio includes the prediction probabilities corresponding to the three categories of non-speech at the incomplete speech and semantics part and non-speech after the complete speech and semantics of this audio frame;

[0021] Using the first voice activity detection model, based on the word-level first speech text features, obtain the first text prediction probability; and / or, using the first voice activity detection model, based on the word-level second speech text features, obtain the second text prediction probability;

[0022] Based on the frame-level category prediction probabilities of the first training audio and the frame-level audio categories annotated for the first training audio, determine the first category prediction loss;

[0023] Based on the first text prediction probability and the text annotated for the first training audio, determine the first text prediction loss; and / or, based on the second text prediction probability and the text annotated for the first training audio, determine the second text prediction loss;

[0024] Based on the first category prediction loss, and simultaneously combining the first text prediction loss and / or the second text prediction loss, perform parameter update on the first voice activity detection model.

[0025] In a possible implementation, using the second training audio annotated with frame-level audio categories, supplemented with the first voice activity detection model, training to obtain a second voice activity detection model capable of capturing semantic information in speech includes:

[0026] Extract audio features from the second training audio to obtain second audio features;

[0027] Use the first voice activity detection model to encode the second audio features to obtain second audio encoded features;

[0028] According to a preset feature selection strategy, select one feature from the second audio features and the second audio encoded features;

[0029] Train the second voice activity detection model by using the selected features and the frame-level audio categories labeled for the second training audio, supplemented with the first voice activity detection model.

[0030] In a possible implementation, the training of the second voice activity detection model by using the selected features and the frame-level audio categories labeled for the second training audio, supplemented with the first voice activity detection model, includes:

[0031] Use the first voice activity detection model to obtain the frame-level first category prediction probabilities of the second training audio according to the second audio coding features, where the first category prediction probabilities of each audio frame of the second training audio include the prediction probabilities corresponding to the three categories of speech, non-speech at the semantically incomplete part, and non-speech after semantically complete speech for this audio frame;

[0032] Use the second voice activity detection model to obtain the frame-level second category prediction probabilities of the second training audio according to the selected features, where the second category prediction probabilities of each audio frame of the second training audio include the prediction probabilities corresponding to the two categories of speech and non-speech for this audio frame;

[0033] Adjust the frame-level second category prediction probabilities of the second training audio according to the frame-level first category prediction probabilities of the second training audio to obtain the adjusted frame-level second category prediction probabilities of the second training audio;

[0034] Determine the second category prediction loss according to the adjusted frame-level second category prediction probabilities of the second training audio and the frame-level audio categories labeled for the second training audio;

[0035] Update the parameters of the second voice activity detection model according to the second category prediction loss.

[0036] In a possible implementation, the adjustment of the frame-level second category prediction probabilities of the second training audio according to the frame-level first category prediction probabilities of the second training audio includes:

[0037] Determine the first audio category of each audio frame of the second training audio according to the frame-level first category prediction probabilities of the second training audio, and determine the second audio category of each audio frame of the second training audio according to the frame-level second category prediction probabilities of the second training audio;

[0038] For each audio frame of the second training audio, adjust the second category prediction probability of this audio frame according to the first audio category and the second audio category of this audio frame.

[0039] In a possible implementation, adjusting the second-class prediction probability of the audio frame according to the first audio category and the second audio category of the audio frame includes:

[0040] If both the first audio category and the second audio category of the audio frame are speech, increase the second-class prediction probability corresponding to the audio frame in terms of speech;

[0041] If the first audio category of the audio frame is non-speech after semantically complete speech and the second audio category of the audio frame is non-speech, increase the second-class prediction probability corresponding to the audio frame in terms of non-speech;

[0042] If the first audio category of the audio frame is non-speech at the semantically incomplete part and the second audio category of the audio frame is non-speech, decrease the second-class prediction probability corresponding to the audio frame in terms of non-speech.

[0043] In a possible implementation, increasing the second-class prediction probability corresponding to the audio frame in terms of speech includes:

[0044] Increase the second-class prediction probability corresponding to the audio frame in terms of speech by multiplying the second-class prediction probability corresponding to the audio frame in terms of speech by a preset first coefficient;

[0045] Increasing the second-class prediction probability corresponding to the audio frame in terms of non-speech includes:

[0046] Increase the second-class prediction probability corresponding to the audio frame in terms of non-speech by multiplying the second-class prediction probability corresponding to the audio frame in terms of non-speech by a preset first coefficient;

[0047] Decreasing the second-class prediction probability corresponding to the audio frame in terms of non-speech includes:

[0048] Decrease the second-class prediction probability corresponding to the audio frame in terms of non-speech by multiplying the second-class prediction probability corresponding to the audio frame in terms of non-speech by a preset second coefficient.

[0049] A second aspect of the present application provides a voice activity detection method, including:

[0050] Obtain a target audio;

[0051] Perform voice activity detection on the target audio by using a target voice activity detection model trained by using any one of the above voice activity detection model training methods.

[0052] A third aspect of the present application provides a voice activity detection model training device, including: a first training module and a second training module;

[0053] The first training module is used to train a first voice activity detection model with a semantic integrity discrimination function by using first training audio labeled with frame-level audio categories. Among them, the audio category of any audio frame of the first training audio is one of speech, non-speech at the place where semantics is incomplete, and non-speech after semantically complete speech;

[0054] The second training module is used to train a second voice activity detection model capable of capturing semantic information in speech by using second training audio labeled with frame-level audio categories and supplemented with the first voice activity detection model. The trained second voice activity detection model is used as the target voice activity detection model. Among them, the audio category of any audio frame of the second training audio is one of speech and non-speech.

[0055] A fifth aspect of the present application provides a voice activity detection device, including: an audio acquisition module and a voice activity detection module;

[0056] The audio acquisition module is used to acquire target audio;

[0057] The voice activity detection module is used to perform voice activity detection on the target audio by using the target voice activity detection model trained by the above voice activity detection model training device.

[0058] A sixth aspect of the present application provides an electronic device, including at least one processor and a memory connected to the processor, where:

[0059] The memory is used to store a computer program;

[0060] The processor is used to execute the computer program so that the electronic device can implement the steps of any of the above voice activity detection model training methods, and / or implement the steps of the above voice activity detection method.

[0061] A seventh aspect of the present application provides a computer storage medium, where the storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can be enabled to implement the steps of any of the above voice activity detection model training methods, and / or implement the steps of the above voice activity detection method.

[0062] An eighth aspect of the present application provides a computer program product, including computer-readable instructions. When the computer-readable instructions run on an electronic device, the electronic device is enabled to implement the steps of any of the above voice activity detection model training methods, and / or implement the steps of the above voice activity detection method.

[0063] With the above technical solution, the method for training a voice activity detection model provided by this application first uses the first training audio labeled with frame-level audio categories (voice / non-voice at the semantically incomplete part / non-voice after the semantically complete voice) to train a first voice activity detection model with a semantic integrity discrimination function. Then, using the second training audio labeled with frame-level audio categories (voice / non-voice), supplemented by the first voice activity detection model, a second voice activity detection model capable of capturing semantic information in the voice is trained. The trained second voice activity detection model is used as the target voice activity detection model to perform voice activity detection on the audio to be detected. The target voice activity detection model trained by the method for training a voice activity detection model provided by this application can capture the semantic information of the voice in the input audio, and then can give a reasonable category for each audio frame of the input audio with reference to the semantic information, so that the voice segments with complete semantics and short pauses in the middle can be segmented into the same voice segment as much as possible, thereby improving the effect of subsequent processing (such as speech recognition). Based on the method for training a voice activity detection model provided by this application, this application also provides a voice activity detection method. Since this voice activity detection method uses the target voice activity detection model trained by the method for training a voice activity detection model provided by this application to perform voice activity detection on the target audio, a more reasonable detection result can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the provided drawings without creative efforts.

[0065] Figure 1 It is a schematic diagram of a system architecture related to this application;

[0066] Figure 2 It is a schematic diagram of a hardware structure of a terminal provided by an embodiment of this application;

[0067] Figure 3 It is a schematic diagram of a hardware structure of a server provided by an embodiment of this application;

[0068] Figure 4 It is a schematic diagram of the flow of the method for training a voice activity detection model provided by an embodiment of this application;

[0069] Figure 5Schematic flow chart of the first voice activity detection model with a semantic integrity discrimination function trained by using the first training audio labeled with frame-level audio categories and corresponding texts, in combination with the voice activity detection task with semantic information and jointly with the speech recognition task, provided by an embodiment of the present application;

[0070] Figure 6 Schematic structural diagram of the first voice activity detection model provided by an embodiment of the present application;

[0071] Figure 7 Schematic flow chart of the second voice activity detection model capable of capturing semantic information in speech trained by using the second training audio labeled with frame-level audio categories, supplemented by the first voice activity detection model, provided by an embodiment of the present application;

[0072] Figure 8 Schematic diagram of training the second voice activity detection model supplemented by the first voice activity detection model provided by an embodiment of the present application;

[0073] Figure 9 Schematic structural diagram of the voice activity detection model training device provided by an embodiment of the present application;

[0074] Figure 10 Schematic structural diagram of the voice activity detection device provided by an embodiment of the present application. Detailed implementation manners

[0075] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. The terms used in the embodiment part of the present application are only used to explain the specific embodiments of the present application, rather than intended to limit the present application.

[0076] The embodiments of the present application will be described below with reference to the accompanying drawings. Those skilled in the art will know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0077] The terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing when describing objects with the same attributes in the embodiments of the present application. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device comprising a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.

[0078] In one possible implementation manner, asFigure 1 As shown in the figure, the system architecture involved in this application may include a terminal 101 and a server 102. The terminal 101 can interact with the server 102 through a network (wired network or wireless network). Among them, the server 102 may include one or more servers ( Figure 1 In the following, an example of including one server is used for illustration). The terminal can obtain training data and transmit the training data to the server through the network. The server uses the voice activity detection model training method provided by this application to train and obtain a target voice activity detection model. In one possible implementation, the server can use the trained target voice activity detection model to perform voice activity detection on the target audio. In another possible implementation, the trained target voice activity detection model can be deployed to the terminal, and the terminal uses the target voice activity detection model to perform voice activity detection on the target audio.

[0079] In another possible implementation, the system architecture involved in this application may include a terminal. The terminal has strong data processing capabilities. The terminal can use the voice activity detection model training method provided by this application to train and obtain a target voice activity detection model, and then use the target voice activity detection model to perform voice activity detection on the target audio.

[0080] Next, the product form of the above terminal is described.

[0081] The above terminal can be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, a robot, an augmented reality (AR) / virtual reality (VR) device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. The embodiments of this application do not impose any restrictions on this.

[0082] Figure 2 Shows an optional schematic diagram of the hardware structure of the terminal.

[0083] Refer to Figure 2 As shown in the figure, the terminal may include a radio frequency unit 210, a memory 220, an input unit 230, a display unit 240, a camera 250 (optional), an audio circuit 260 (optional), a speaker 261 (optional), a microphone 262 (optional), a headphone jack 263 (optional), a processor 270, an external interface 280, a power supply 290, and other components. Those skilled in the art can understand that, Figure 2The above are merely examples of terminals and do not constitute a limitation thereto. A terminal may include more or fewer components than those shown in the figures, or combine certain components, or have different components.

[0084] The input unit 230 can be used to receive input numerical or character information, and generate key signal inputs related to user settings and function controls of the terminal. Specifically, the input unit 230 may include a touch screen 231 (optional) and / or other input devices 232. The touch screen 231 can collect touch operations of a user thereon or nearby (such as operations of the user using any suitable object such as a finger, a joint, a stylus, etc. on or near the touch screen), and drive corresponding connection devices according to a pre-set program. The touch screen can detect a touch action of the user on the touch screen, convert the touch action into a touch signal and send it to the processor 270, and can receive and execute commands sent by the processor 270; the touch signal at least includes contact coordinate information. The touch screen 231 can provide an input interface and an output interface between the terminal and the user. In addition, various types such as a resistive type, a capacitive type, an infrared type, and a surface acoustic wave type can be used to implement the touch screen. In addition to the touch screen 231, the input unit 230 may further include other input devices. Specifically, the other input devices 232 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, a joystick, etc.

[0085] The display unit 240 can be used to display information input by the user or information provided to the user, various menus of the terminal, an interactive interface, file display, and / or playback of any one multimedia file.

[0086] The memory 220 can be used to store instructions and data. The memory 220 mainly includes a storage instruction area and a storage data area. The storage data area can store various data, such as multimedia files, texts, etc.; the storage instruction area can store software units such as an operating system, an application, instructions required for at least one function, or their subsets or extended sets. It may also include a non-volatile random access memory; it provides the processor 270 with management of hardware, software, and data resources in the computing processing device, supports control software and applications. It is also used for storage of multimedia files, and storage of running programs and applications.

[0087] The processor 270 is the control center of the terminal. It connects various parts of the entire terminal using various interfaces and circuits. By running or executing the instructions stored in the memory 220 and invoking the data stored in the memory 220, it performs various functions of the terminal and processes data, thereby exercising overall control over the terminal. Optionally, the processor 270 may include one or more processing units; preferably, the processor 270 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 270 either. In some embodiments, the processor and the memory may be implemented on a single chip, and in some embodiments, they may also be separately implemented on independent chips. The processor 270 can also be used to generate corresponding operation control signals, send them to the corresponding components of the computing and processing device, read and process the data in the software, especially read and process the data and programs in the memory 220, so that each functional module therein executes the corresponding functions, thereby controlling the corresponding components to act according to the requirements of the instructions.

[0088] Among them, the memory 220 can be used to store software codes related to the voice activity detection model training method and the voice activity detection method. The processor 270 can execute the software codes in the memory 220 or can also schedule other units (such as the above-mentioned input unit 230 and display unit 240) to implement the corresponding functions.

[0089] The radio frequency unit 210 (optional) can be used for receiving and transmitting information or signals during a call. For example, after receiving the downlink information from the base station, it is sent to the processor 270 for processing. Additionally, the data designed for uplink is sent to the base station. Generally, the radio frequency unit 210 includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the radio frequency unit 210 can also communicate with network devices and other devices through wireless communication. This wireless communication can use any communication standard or protocol, including but not limited to Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0090] Among them, in the embodiments of the present application, the radio frequency unit 210 can send data to other devices and can also receive data sent by other devices. It should be understood that the radio frequency unit 210 is optional and can be replaced by other communication interfaces, such as a network interface.

[0091] The terminal also includes a power supply 290 (such as a battery) for powering each component. Preferably, the power supply can be logically connected to the processor 270 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system.

[0092] The terminal also includes an external interface 280. This external interface can be a standard Micro USB interface or a multi-pin connector, and can be used to connect the terminal to other devices for communication and can also be used to connect a charger to charge the terminal.

[0093] Although not shown, the terminal may also include a flashlight, a wireless fidelity (WiFi) module, a Bluetooth module, sensors with different functions, etc., which will not be elaborated here.

[0094] Next, the product form of the above server will be described.

[0095] Figure 3 A schematic structural diagram of the above server is provided, as Figure 3As shown, the server may include a bus 301, a processor 302, a communication interface 303, and a memory 304. The processor 302, the memory 304, and the communication interface 303 communicate with each other via the bus 301.

[0096] The bus 301 may be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 3 only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0097] The processor 302 can be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0098] The memory 304 may include volatile memory, such as random access memory (RAM). The memory 304 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0099] The memory 304 can be used to store software codes related to the voice activity detection model training method and the voice activity detection method. The processor 302 can call the software codes stored in the memory 304 or schedule other units to implement the corresponding functions.

[0100] The processors in the above-mentioned terminal and server (such as processor 270 and processor 302) can be hardware circuits (such as application specific integrated circuit (ASIC), field-programmable gate array (FPGA), general-purpose processor, digital signal processing (DSP), microprocessor or microcontroller, etc.), or a combination of these hardware circuits. For example, the processor can be a hardware system with the function of executing instructions, such as CPU, DSP, etc., or a hardware system without the function of executing instructions, such as ASIC, FPGA, etc., or a combination of the above hardware systems without the function of executing instructions and hardware systems with the function of executing instructions.

[0101] Voice activity detection is widely used in various scenarios, such as speech recognition scenarios, voice communication scenarios, etc.

[0102] Taking the speech recognition scenario as an example, in a speech recognition system, voice activity detection, as a key preprocessing module, can effectively identify speech segments in a speech signal and filter out non-speech signals (such as silence, ambient noise, etc.). By using voice activity detection, the speech recognition system can obtain multiple advantages: First, voice activity detection helps reduce the waste of computing resources. The system only processes when voice activity is detected, significantly reducing the computing burden of the system. Second, voice activity detection can filter out background noise and silence parts, preventing these invalid signals from interfering with the speech recognition process, thereby improving the recognition accuracy. Third, voice activity detection can also save storage and bandwidth resources in long-term speech recording or real-time processing applications, enabling the system to only store and transmit valid speech segments. Fourth, voice activity detection also has the ability to adapt to different noise environments, ensuring that the system can still accurately recognize speech in a noisy or complex acoustic environment, thereby enhancing the robustness and practicality of the system. In summary, voice activity detection plays a crucial role in speech recognition, and it can effectively improve the overall performance of the system.

[0103] Existing voice activity detection methods include rule-based voice activity detection methods and voice activity detection methods based on machine learning and deep learning.

[0104] Among them, rule-based voice activity detection methods include energy-based voice activity detection methods, zero-crossing rate (ZCR)-based voice activity detection methods, spectral entropy-based voice activity detection methods, etc.

[0105] The energy-based voice activity detection method distinguishes speech from non-speech based on the energy level of the audio. Specifically, the energy value of each audio frame is calculated and compared with a set threshold. When the energy value exceeds the set threshold, it is judged as speech.

[0106] The zero-crossing rate-based voice activity detection method distinguishes speech from non-speech based on the zero-crossing rate distribution of the audio. The zero-crossing rate refers to the number of times the signal waveform changes from positive to negative or from negative to positive on the time axis. The zero-crossing rate distributions of speech and noise are usually different. Therefore, the zero-crossing rate can be used to detect voice activity.

[0107] The spectral entropy-based voice activity detection method distinguishes speech from non-speech by analyzing the spectral entropy value of the audio signal. This method can distinguish relatively complex speech signals from more random noise signals.

[0108] The above rule-based voice activity detection methods mainly rely on simple statistical features such as the energy, zero-crossing rate, and spectral entropy of the signal. These methods distinguish speech from non-speech by setting fixed thresholds and are suitable for low-noise environments, but their performance is limited in complex noise scenarios and they are prone to false detections and missed detections.

[0109] The machine learning-based voice activity detection method extracts various features of the audio (such as Mel Frequency Cepstral Coefficients MFCC, Linear Prediction Coefficients LPC, etc.) and uses classifiers (such as Support Vector Machine SVM, Decision Tree, etc.) to distinguish speech from non-speech. Compared with the rule-based voice activity detection method, the machine learning-based voice activity detection method can better adapt to different noise environments.

[0110] The deep learning-based voice activity detection method uses deep learning models such as Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), and Long Short-Term Memory Network (LSTM) to implement voice activity detection. The deep learning model can automatically learn the complex features of speech and noise from a large amount of labeled data, significantly improving the accuracy and robustness of the voice activity detection method. Especially when facing complex acoustic scenarios, the deep learning-based voice activity detection method has strong adaptability.

[0111] In the process of implementing this case, the inventors of this case found that whether it is a rule-based voice activity detection method or a voice activity detection method based on machine learning and deep learning, when performing voice activity detection, it is judged based on the difference between voice and non-voice. However, judging based on the difference between voice and non-voice may split a semantically complete voice into two parts. When such a result is sent to the subsequent speech recognition system, it may lead to incorrect recognition results. For example, the content of an audio is "An atom contains neutrons and protons". If there is a short pause after "In an atom", the current voice activity detection method may split this audio into two audio segments with the content of "In an atom" and the content of "contains neutrons and protons". When these two audio segments are sent to the subsequent speech recognition system, the audio segment with the content of "In an atom" may be recognized as "atomic clock".

[0112] In view of the problems existing in the existing voice activity detection methods, the inventors of this case conducted research. Through continuous research, they finally proposed a method for training a voice activity detection model and a voice activity detection method. Next, the voice activity detection model training method and the voice activity detection method provided by this application will be introduced through the following embodiments.

[0113] Please refer to Figure 4 , which shows a schematic flowchart of the voice activity detection model training method provided by an embodiment of this application. The voice activity detection model training method may include:

[0114] Step S401: Use the first training audio labeled with frame-level audio categories to train a first voice activity detection model with a semantic integrity discrimination function.

[0115] Among them, the audio category labeled for any audio frame of the first training audio is one of voice, non-voice at the semantically incomplete part, and non-voice after the semantically complete voice.

[0116] The first training audio in this embodiment is the training audio used to train the first voice activity detection model. When annotating each audio frame of the first training audio, for non-voice frames, instead of simply labeling their category as non-voice, semantic information is considered, and their category is labeled as one of non-voice at the semantically incomplete part and non-voice after the semantically complete voice.

[0117] In a possible implementation, the audio category of the first training audio annotation can be obtained in the following way: obtain the text corresponding to the first training audio; use a large model to score the integrity of the text corresponding to the first training audio to obtain the integrity score of the text corresponding to the first training audio; determine whether the text corresponding to the first training audio is complete according to the integrity score of the text corresponding to the first training audio; for the first training audio, obtain the speech / non-speech category of each audio frame through forced alignment, and for each audio frame with the category of non-speech, further determine the final category of the audio frame from the two categories of non-speech at the semantically incomplete part and non-speech after the semantically complete speech according to the integrity determination result of the text corresponding to the first training audio.

[0118] There are various implementation ways to train a first voice activity detection model with semantic integrity discrimination function by using the first training audio annotated with frame-level audio categories. The present embodiment provides the following two implementation ways:

[0119] The first implementation way: The first training audio annotated with frame-level audio categories can be used to train a first voice activity detection model with semantic integrity discrimination function in the voice activity detection task combined with semantic information. In this implementation way, the training objective of the first voice activity detection model is to make the audio category predicted by the first voice activity detection model for each audio frame of the first training audio tend to be consistent with the corresponding annotated category.

[0120] The second implementation way: In order to obtain a first voice activity detection model with better performance, the first training audio annotated with frame-level audio categories and the corresponding text can be used to jointly train a first voice activity detection model with semantic integrity discrimination function in the voice activity detection task combined with semantic information and the speech recognition task at the same time. In this implementation way, the training objectives of the first voice activity detection model are to make the audio category predicted by the first voice activity detection model for each audio frame of the first training audio tend to be consistent with the corresponding annotated category, and to make the text predicted by the first voice activity detection model for the first training audio tend to be consistent with the text annotated in the first training audio.

[0121] Step S402: Use the second training audio annotated with frame-level audio categories, supplemented by the first voice activity detection model, to train a second voice activity detection model capable of capturing semantic information in speech. The trained second voice activity detection model is used as the target voice activity detection model.

[0122] Among them, the audio category of any audio frame of the second training audio is one of speech and non-speech.

[0123] The second training audio in this embodiment is the training audio for training the second voice activity detection model. When annotating each audio frame of the second training audio, its category is annotated as one of voice and non-voice.

[0124] In order to enable the second voice activity detection model to capture the semantic information in the voice, and then be able to give a reasonable category for each audio frame of the input audio with reference to the semantic information, this embodiment uses the second training audio annotated with frame-level audio categories, and at the same time supplements the second voice activity detection model with the first voice activity detection model for training. The trained second voice activity detection model is used as the target voice activity detection model for finally performing voice activity detection on the audio to be detected.

[0125] The voice activity detection model training method provided by the embodiments of this application first uses the first training audio annotated with frame-level audio categories (voice / non-voice at the incomplete semantic part / non-voice after the voice with complete semantics) to train the first voice activity detection model with the function of judging semantic integrity. Then, it uses the second training audio annotated with frame-level audio categories (voice / non-voice), supplemented by the first voice activity detection model, to train the second voice activity detection model that can capture the semantic information in the voice. The trained second voice activity detection model is used as the target voice activity detection model to perform voice activity detection on the audio to be detected. The target voice activity detection model trained by the voice activity detection model training method provided by the embodiments of this application can capture the semantic information of the voice in the input audio, and then be able to give a reasonable category for each audio frame of the input audio with reference to the semantic information, so that the voice segments with complete semantics and short pauses in the middle can be segmented into the same voice segment as much as possible, thereby improving the effect of subsequent processing (such as speech recognition).

[0126] In another embodiment of this application, the specific implementation process of "step S401: Use the first training audio annotated with frame-level audio categories to train the first voice activity detection model with the function of judging semantic integrity" in the above embodiment is introduced.

[0127] As mentioned in the above embodiment, in order to obtain a first voice activity detection model with better performance, the first training audio annotated with frame-level audio categories and the corresponding text can be used. In the voice activity detection task combined with semantic information, the first voice activity detection model with the function of judging semantic integrity can be trained by jointly performing the speech recognition task. Next, this process will be introduced.

[0128] Please refer to Figure 5, which shows a schematic flow chart of training a first voice activity detection model with semantic integrity discrimination function by using the first training audio labeled with frame-level audio categories and corresponding texts, in the voice activity detection task combined with semantic information and simultaneously jointly with the speech recognition task, may include:

[0129] Step S501: Extract audio features from the first training audio to obtain first audio features.

[0130] In a possible implementation, the process of extracting audio features from the first training audio may include: obtaining the waveform diagram corresponding to the first training audio; adopting a strategy of frame segmentation and windowing to obtain the Fbank features corresponding to the waveform diagram as the audio features corresponding to the first training audio, that is, the first audio features.

[0131] Step S502: Use the first voice activity detection model to encode the first audio features to obtain first audio encoded features.

[0132] As Figure 6 shown, the first voice activity detection model may include an encoder (such as a streaming encoder) F E , and the first audio features can be input into the encoder F E of the first voice activity detection model for encoding to obtain the first audio encoded features A E output by the encoder F E . If the first audio features are represented as X, then the first audio encoded features A E can be represented as:

[0133] A E = F E (X) (1)

[0134] The first voice activity detection model in this embodiment may be, but is not limited to, a voice activity detection model based on the Conformer structure.

[0135] Step S503: Use the first voice activity detection model to obtain the word-level speech text features corresponding to the first training audio according to the first audio encoded features.

[0136] In a possible implementation, the process of using the first voice activity detection model to obtain the word-level speech text features corresponding to the first training audio according to the first audio encoded features may include:

[0137] Step S503a: Use the connectionist temporal classification CTC module of the first voice activity detection model to obtain the word-level first speech text features corresponding to the first training audio according to the first audio encoded features.

[0138] As Figure 6As shown, the first voice activity detection model may include, in addition to the encoder F E also include the CTC module F C The first audio encoding feature A E can be input into the CTC module F C for processing. The CTC module F C outputs the word-level first voice text feature A corresponding to the first training audio C . The word-level first voice text feature A corresponding to the first training audio C can be expressed as:

[0139] A C = F C (A E ) (2)

[0140] Step S503b: Using the autoregressive decoder of the first voice activity detection model, according to the first audio encoding feature, obtain the word-level second voice text feature corresponding to the first training audio.

[0141] As Figure 6 shown, the first voice activity detection model may also include the autoregressive decoder F D . The first audio encoding feature A E can be input into the autoregressive decoder F of the first voice activity detection model D for processing. The autoregressive decoder F D outputs the word-level second voice text feature A corresponding to the first training audio D . The word-level second voice text feature A corresponding to the first training audio D can be expressed as:

[0142] A D = F D (A E ) (3)

[0143] Step S504: Using the first voice activity detection model, process the first audio encoding feature and the word-level voice text feature corresponding to the first training audio into a feature containing the semantic information of the first training audio to obtain the target feature.

[0144] As Figure 6 shown, the first voice activity detection model may also include a feature processing module (such as a long history Conformer module) F V , which can splice the first audio encoding feature A E with the word-level first voice text feature A corresponding to the first training audio C , input the spliced feature into the feature processing module F of the first voice activity detection model V , and the feature processing module F VObtain the features containing the semantic information of the first training audio by processing the spliced features as the target feature A V The target feature A V can be expressed as:

[0145] A V =F V (concat(A E ,A C )) (4)

[0146] Step S505: Update the parameters of the first voice activity detection model with the goal of making the frame-level audio category predicted by the first voice activity detection model according to the target feature tend to be consistent with the frame-level audio category annotated in the first training audio, and making the text predicted by the first voice activity detection model according to the word-level speech text feature corresponding to the first training audio tend to be consistent with the text annotated in the first training audio

[0147] In a possible implementation, the process of updating the parameters of the first voice activity detection model with the goal of making the frame-level audio category predicted by the first voice activity detection model according to the target feature tend to be consistent with the frame-level audio category annotated in the first training audio, and making the text predicted by the first voice activity detection model according to the word-level speech text feature corresponding to the first training audio tend to be consistent with the text annotated in the first training audio may include:

[0148] Step S5051a: Use the first voice activity detection model to obtain the frame-level category prediction probability of the first training audio according to the target feature

[0149] Among them, the category prediction probability of each audio frame of the first training audio includes the prediction probabilities corresponding to the three categories of non-speech at the incomplete speech and semantics and non-speech after the complete speech and semantics of the audio frame

[0150] Use the first voice activity detection model, according to the target feature A V to predict the prediction probabilities corresponding to the three categories of non-speech at the incomplete speech and semantics and non-speech after the complete speech and semantics of each audio frame of the first training audio, and obtain the frame-level category prediction probability of the first training audio

[0151] As Figure 6 shown, the first voice activity detection model can also have a first classification module, which can input the target feature A V into the first classification module of the first voice activity detection model. The first classification module is based on the target feature A VPredict the predicted probabilities corresponding to each audio frame of the first training audio in the three categories of non-speech at the incomplete speech and semantic parts and non-speech after the complete semantic speech, and output the frame-level category prediction probabilities of the first training audio.

[0152] Step S5051b: Use the first voice activity detection model to obtain the first text prediction probability according to the character-level first voice text features corresponding to the first training audio.

[0153] Use the first voice activity detection model, according to the character-level first voice text feature A corresponding to the first training audio C , predict the text corresponding to the first training audio, and obtain the first text prediction probability.

[0154] As Figure 6 shown, the first voice activity detection model can also have a second classification module, which can input the character-level first voice text feature A corresponding to the first training audio C into the second classification module of the first voice activity detection model. The second classification module predicts the text corresponding to the first training audio according to the character-level first voice text feature A C , and outputs the first text prediction probability.

[0155] Step S5051c: Use the first voice activity detection model to obtain the second text prediction probability according to the character-level second voice text features corresponding to the first training audio.

[0156] Use the first voice activity detection model, according to the character-level second voice text feature A corresponding to the first training audio D , predict the text corresponding to the first training audio, and obtain the second text prediction probability.

[0157] As Figure 6 shown, the first voice activity detection model can also have a third classification module, which can input the character-level second voice text feature A corresponding to the first training audio D into the third classification module of the first voice activity detection model. The third classification module predicts the text corresponding to the first training audio according to the character-level second voice text feature A D , and outputs the second text prediction probability.

[0158] Step S5052a: Determine the first category prediction loss according to the frame-level category prediction probabilities of the first training audio and the frame-level audio categories annotated for the first training audio.

[0159] The first category prediction loss LOSS_3 can be calculated by the following formula:

[0160] (5)

[0161] Among them, P V represents the frame-level class prediction probability of the first training audio, and Target3 represents the frame-level audio class annotated for the first training audio.

[0162] Step S5052b: Determine the first text prediction loss according to the first text prediction probability and the text annotated for the first training audio.

[0163] The first text prediction loss CTC_LOSS can be calculated by the following formula:

[0164] (6)

[0165] Among them, P C represents the first text prediction probability of the first training audio, and Target w represents the text annotated for the first training audio.

[0166] Step S5052c: Determine the second text prediction loss according to the second text prediction probability and the text annotated for the first training audio.

[0167] The second text prediction loss ED_LOSS can be calculated by the following formula:

[0168] (7)

[0169] Among them, P D represents the second text prediction probability of the first training audio, and Target w represents the text annotated for the first training audio.

[0170] Step S5053: Update the parameters of the first voice activity detection model according to the first class prediction loss, and simultaneously combining the first text prediction loss and the second text prediction loss.

[0171] After calculating the first class prediction loss LOSS_3, the first text prediction loss CTC_LOSS, and the second text prediction loss ED_LOSS, update the parameters of the first voice activity detection model according to the three losses.

[0172] Use the training data in the first training dataset (the first training dataset includes multiple pieces of training data, and each piece of training data is the first training audio annotated with the frame-level audio class and the corresponding text), and perform multiple iterative trainings on the first voice activity detection model according to the above process until the training end condition is met (such as the model converges, reaches the preset number of training iterations, etc.).

[0173] In another embodiment of the present application, the specific implementation process of "step S402: using the second training audio labeled with frame-level audio categories, supplemented by the first voice activity detection model, to train the second voice activity detection model capable of capturing semantic information in speech" in the above embodiment is introduced.

[0174] Please refer to Figure 7 , which shows a schematic flowchart of training the second voice activity detection model capable of capturing semantic information in speech by using the second training audio labeled with frame-level audio categories and supplemented by the first voice activity detection model, and may include:

[0175] Step S701: Extract audio features from the second training audio to obtain the second audio features.

[0176] Specifically, the process of extracting audio features from the second training audio may include: obtaining the waveform diagram corresponding to the second training audio; adopting a frame segmentation and windowing strategy to obtain the Fbank features corresponding to the waveform diagram as the audio features corresponding to the second training audio, that is, the second audio features.

[0177] Step S702: Encode the second audio features by using the first voice activity detection model to obtain the second audio encoded features.

[0178] As mentioned in the above embodiment, the first voice activity detection model may include an encoder, and the second audio features can be input into the encoder of the first voice activity detection model for encoding to obtain the second audio encoded features.

[0179] Step S703: Select one feature from the second audio features and the second audio encoded features according to a preset feature selection strategy.

[0180] In a possible implementation manner, the probability of selecting the second audio encoded features can be preset, and then, features are selected from the second audio features and the second audio encoded features according to the preset probability.

[0181] Step S704: Use the selected features and the frame-level audio categories labeled by the second training audio, supplemented by the first voice activity detection model, to train the second voice activity detection model.

[0182] It should be noted that selecting the second audio encoded features to train the second voice activity detection model can enable the second voice activity detection model to learn the semantic information in speech.

[0183] Specifically, the process of using the selected features and the frame-level audio categories labeled by the second training audio, supplemented by the first voice activity detection model, to train the second voice activity detection model may include:

[0184] Step S7041a: Using the first voice activity detection model, obtain the frame-level first-category prediction probabilities of the second training audio according to the second audio coding features.

[0185] Among them, the first-category prediction probabilities of each audio frame of the second training audio include the prediction probabilities corresponding to the audio frame in the three categories of non-speech at the speech and semantically incomplete parts and non-speech after semantically complete speech.

[0186] Specifically, input the second audio coding features into the CTC module F of the first voice activity detection model. C Then, splice the second audio coding features with the features output by the CTC module F. C After that, input the spliced features into the feature processing module (such as the long history Conformer module) F of the first voice activity detection model. V Next, input the features output by the feature processing module F. V into the first classification module of the first voice activity detection model to obtain the frame-level first-category prediction probabilities of the second training audio output by the first classification module.

[0187] Step S7041b: Using the second voice activity detection model, obtain the frame-level second-category prediction probabilities of the second training audio according to the selected features.

[0188] Among them, the second-category prediction probabilities of each audio frame of the second training audio include the prediction probabilities corresponding to the audio frame in the two categories of speech and non-speech.

[0189] In a possible implementation, as Figure 8 shown, the second voice activity detection model includes a feature processing module and a classification module. The selected features can be input into the feature processing module of the second voice activity detection model, and the features output by the feature processing module are input into the classification module. The classification module outputs the prediction probabilities corresponding to the audio frame of the second training audio in the two categories of speech and non-speech respectively, that is, the frame-level second-category prediction probabilities of the second training audio.

[0190] In a possible implementation, the feature processing module of the second voice activity detection model may sequentially include a first convolutional module, an LSTM module, and a second convolutional module.

[0191] Step S7042: Adjust the frame-level second-category prediction probabilities of the second training audio according to the frame-level first-category prediction probabilities of the second training audio to obtain the adjusted frame-level second-category prediction probabilities of the second training audio.

[0192] Specifically, the process of adjusting the frame-level second-category prediction probabilities of the second training audio according to the frame-level first-category prediction probabilities of the second training audio may include:

[0193] Step a1: Determine the first audio category of each audio frame of the second training audio according to the frame-level first category prediction probability of the second training audio, and determine the second audio category of each audio frame of the second training audio frame according to the frame-level second category prediction probability of the second training audio.

[0194] The frame-level first category prediction probability of the second training audio includes the prediction probabilities corresponding to each of the three categories of non-speech at the speech and semantic incomplete parts and non-speech after the semantic complete speech of each audio frame of the second training audio. For each audio frame of the second training audio, the category corresponding to the maximum probability in the first category prediction probability of the audio frame can be determined as the first audio category of the audio frame.

[0195] The frame-level second category prediction probability of the second training audio includes the prediction probabilities corresponding to each of the two categories of speech and non-speech of each audio frame of the second training audio. For each audio frame of the second training audio, the category corresponding to the maximum probability in the second category prediction probability of the audio frame can be determined as the second audio category of the audio frame.

[0196] Step a2: For each audio frame of the second training audio, adjust the second category prediction probability of the audio frame according to the first audio category and the second audio category of the audio frame.

[0197] Specifically, the process of adjusting the second category prediction probability of the audio frame according to the first audio category and the second audio category of the audio frame may include: if both the first audio category and the second audio category of the audio frame are speech, increase the second category prediction probability corresponding to the speech of the audio frame; if the first audio category of the audio frame is non-speech after the semantic complete speech and the second audio category of the audio frame is non-speech, increase the second category prediction probability corresponding to the non-speech of the audio frame; if the first audio category of the audio frame is non-speech at the semantic incomplete part and the second audio category of the audio frame is non-speech, decrease the second category prediction probability corresponding to the non-speech of the audio frame.

[0198] In a possible implementation, if both the first audio category and the second audio category of the audio frame are speech, the predicted probability of the second category corresponding to the audio frame in terms of speech can be increased by multiplying it by a preset first coefficient (a relatively large coefficient, such as 1.5). If the first audio category of the audio frame is non-speech after semantically complete speech and the second audio category of the audio frame is non-speech, the predicted probability of the second category corresponding to the audio frame in terms of non-speech can be increased by multiplying it by a preset first coefficient (a relatively large coefficient). If the first audio category of the audio frame is non-speech at a semantically incomplete part and the second audio category of the audio frame is non-speech, the predicted probability of the second category corresponding to the audio frame in terms of non-speech can be decreased by multiplying it by a preset second coefficient (a relatively small coefficient, such as 0.5).

[0199] Among them, the first coefficient can take values within the interval [1.5, 2], and the second coefficient can take values within the interval [0.5, 0.7]. It should be noted that the above intervals are only examples, and the value ranges and specific values of the first coefficient and the second coefficient can be set according to the actual application scenario.

[0200] In addition, it should be noted that the embodiments of the present application do not limit "increasing the predicted probability of the second category corresponding to the audio frame in terms of speech by multiplying the predicted probability of the second category corresponding to the audio frame in terms of speech by a preset first coefficient", and other methods can also be used to increase the predicted probability of the second category corresponding to the audio frame in terms of speech. For example, a set value can be added to the predicted probability of the second category corresponding to the audio frame in terms of speech. This embodiment also does not limit "increasing the predicted probability of the second category corresponding to the audio frame in terms of non-speech by multiplying the predicted probability of the second category corresponding to the audio frame in terms of non-speech by a preset first coefficient (multiplying by a relatively large coefficient)", for example, a set value can be added to the predicted probability of the second category corresponding to the audio frame in terms of non-speech. Similarly, this embodiment does not limit "decreasing the predicted probability of the second category corresponding to the audio frame in terms of non-speech by multiplying the predicted probability of the second category corresponding to the audio frame in terms of non-speech by a preset second coefficient (multiplying by a relatively small coefficient)", for example, a set value can be subtracted from the predicted probability of the second category corresponding to the audio frame in terms of non-speech.

[0201] Step S7043: Determine the second category prediction loss according to the adjusted frame-level second category prediction probability of the second training audio and the frame-level audio category annotated for the second training audio.

[0202] If the adjusted frame-level second category prediction probability of the second training audio is represented as P' VAD, represent the frame-level audio categories of the second training audio annotations as Target2, then the calculation formula of the second category prediction loss LOSS_VAD is as follows:

[0203] (8)

[0204] Step S7044: Update the parameters of the second voice activity detection model according to the second category prediction loss.

[0205] Use the training data in the second training dataset (the second training dataset includes multiple second training audios annotated with frame-level audio categories), and perform multiple iterative trainings on the second voice activity detection model according to the above process until the training end condition is met (such as model convergence, reaching the preset number of training iterations, etc.).

[0206] It should be noted that in the above step S703, features will be selected from the second audio features and the second audio encoding features according to a preset feature selection strategy. In one possible implementation, in the early stage of training (such as the first 50,000 steps), the second audio encoding features are selected with a probability of P1 (such as 0.8), in the middle stage of training (such as 50,000 steps - 100,000 steps), the second audio encoding features are selected with a probability of P2 (such as 0.5), and in the later stage of training (such as after 100,000 steps), the second audio encoding features are selected with a probability of P3 (such as 0.2), forcing the second voice activity detection model to still be able to directly perform speech and non-speech predictions based on the differences between speech and non-speech after learning semantic information. Among them, P1 > P2 > P3, and the specific values of P1, P2, and P3 can be set according to the actual application situation.

[0207] Based on the voice activity detection model training method provided in the above embodiments, the embodiments of the present application also provide a voice activity detection method, which includes:

[0208] Obtain the target audio; use the target voice activity detection model trained by the voice activity detection model training method provided in the above embodiments to perform voice activity detection on the target audio.

[0209] Through the voice activity detection model provided by the embodiments of the present application, reasonable voice activity detection results can be obtained.

[0210] The embodiments of the present application also provide a device corresponding to the voice activity detection model training method provided in the above embodiments. Please refer to Figure 9 , Figure 9 is a schematic structural diagram of a voice activity detection model training device provided by an embodiment of the present application. The voice activity detection model training device may include: a first training module 901 and a second training module 902.

[0211] The first training module 901 is configured to use first training audio labeled with frame-level audio categories to train a first voice activity detection model with a semantic integrity discrimination function. Among them, the audio category of any audio frame of the first training audio is one of speech, non-speech at the non-semantic-integrity part, and non-speech after semantically complete speech.

[0212] The second training module 902 is configured to use second training audio labeled with frame-level audio categories, supplemented by the first voice activity detection model, to train a second voice activity detection model capable of capturing semantic information in speech. The trained second voice activity detection model is used as the target voice activity detection model. Among them, the audio category of any audio frame of the second training audio is one of speech and non-speech.

[0213] In a possible implementation manner, when the first training module 901 uses the first training audio labeled with frame-level audio categories to train a first voice activity detection model with a semantic integrity discrimination function, it is specifically configured to:

[0214] Use the first training audio labeled with frame-level audio categories and corresponding texts, and simultaneously combine the speech recognition task in the speech activity detection task that combines semantic information to train a first voice activity detection model with a semantic integrity discrimination function.

[0215] In a possible implementation manner, when the first training module 901 uses the first training audio labeled with frame-level audio categories and corresponding texts, and simultaneously combines the speech recognition task in the speech activity detection task that combines semantic information to train a first voice activity detection model with a semantic integrity discrimination function, it is specifically configured to:

[0216] Extract audio features from the first training audio to obtain first audio features;

[0217] Use the first voice activity detection model to encode the first audio features to obtain first audio encoded features;

[0218] Use the first voice activity detection model to obtain word-level speech text features corresponding to the first training audio according to the first audio encoded features;

[0219] Use the first voice activity detection model to process the first audio encoded features and the word-level speech text features corresponding to the first training audio into features containing the semantic information of the first training audio to obtain target features;

[0220] The parameter update of the first voice activity detection model is targeted at making the frame-level audio categories predicted by the first voice activity detection model according to the target features tend to be consistent with the frame-level audio categories annotated in the first training audio, and making the text predicted by the first voice activity detection model according to the word-level speech text features corresponding to the first training audio tend to be consistent with the text annotated in the first training audio.

[0221] In a possible implementation, when the first training module 901 uses the first voice activity detection model to obtain the word-level speech text features corresponding to the first training audio according to the first audio coding features, it is specifically used for:

[0222] Using the connectionist temporal classification (CTC) module of the first voice activity detection model to obtain the word-level first speech text features corresponding to the first training audio according to the first audio coding features;

[0223] And / or, using the autoregressive decoder of the first voice activity detection model to obtain the word-level second speech text features corresponding to the first training audio according to the first audio coding features.

[0224] In a possible implementation, when the first training module 901 updates the parameters of the first voice activity detection model with the goal of making the frame-level audio categories predicted by the first voice activity detection model according to the target features tend to be consistent with the frame-level audio categories annotated in the first training audio, and making the text predicted by the first voice activity detection model according to the word-level speech text features corresponding to the first training audio tend to be consistent with the text annotated in the first training audio, it is specifically used for:

[0225] Using the first voice activity detection model to obtain the frame-level category prediction probabilities of the first training audio according to the target features, where the category prediction probabilities of each audio frame of the first training audio include the prediction probabilities corresponding to the three categories of non-speech at the speech and semantic incomplete parts and non-speech after the semantic complete speech of the audio frame;

[0226] Using the first voice activity detection model to obtain the first text prediction probability according to the word-level first speech text features; and / or, using the first voice activity detection model to obtain the second text prediction probability according to the word-level second speech text features;

[0227] Determining the first category prediction loss according to the frame-level category prediction probabilities of the first training audio and the frame-level audio categories annotated in the first training audio;

[0228] Determining the first text prediction loss according to the first text prediction probability and the text annotated in the first training audio; and / or, determining the second text prediction loss according to the second text prediction probability and the text annotated in the first training audio;

[0229] Predict the loss according to the first category, and update the parameters of the first voice activity detection model by combining the first text prediction loss and / or the second text prediction loss.

[0230] In a possible implementation, when the second training module 902 uses the second training audio labeled with frame-level audio categories and supplements it with the first voice activity detection model to train the second voice activity detection model capable of capturing semantic information in speech, it is specifically used for:

[0231] Extract audio features from the second training audio to obtain second audio features;

[0232] Use the first voice activity detection model to encode the second audio features to obtain second audio encoded features;

[0233] Select a feature from the second audio features and the second audio encoded features according to a preset feature selection strategy;

[0234] Use the selected feature and the frame-level audio categories labeled by the second training audio, supplemented by the first voice activity detection model, to train the second voice activity detection model.

[0235] In a possible implementation, when the second training module 902 uses the selected feature and the frame-level audio categories labeled by the second training audio, supplemented by the first voice activity detection model, to train the second voice activity detection model, it is specifically used for:

[0236] Use the first voice activity detection model to obtain the frame-level first category prediction probabilities of the second training audio according to the second audio encoded features, where the first category prediction probabilities of each audio frame of the second training audio include the prediction probabilities corresponding to the three categories of non-speech at the speech and semantically incomplete parts and non-speech after semantically complete speech of the audio frame;

[0237] Use the second voice activity detection model to obtain the frame-level second category prediction probabilities of the second training audio according to the selected feature, where the second category prediction probabilities of each audio frame of the second training audio include the prediction probabilities corresponding to the two categories of speech and non-speech of the audio frame;

[0238] Adjust the frame-level second category prediction probabilities of the second training audio according to the frame-level first category prediction probabilities of the second training audio to obtain the adjusted frame-level second category prediction probabilities of the second training audio;

[0239] Determine the second category prediction loss according to the adjusted frame-level second category prediction probabilities of the second training audio and the frame-level audio categories labeled by the second training audio;

[0240] Update the parameters of the second voice activity detection model according to the predicted loss of the second category.

[0241] In a possible implementation, when the second training module 902 adjusts the frame-level second-category prediction probability of the second training audio according to the frame-level first-category prediction probability of the second training audio, it is specifically used for:

[0242] Determine the first audio category of each audio frame of the second training audio according to the frame-level first-category prediction probability of the second training audio, and determine the second audio category of each audio frame of the second training audio according to the frame-level second-category prediction probability of the second training audio;

[0243] For each audio frame of the second training audio, adjust the second-category prediction probability of the audio frame according to the first audio category and the second audio category of the audio frame.

[0244] In a possible implementation, when the second training module 902 adjusts the second-category prediction probability of the audio frame according to the first audio category and the second audio category of the audio frame, it is specifically used for:

[0245] If both the first audio category and the second audio category of the audio frame are speech, increase the second-category prediction probability corresponding to the audio frame in speech;

[0246] If the first audio category of the audio frame is non-speech after semantically complete speech and the second audio category of the audio frame is non-speech, increase the second-category prediction probability corresponding to the audio frame in non-speech;

[0247] If the first audio category of the audio frame is non-speech at the semantically incomplete part and the second audio category of the audio frame is non-speech, decrease the second-category prediction probability corresponding to the audio frame in non-speech.

[0248] In a possible implementation, when the second training module 902 increases the second-category prediction probability corresponding to the audio frame in speech, it is specifically used for:

[0249] Increase the second-category prediction probability corresponding to the audio frame in speech by multiplying the second-category prediction probability corresponding to the audio frame in speech by a preset first coefficient;

[0250] When the second training module 902 increases the second-category prediction probability corresponding to the audio frame in non-speech, it is specifically used for:

[0251] Increase the second-category prediction probability corresponding to the audio frame in non-speech by multiplying the second-category prediction probability corresponding to the audio frame in non-speech by a preset first coefficient;

[0252] When the second training module 902 reduces the predicted probability of the second category corresponding to the audio frame on non-speech, it is specifically configured to:

[0253] Reduce the predicted probability of the second category corresponding to the audio frame on non-speech by multiplying the predicted probability of the second category corresponding to the audio frame on non-speech by a preset second coefficient.

[0254] The target voice activity detection model trained by the voice activity detection model training device provided by the embodiments of the present application can capture the semantic information of the voice in the input audio, and then can give a reasonable category for each audio frame of the input audio with reference to the semantic information, so that the voice segments with complete semantics and short intermediate pauses can be segmented in the same voice segment as much as possible, thereby improving the effect of subsequent processing (such as speech recognition).

[0255] The embodiments of the present application also provide a voice activity detection device, as Figure 10 shown, the voice activity detection device may include an audio acquisition module 1001 and a voice activity detection module 1002.

[0256] The audio acquisition module 1001 is configured to acquire a target audio. Wherein, the target audio is an audio to be detected.

[0257] The voice activity detection module 1002 is configured to perform voice activity detection on the target audio by using the target voice activity detection model trained by the voice activity detection model training device provided by the above embodiments.

[0258] Through the voice activity detection device provided by the embodiments of the present application, reasonable voice activity detection results can be obtained.

[0259] The embodiments of the present application also provide an electronic device, which may include: at least one processor, at least one communication interface, at least one memory, and at least one communication bus.

[0260] In the embodiments of the present application, the number of the processor, the communication interface, the memory, and the communication bus is at least one, and the processor, the communication interface, and the memory complete communication with each other through the communication bus;

[0261] The processor may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application, etc.;

[0262] The memory may include high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory;

[0263] Among them, the memory stores a program, and the processor can call the program stored in the memory. The program is used to implement the steps of the voice activity detection model training method provided in the above embodiments, and / or implement the steps of the voice activity detection method provided in the above embodiments.

[0264] The embodiments of the present application also provide a computer storage medium. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can be enabled to implement the steps of the voice activity detection model training method provided in the above embodiments, and / or implement the steps of the voice activity detection method provided in the above embodiments.

[0265] The embodiments of the present application also provide a computer program product, including computer-readable instructions. When the computer-readable instructions run on an electronic device, the electronic device is enabled to implement the steps of the voice activity detection model training method provided in the above embodiments, and / or implement the steps of the voice activity detection method provided in the above embodiments.

[0266] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided in the present application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.

[0267] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits, etc. However, for this application, software program implementation is a better embodiment in more cases. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, etc., and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.

[0268] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0269] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training device, or data center to another website, computer, training device, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

Claims

1. A method for training a voice activity detection model, characterized in that: include: Using a first training audio marked with a frame-level audio category, a first voice activity detection model with a semantic completeness discrimination function is trained, wherein the audio category of any audio frame of the first training audio is one of speech, non-speech at a semantically incomplete location, and non-speech after a semantically complete speech; Using a second training audio labeled with frame-level audio categories and assisted by the first voice activity detection model, a second voice activity detection model capable of capturing semantic information in speech is trained, and the trained second voice activity detection model is used as a target voice activity detection model, wherein the audio category of any audio frame of the second training audio is one of speech and non-speech.

2. The method for training a voice activity detection model according to claim 1, characterized in that: The method of using the first training audio marked with the frame-level audio category to train a first voice activity detection model with a semantic integrity discrimination function includes: By using the first training audio labeled with frame-level audio categories and corresponding texts, a first voice activity detection model with semantic integrity discrimination function is trained on a voice activity detection task combining semantic information and a speech recognition task.

3. The method for training a voice activity detection model according to claim 2, characterized in that: The first training audio with the frame-level audio category and the corresponding text annotated therewith is trained to obtain a first voice activity detection model with the function of distinguishing semantic integrity in the voice activity detection task combined with semantic information and the voice recognition task, including: Extracting audio features from the first training audio to obtain first audio features; Encoding the first audio feature using a first voice activity detection model to obtain a first audio encoding feature; Using a first voice activity detection model, and based on the first audio coding feature, obtaining a word-level speech text feature corresponding to the first training audio; Using a first voice activity detection model, processing the first audio coding feature and the word-level speech text feature corresponding to the first training audio into a feature containing semantic information of the first training audio to obtain a target feature; The parameters of the first voice activity detection model are updated with the goal of making the frame-level audio category predicted by the first voice activity detection model based on the target feature consistent with the frame-level audio category annotated by the first training audio, and making the text predicted by the first voice activity detection model based on the word-level speech text feature corresponding to the first training audio consistent with the text annotated by the first training audio.

4. The method for training a voice activity detection model according to claim 3, characterized in that: The step of using the first voice activity detection model to obtain the word-level speech text feature corresponding to the first training audio according to the first audio coding feature includes: Using a connection temporal classification (CTC) module of a first voice activity detection model, obtaining a word-level first speech text feature corresponding to the first training audio according to the first audio coding feature; And / or, using an autoregressive decoder of a first voice activity detection model, obtaining a word-level second speech text feature corresponding to the first training audio according to the first audio coding feature.

5. The method for training a voice activity detection model according to claim 4, characterized in that: The updating of parameters of the first voice activity detection model with the goal of making the frame-level audio category predicted by the first voice activity detection model according to the target feature consistent with the frame-level audio category annotated by the first training audio, and making the text predicted by the first voice activity detection model according to the word-level speech text feature corresponding to the first training audio consistent with the text annotated by the first training audio, comprises: Using the first voice activity detection model, according to the target feature, obtaining a frame-level category prediction probability of the first training audio, wherein the category prediction probability of each audio frame of the first training audio includes prediction probabilities corresponding to three categories of the audio frame: speech, non-speech at a semantically incomplete location, and non-speech after a semantically complete speech; Using the first voice activity detection model, according to the word-level first voice-text feature, obtaining a first text prediction probability; and / or using the first voice activity detection model, according to the word-level second voice-text feature, obtaining a second text prediction probability; Determining a first category prediction loss according to the frame-level category prediction probability of the first training audio and the frame-level audio category annotated by the first training audio; Determine a first text prediction loss based on the first text prediction probability and the text annotated by the first training audio; and / or determine a second text prediction loss based on the second text prediction probability and the text annotated by the first training audio; According to the first category prediction loss and in combination with the first text prediction loss and / or the second text prediction loss, parameters of the first voice activity detection model are updated.

6. The method for training a voice activity detection model according to claim 1, wherein: The method of using the second training audio marked with the frame-level audio category and the first voice activity detection model to train a second voice activity detection model capable of capturing semantic information in the voice includes: Extracting audio features from the second training audio to obtain second audio features; Encoding the second audio feature using the first voice activity detection model to obtain a second audio encoding feature; Selecting a feature from the second audio feature and the second audio coding feature according to a preset feature selection strategy; The second voice activity detection model is trained using the selected features and the frame-level audio categories annotated by the second training audio, assisted by the first voice activity detection model.

7. The method for training a voice activity detection model according to claim 6, wherein: The step of training the second voice activity detection model by using the selected features and the frame-level audio category annotated by the second training audio, supplemented by the first voice activity detection model, includes: Using the first voice activity detection model, according to the second audio coding feature, obtaining a frame-level first category prediction probability of the second training audio, wherein the first category prediction probability of each audio frame of the second training audio includes prediction probabilities corresponding to three categories of the audio frame: speech, non-speech at a semantically incomplete location, and non-speech after a semantically complete speech; Using the second voice activity detection model, according to the selected features, obtaining a frame-level second category prediction probability of the second training audio, wherein the second category prediction probability of each audio frame of the second training audio includes prediction probabilities corresponding to the audio frame in two categories, speech and non-speech; According to the frame-level first category prediction probability of the second training audio, adjusting the frame-level second category prediction probability of the second training audio to obtain an adjusted frame-level second category prediction probability of the second training audio; Determine a second category prediction loss according to the adjusted frame-level second category prediction probability of the second training audio and the frame-level audio category annotated by the second training audio; According to the second category prediction loss, parameters of the second voice activity detection model are updated.

8. The method for training a voice activity detection model according to claim 7, characterized in that: The adjusting the frame-level second-category prediction probability of the second training audio according to the frame-level first-category prediction probability of the second training audio includes: Determine a first audio category for each audio frame of the second training audio according to the frame-level first category prediction probability of the second training audio, and determine a second audio category for each audio frame of the second training audio according to the frame-level second category prediction probability of the second training audio; For each audio frame of the second training audio, the second category prediction probability of the audio frame is adjusted according to the first audio category of the audio frame and the second audio category of the audio frame.

9. The method for training a voice activity detection model according to claim 8, characterized in that: The adjusting the second category prediction probability of the audio frame according to the first audio category of the audio frame and the second audio category of the audio frame includes: If the first audio category of the audio frame and the second audio category of the audio frame are both speech, increasing the prediction probability of the second category corresponding to the speech of the audio frame; If the first audio category of the audio frame is non-speech after semantically complete speech, and the second audio category of the audio frame is non-speech, then the prediction probability of the second category corresponding to the non-speech of the audio frame is increased; If the first audio category of the audio frame is non-speech at a semantically incomplete location, and the second audio category of the audio frame is non-speech, then the prediction probability of the second category corresponding to the non-speech of the audio frame is reduced.

10. The method for training a voice activity detection model according to claim 9, wherein: The step of increasing the prediction probability of the second category corresponding to the audio frame in speech comprises: Increasing the second category prediction probability corresponding to the audio frame in speech by multiplying the second category prediction probability corresponding to the audio frame in speech by a preset first coefficient; The step of increasing the prediction probability of the second category corresponding to the audio frame in non-speech mode comprises: Increasing the second category prediction probability corresponding to the audio frame in non-speech by multiplying the second category prediction probability corresponding to the audio frame in non-speech by a preset first coefficient; The reducing the prediction probability of the second category corresponding to the audio frame in non-speech includes: The second category prediction probability corresponding to the audio frame in non-speech is reduced by multiplying the second category prediction probability corresponding to the audio frame in non-speech by a preset second coefficient.

11. A voice activity detection method, characterized in that: include: Get the target audio; Using a target voice activity detection model trained by the voice activity detection model training method according to any one of claims 1 to 10, voice activity detection is performed on the target audio.

12. A voice activity detection model training device, characterized in that: include: a first training module and a second training module; The first training module is used to train a first voice activity detection model with a semantic integrity discrimination function using a first training audio marked with a frame-level audio category, wherein the audio category of any audio frame of the first training audio is one of speech, non-speech at a semantically incomplete location, and non-speech after a semantically complete speech; The second training module is used to use the second training audio marked with frame-level audio categories, assisted by the first voice activity detection model, to train a second voice activity detection model that can capture semantic information in the voice, and the trained second voice activity detection model is used as the target voice activity detection model, wherein the audio category of any audio frame of the second training audio is one of speech and non-speech.

13. A voice activity detection device, characterized in that: include: Audio acquisition module and voice activity detection module; The audio acquisition module is used to acquire target audio; The voice activity detection module is used to perform voice activity detection on the target audio using a target voice activity detection model trained by the voice activity detection model training device according to claim 12.

14. An electronic device, characterized in that: The method comprises at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the electronic device can implement the steps of the voice activity detection model training method as described in any one of claims 1 to 10, and / or implement the steps of the voice activity detection method as described in claim 11.

15. A computer storage medium, characterized in that: The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement the steps of the voice activity detection model training method as described in any one of claims 1 to 10, and / or implement the steps of the voice activity detection method as described in claim 11.

16. A computer program product, characterized in that The method comprises computer-readable instructions, which, when executed on an electronic device, enable the electronic device to implement the steps of the voice activity detection model training method as described in any one of claims 1 to 10, and / or implement the steps of the voice activity detection method as described in claim 11.