Model training method, voice detection method, device, equipment, medium and product
By constructing a sample set containing reverberation and noise samples, the speech detection model is trained, which solves the problem of insufficient accuracy in speech detection under complex noise environments and achieves more efficient speech recognition results.
Patent Information
- Application Number
- CN202410573134.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-09
- Publication Date
- 2025-11-11
AI Technical Summary
Existing technologies lack accuracy in speech detection in complex noisy environments, making it difficult to effectively distinguish between valid and invalid speech, especially in reverberant and noisy environments.
By constructing a sample set containing reverberation and noise samples, and iteratively training the initial speech detection model, a speech detection model capable of recognizing both valid and invalid speech is generated. This includes reverberation processing and noise simulation of the samples, as well as room impulse response library and signal-to-noise ratio adjustment.
It improves the accuracy and practicality of speech detection, enabling more accurate identification of valid human voices and background noise in complex noisy environments, thus enhancing the recognition capabilities of the speech detection model.
Smart Images

Figure CN120932632A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of cloud technology, artificial intelligence and other technical fields, and to a model training method, a speech detection method, a device, an equipment, a medium and a product. Background Technology
[0002] Effective speech detection (VAD) technology detects the presence of valid speech in an audio signal by analyzing its features. In this field, speech detection is typically based on the energy of the audio signal. Specifically, this involves comparing the short-time energy of the audio signal with an energy threshold; if the short-time energy exceeds the energy threshold, the audio is considered speech; otherwise, it is considered non-speech. However, this method, relying solely on energy and thresholds, is only suitable for simple environments. Therefore, improving audio processing to handle speech detection in complex, noisy environments remains a pressing technical problem to be solved in this field. Summary of the Invention
[0003] This application provides a model training method, a speech detection method, a device, an equipment, a medium, and a product, which can improve the accuracy and practicality of speech detection. The technical solution is as follows: On the one hand, this application provides a model training method, the method comprising: A first sample set and a second sample set are obtained, and the sample labels corresponding to the first sample set and the second sample set respectively represent invalid speech and valid speech; the first sample set includes at least one first reverberation sample and a noise sample, and the second sample set includes at least one valid speech sample. The effective speech samples include effective speech data from human voice sources; the first reverberation sample is obtained by propagating the effective speech data in sample rooms with different reverberation intensities; and the noise samples include environmental noise data from non-human voice sources. Based on the audio features of each sample in the first and second sample sets, an initial speech detection model is used to perform speech detection on each sample to obtain the sample detection results of each sample. The sample detection results represent whether the corresponding sample belongs to valid speech. Based on the sample detection results and sample labels of each sample, the total training loss corresponding to each sample is determined, and the initial speech detection model is iteratively trained based on the total training loss corresponding to each sample to obtain the speech detection model.
[0004] On the other hand, this application provides a speech detection method, the method comprising: Acquire at least one frame of audio signal from the target room where the target object is located; The trained speech detection model is used to detect each frame of audio signal to obtain the detection result of each frame of audio signal. The detection result indicates whether the corresponding frame of audio signal belongs to valid speech. The speech detection model is trained using the model training method described above.
[0005] On the other hand, this application provides a model training apparatus, the apparatus comprising: The sample acquisition module is used to acquire a first sample set and a second sample set, wherein the sample labels corresponding to the first sample set and the second sample set respectively represent invalid speech and valid speech; the first sample set includes at least one first reverberation sample and a noise sample, and the second sample set includes at least one valid speech sample. The effective speech samples include effective speech data from human voice sources; the first reverberation sample is obtained by propagating the effective speech data in sample rooms with different reverberation intensities; and the noise samples include environmental noise data from non-human voice sources. The first detection module is used to perform speech detection on each sample based on the sample audio features of each sample in the first sample set and the second sample set, using an initial speech detection model to obtain the sample detection result of each sample. The sample detection result represents whether the corresponding sample belongs to valid speech. The total loss determination module is used to determine the total training loss for each sample based on the sample detection results and sample labels of each sample. The training module is used to iteratively train the initial speech detection model based on the total training loss corresponding to each sample to obtain the speech detection model.
[0006] In one possible implementation, the second sample set further includes at least one second reverberation sample; the sample acquisition module, when determining at least one first reverberation sample and the second reverberation sample, includes any one of the following: The first determining unit is configured to perform reverberation processing on valid speech data based on at least one first impulse response to obtain at least one first reverberation sample, and to perform reverberation processing on valid speech data based on at least one second impulse response to obtain at least one second reverberation sample. Wherein, the first impulse response or the second impulse response characterizes the propagation characteristics of effective speech in the corresponding sample room, and the reverberation intensity corresponding to the first impulse response is greater than that of the second impulse response; The second determining unit is configured to perform reverberation processing on at least one valid speech data from the target human voice source based on the target human voice source corresponding to each valid speech sample in the second sample set, to obtain at least one second reverberation sample, and to perform reverberation processing on at least one valid speech data from a non-target human voice source, to obtain at least one first reverberation sample.
[0007] In one possible implementation, the first determining unit is configured to: Convolution processing is performed on each first impulse response and effective speech data to generate initial reverberation data corresponding to each first impulse response; Noise data is obtained from a pre-configured noise library, and the initial reverberation data corresponding to each first impulse response and the obtained noise data are synthesized according to at least one signal-to-noise ratio to obtain the at least one first reverberation sample.
[0008] In one possible implementation, the sample acquisition module, when determining at least one first reverberation sample and a second reverberation sample, further includes at least one of the following: The third determining unit is used to determine the sample label of the first reverberation sample and the sample label of the second reverberation sample based on the reverberation intensity corresponding to the first reverberation sample and the second reverberation sample, respectively. The fourth determining unit is used to determine the sample label of the first reverberation sample and the sample label of the second reverberation sample based on the signal-to-noise ratio of the first reverberation sample and the second reverberation sample, respectively.
[0009] In one possible implementation, the third determining unit is used for at least one of the following: The first category label representing non-effective speech is used as the sample label for the first reverberation sample, and the second category label representing effective speech is used as the sample label for the second reverberation sample. Based on the reverberation intensity corresponding to the first reverberation sample and the second reverberation sample respectively, the effective speech confidence scores corresponding to the first reverberation sample and the second reverberation sample are obtained respectively, and the effective speech confidence scores corresponding to the first reverberation sample and the second reverberation sample are used as the sample labels of the first reverberation sample and the second reverberation sample respectively. The reverberation intensity is negatively correlated with the effective speech confidence score.
[0010] In one possible implementation, the second sample set further includes at least one second reverberation sample; the total loss determination module is configured to: Based on the similarity between the sample detection results and sample labels of each sample, the first training loss corresponding to each sample is determined; Based on the sample labels of each sample, at least one triplet is determined, and each triplet includes a valid speech sample, a first reverberation sample and a second reverberation sample. For each triplet, a first feature distance is determined between the audio features of the valid speech sample in the triplet and the audio features of the first reverberation sample; and a second feature distance is determined between the audio features of the valid speech sample in the triplet and the audio features of the second reverberation sample. Based on the first feature distance and the second feature distance corresponding to each triplet, the second training loss corresponding to each triplet is determined; The total training loss is obtained based on the first training loss corresponding to each sample and the second training loss corresponding to each triplet.
[0011] In one possible implementation, the sample label of the noise sample further characterizes ambient noise in the invalid speech; the sample label of the first reverberation sample further characterizes background reverberant speech in the invalid speech. The sample detection result includes at least a first result indication information, which indicates whether the sample belongs to valid speech. If the first result indication information indicates that the sample does not belong to valid speech, then the sample detection result also includes invalid speech reference information, which indicates that the sample belongs to background reverberant speech or environmental noise in invalid speech.
[0012] In one possible implementation, the sample acquisition module, when acquiring at least one valid speech sample, is configured to: Obtain at least one valid speech data from a pre-configured valid speech library, and obtain at least one noise data from a pre-configured noise library; According to at least one signal-to-noise ratio, the acquired effective speech data and noise data are synthesized to obtain at least one effective speech sample.
[0013] On the other hand, this application provides a voice detection device, the device comprising: The audio signal acquisition module is used to acquire at least one frame of audio signal from the target room where the target object is located. The second detection module is used to detect each frame of audio signal based on the trained speech detection model, and obtain the detection result of each frame of audio signal. The detection result indicates whether the corresponding frame of audio signal belongs to valid speech. The speech detection model is trained using the model training method described above.
[0014] In one possible implementation, the detection result includes audio indication information and target reference information, wherein the audio indication information indicates whether the corresponding frame audio signal belongs to valid speech, and the target reference information indicates whether the corresponding frame audio signal belongs to background reverberant speech or environmental noise in invalid speech; the device further includes any one of the following: The cut-off module is used to cut off the first audio signal corresponding to the background reverberant speech or environmental noise that belongs to non-effective speech from each frame of audio signal based on the audio indication information of each frame of audio signal; The recognition module is used to determine the second audio signal that belongs to valid speech and the third audio signal that belongs to background reverberation speech in each frame of audio signal based on the audio indication information and target reference information of each frame of audio signal respectively; and to perform voiceprint recognition on each frame of audio signal based on the audio features of the second audio signal and the third audio signal to obtain the audio signals from the target object and other objects outside the target object in each frame of audio signal respectively.
[0015] On the other hand, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the above-described model training method or speech detection method.
[0016] On the other hand, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described model training method or speech detection method.
[0017] On the other hand, this application provides a computer program product, including a computer program that, when executed by a processor, implements the above-described model training method or speech detection method.
[0018] The beneficial effects of the technical solutions provided in this application are: The model training method provided in this application obtains a first sample set of ineffective speech and a second sample set of effective speech. The first sample set includes a first reverberation sample and a noise sample. The effective speech sample includes effective speech data from human voice sources. The first reverberation sample is obtained by propagating the effective speech data in sample rooms with different reverberation intensities. The noise sample includes environmental noise data from non-human voice sources. Based on this, the speech of background human voice sources can be simulated using the first reverberation sample. The first and second sample sets are used for training, so that the trained speech detection model can not only accurately identify effective human voices and ineffective environmental noise, but also effectively identify background noisy human voices that belong to ineffective speech, thereby improving the accuracy and practicality of speech detection. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.
[0020] Figure 1 A schematic diagram illustrating the implementation environment of a speech detection method provided in this application embodiment; Figure 2 A schematic flowchart illustrating a model training method provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a speech detection model provided in an embodiment of this application; Figure 4 A schematic diagram of a triplet training process provided in an embodiment of this application; Figure 5 A schematic diagram illustrating a model training process provided in an embodiment of this application; Figure 6 A flowchart illustrating a speech detection method provided in an embodiment of this application; Figure 7 A schematic diagram illustrating an application example of speech detection provided in this application embodiment; Figure 8 This is a schematic diagram of the structure of a model training device provided in an embodiment of this application; Figure 9 This is a schematic diagram of the structure of a voice detection device provided in an embodiment of this application; Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0021] The embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions of the embodiments of this application.
[0022] Those skilled in the art will understand that, unless otherwise stated, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. The terms “comprising” and “including” as used in the embodiments of this application mean that the corresponding feature can be implemented as the presented feature, information, data, step, or operation, but do not exclude implementation as other features, information, data, steps, or operations supported by this art.
[0023] It is understood that, in the specific embodiments of this application, any user-related data, such as user audio signals and the process of collecting audio from the user's room, requires user permission or consent when applied to specific products or technologies. Furthermore, the collection, use, and processing of this data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. In other words, if any user-related data is involved in the embodiments of this application, such data must be obtained with the user's authorization and consent, and in accordance with the relevant laws, regulations, and standards of the country and region.
[0024] Figure 1 This is a schematic diagram illustrating the implementation environment of a speech detection method provided in this application. For example... Figure 1 As shown, the implementation environment includes a server 101 and a terminal 102. The server 101 can be a backend server for an application. The terminal 102 has an application installed, and the terminal 102 and the server 101 can interact with each other based on the application.
[0025] This application can have a voice detection function. For example, for business scenarios supported by the application, such as online meetings, online video or voice, online live streaming, and online teaching, the application can perform voice detection on the collected audio signals to detect valid human voices that are valid speech, as well as background noise and other noises that are invalid speech.
[0026] The server 101 can pre-train a speech detection model using a large number of samples and then distribute the speech detection model to the terminal 102, enabling the terminal 102 to use the speech detection model to detect valid speech, background voices, environmental noise, etc. in an audio segment. Alternatively, the terminal 102 can directly send the collected audio signal to the server 101, which will then use the trained speech detection model for detection.
[0027] It should be noted that server 101 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server or server cluster providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The aforementioned networks can include, but are not limited to, wired networks and wireless networks. Wired networks include local area networks (LANs), metropolitan area networks (MANs), and wide area networks (WANs). Wireless networks include Bluetooth, Wi-Fi, and other networks that enable wireless communication. Terminal 102 can be an in-vehicle terminal, intelligent voice interaction device, smart home appliance, smartphone (such as Android phones, iOS phones, etc.), tablet computer, laptop computer, digital broadcast receiver, MID (Mobile Internet Devices), PDA (Personal Digital Assistant), desktop computer, in-vehicle terminal (such as in-vehicle navigation terminal, in-vehicle computer, etc.), smart speaker, smartwatch, etc. Terminals and servers can be directly or indirectly connected via wired or wireless communication, but are not limited to these methods. The specific application can be determined based on the actual application scenario requirements, and is not limited here. The embodiments of this application can be applied to the transportation and autonomous driving fields; for example, they can be used in various scenarios such as cloud technology, artificial intelligence, in-vehicle scenarios, assisted driving, and intelligent transportation.
[0028] Figure 2 This is a flowchart illustrating a model training method provided in an embodiment of this application. The subject executing this method can be an electronic device. Figure 2 As shown, the method includes the following steps 201-204.
[0029] Step 201: The electronic device acquires the first sample set and the second sample set.
[0030] The sample labels corresponding to the first sample set and the second sample set respectively represent invalid speech and valid speech; the first sample set includes at least one first reverberation sample and a noise sample, and the second sample set includes at least one valid speech sample.
[0031] The effective speech sample includes effective speech data from human voice sources; the first reverberation sample is obtained by propagating the effective speech data in sample rooms with different reverberation intensities; and the noise sample includes environmental noise data from non-human voice sources.
[0032] For example, valid speech refers to speech from a valid human voice source. Invalid speech refers to speech or environmental noise other than valid human voice sources. A valid human voice source can be understood as a valid, active human voice source within the current environment, such as a human voice source that continuously and stably emits speech.
[0033] In one possible scenario example, in a meeting setting, the speaker currently speaking in the meeting room is the valid human voice source, and the speech emitted by the speaker is valid speech. The whispers of other people in the meeting room besides the speaker are considered background noise, while sounds such as doors opening and chairs being moved are considered environmental noise. Both background noise and environmental noise are considered invalid speech.
[0034] In one possible implementation, the second sample set further includes at least one second reverberation sample, and the determination of the at least one first reverberation sample and the second reverberation sample includes at least one of the following steps A1 and A2: Step A1: Perform reverberation processing on the effective speech data based on at least one first impulse response to obtain at least one first reverberation sample; and perform reverberation processing on the effective speech data based on at least one second impulse response to obtain at least one second reverberation sample.
[0035] The first impulse response or the second impulse response characterizes the propagation characteristics of effective speech in the corresponding sample room, and the reverberation intensity corresponding to the first impulse response is greater than that of the second impulse response.
[0036] The first and second reverberation samples are reverberant speech samples obtained by processing effective speech data based on the room impulse response of at least one sample room. These first and second impulse responses characterize the acoustic properties of different effective speech data propagating within the sample room. In this embodiment, the room impulse response can be used to process the effective speech data to simulate the acoustic properties of the effective speech data propagating within the sample room, thereby obtaining reverberant speech samples. For example, when an audio signal propagates within a sample room, reflection and refraction may occur, resulting in reverberation during the audio signal propagation process. The room impulse response can be used to simulate the system response of the sample room to speech propagation.
[0037] Different impulse responses correspond to different acoustic characteristics, resulting in different reverberation intensities for the obtained reverberant speech samples. Specifically, the reverberation intensity of the second impulse response corresponding to the second reverberant sample is less than the reverberation intensity of the first impulse response corresponding to the first reverberant sample.
[0038] In one possible scenario, in some conference rooms, the speaker (the target speaker, a valid human voice source) is generally closer to the audio acquisition device, while others (non-valid human voice sources, such as background noise) are farther away. Therefore, the speaker's voice acquired by the audio acquisition device is clearer and has lower reverberation, while the voices of others are more blurred and have higher reverberation. In this application, by constructing a first reverberation sample with higher reverberation intensity into a first sample set of non-valid voices, and a second reverberation sample with lower reverberation intensity into a second sample set of valid voices, the initial speech detection model is subsequently trained based on the first and second sample sets. This allows the initial speech detection model to learn the difference between valid human voice sources with lower reverberation intensity and background noise with higher reverberation intensity, enhancing its knowledge of reverberation intensity in relation to valid human voice sources and background noise. This enables the trained speech detection model to more accurately identify the speaker's valid voice and background noise in real-world scenarios, improving the accuracy of speech recognition.
[0039] In this embodiment, a pre-built RIR (Room Impulse Response) library can be constructed, which includes room impulse responses corresponding to multiple sample rooms. These sample rooms can correspond to different room sizes, different room impulse responses, and different reverberation intensities. For example, this RIR library characterizes the room system responses of various sample rooms with different sizes and reverberation intensities; for instance, this RIR library can be generated using the mirror method and simulation tools.
[0040] For example, reverberation intensity characterizes the relative strength between reverberant sound and direct sound generated during audio propagation. The greater the reverberation intensity, the stronger the reverberant sound and the weaker the direct sound; conversely, the smaller the reverberation intensity, the stronger the direct sound and the weaker the reverberant sound. Reverberant sound can be sound waves that have undergone reflection, refraction, etc., in audio, while direct sound can be understood as sound waves that arrive directly at the ear from the sound source without reflection. The reverberation intensity corresponding to each first reverberation sample exceeds the target intensity, while the reverberation intensity corresponding to each second reverberation sample does not exceed the target intensity. That is, the first reverberation sample can be obtained using the room impulse response with a reverberation intensity exceeding the target intensity, and the second reverberation sample can be obtained using the room impulse response with a reverberation intensity not exceeding the target intensity.
[0041] Step A2: Based on the target human voice source corresponding to each valid speech sample in the second sample set, perform reverberation processing on at least one valid speech data from the target human voice source to obtain at least one second reverberation sample, and perform reverberation processing on at least one valid speech data from a non-target human voice source to obtain at least one first reverberation sample.
[0042] In this step, based on the target human voice source to which the valid speech data in the valid speech samples belong, the second reverberation sample of the same human voice source as the valid speech sample can be constructed into the second sample set; the first reverberation sample of the human voice source that is different from the valid speech sample can be constructed into the first sample set.
[0043] For example, based on the room impulse response of at least one sample room, reverberation processing can be performed on valid speech data from each target human voice source to obtain at least one second reverberation sample. Also, based on the room impulse response of at least one sample room, reverberation processing can be performed on valid speech data from each other human voice source besides the target human voice source to obtain at least one first reverberation sample.
[0044] In one possible scenario, a meeting setting, the speaker (the target speaker belonging to the effective human voice source) is a fixed speaker. The speaker's voice source is the effective speech, while other voice sources are background noise. Therefore, speech from the same speaker as the effective speech sample can be reverberated and used as the second reverberation sample of the effective speech. Speech from different speakers than the effective speech sample can be reverberated and used as the first reverberation sample of the background noise.
[0045] In one possible implementation, noise can be further added to the reverberated speech data. Accordingly, step A1 can be implemented by including the following steps A11-A12: Step A11: Perform convolution processing on each first impulse response and effective speech data to generate initial reverberation data corresponding to each first impulse response; Step A12: Obtain noise data from a pre-configured noise library, and synthesize the initial reverberation data corresponding to each first impulse response and the obtained noise data according to at least one signal-to-noise ratio to obtain the at least one first reverberation sample.
[0046] For example, the room impulse response can be represented by the corresponding system response function, and convolution processing can be performed using the system response function corresponding to the room impulse response and the effective speech data.
[0047] For any sample room, the room impulse response can be represented by the corresponding RIR function, for example, the RIR function is represented as h(n). The expansion of the RIR function h(n) refers to representing the RIR function as a discrete-time sequence. The form of the expansion of h(n) can be expressed as: h(n) = δ(n) + α1δ(n- T1) + α2δ(n-T2) +……+ α N δ(n -T N); Where δ(n) represents the unit impulse response function, which can be represented by a pulse signal at time n. α1, α2, ..., α N The attenuation coefficients corresponding to each pulse signal are T1, T2, ..., T. N This indicates the delay time corresponding to each pulse signal.
[0048] It should be noted that each term in the expansion of h(n) represents a pulse signal, corresponding to the propagation of effective speech along different paths. The first term δ(n) represents the direct path, that is, the effective speech propagates directly from the sound source to the human ear or a receiver such as a speech acquisition device; the subsequent terms represent multiple reflection paths such as reflection, refraction, and scattering.
[0049] For example, valid speech data may include audio features corresponding to the valid speech at each sampling point, such as the amplitude of the valid speech sampled at each sampling point.
[0050] In one possible example, the effective speech data is represented by x(n). Accordingly, in step A11, the room impulse response and the effective speech data can be convolved using the following formula to obtain the initial reverberation data: y(n) = x(n) * h(n); Where y(n) represents the initial reverberation data, and * indicates the convolution calculation of the effective speech data x(n) and the RIR function h(n) of the room impulse response.
[0051] For example, the electronic device can further mix environmental noise data into the initial reverberation data to more realistically reproduce the actual speech detection scenario. In one possible example, in step A12, the initial reverberation data and noise data can be synthesized according to the signal-to-noise ratio using the following formula 1 to obtain the first reverberation sample: Formula 1: n = γ(a*r + βb); Where n represents the first reverberation sample; a represents the effective speech data, specifically represented by x(n); r represents the first impulse response, specifically represented by the RIR function h(n) of the room impulse response. β represents the signal-to-noise ratio; for example, β ranges from 0 to 1 and can be used to simulate different noise levels. b represents the noise data. γ represents the volume intensity; for example, γ ranges from 0 to 1 and can be used to simulate different volume levels.
[0052] It should be noted that one or more noise data can be synthesized on the initial reverberation data. In Formula 1, noise data b can represent one or more noise data; if it represents multiple noise data, then βb = β1b1 + β2b2 + ... + β m b mDifferent noise data may correspond to different signal-to-noise ratios.
[0053] For example, in step A12, for each initial reverberant speech data, at least one noise data is obtained from a pre-built noise library, and based on the timestamp information of the initial reverberant speech data and the timestamp information of the at least one noise data, the initial reverberant speech data and the at least one noise data are synthesized to obtain at least one reverberant speech data. For example, the initial reverberant speech data and noise data at corresponding sampling points on the two time axes can be synthesized according to each moment on the time axis of the initial reverberant speech data and each moment on the time axis of the noise data; for example, for a sampling point at a certain moment t1, the amplitude of the initial reverberant speech data a at moment t1 is a1, and the amplitude of the noise data b at moment t1 is b1. After synthesis, the amplitude of the synthesized reverberant speech data at moment t1 is a1+b1.
[0054] It should be noted that multiple second reverberation samples can also be constructed in the manner described in steps A11-A12. That is, using the above formula 1, the impulse response function corresponding to the second impulse response and the effective speech data are convolved, and noise is further increased according to at least one signal-to-noise ratio to obtain multiple second reverberation samples. The method for determining the first reverberation sample is the same, and will not be elaborated here. In addition, step A2 can also be implemented in the same way as steps A11-A12. That is, using the above formula 1, the effective speech data from the non-target human voice source and the corresponding impulse response function are convolved, and the effective speech data from the target human voice source and the corresponding impulse response function are convolved; and noise is further increased according to at least one signal-to-noise ratio to obtain each first reverberation sample and second reverberation sample. The method for determining the first reverberation sample is the same, and will not be elaborated here.
[0055] In one possible approach, steps A1 and A2 above can be combined to determine the first reverberation sample and the second reverberation sample. For example, for valid speech data from the target human voice source, a second impulse response with a smaller reverberation intensity can be used for reverberation processing to obtain each second reverberation sample; for valid speech data from a non-target human voice source, a first impulse response with a larger reverberation intensity can be used for reverberation processing to obtain each first reverberation sample.
[0056] In this application, by utilizing multiple signal-to-noise ratios and various room impulse responses, a first reverberation sample with a larger reverberation intensity and a second reverberation sample with a smaller reverberation intensity are constructed. These samples simulate background noisy human voices and the speech of the speaker from an effective human voice source in the actual speech detection environment, respectively. This improves the sample richness of background noisy human voices in ineffective speech and makes the second sample set closer to the effective human voice source in the actual speech environment, thereby improving the accuracy of training.
[0057] In one possible implementation, the construction process of each reverberant speech sample further includes a sample label determination step. Correspondingly, the determination of the at least one first reverberant sample and the second reverberant sample further includes at least one of the following steps B1 and B2: Step B1: Determine the sample labels of the first reverberation sample and the second reverberation sample based on the reverberation intensity corresponding to each of the first and second reverberation samples respectively; Step B2: Determine the sample labels of the first reverberation sample and the second reverberation sample based on their respective signal-to-noise ratios.
[0058] For example, in Formula 1, a larger reverberation intensity corresponding to the room impulse response r indicates a stronger relative reverberation sound compared to the direct sound in the reverberant speech sample. When the reverberation intensity exceeds the target intensity, the sample label representation of the reverberant speech sample can be determined to be invalid speech. Conversely, if the reverberation intensity does not exceed the target intensity, the sample label representation of the reverberant speech sample can be determined to be valid speech. For instance, a reverberant speech sample with a large reverberation intensity exceeding the target intensity can be used to simulate background noise in a real speech detection environment.
[0059] For example, in Formula 1, the signal-to-noise ratio (SNR) represents the ratio between effective speech and ambient noise. A higher SNR indicates stronger effective speech and lower noise. If the SNR exceeds a target threshold, the sample label of the reverberant speech sample can be determined to be effective speech; if the SNR does not exceed the target threshold, the sample label of the reverberant speech sample can be determined to be ineffective speech. For instance, a reverberant speech sample with a higher SNR can be used to simulate effective speech, while a reverberant speech sample with a lower SNR can be used to simulate background noise such as human voices or ambient noise.
[0060] Of course, steps B1 and B2 can also be used together to determine the sample label. For example, when either the reverberation intensity exceeds the target intensity or the signal-to-noise ratio exceeds the target threshold, the sample label of the reverberant speech sample can be determined to be ineffective speech.
[0061] In one possible approach, the sample label can be a category label representing a class, or it can be a label representing different values of the effective speech confidence of the reverberant speech sample. For example, taking reverberation intensity as an example, step B1 can be implemented in a manner that includes at least one of the following methods 1 and 2: Method 1: Use the first category label representing non-effective speech as the sample label of the first reverberation sample, and use the second category label representing effective speech as the sample label of the second reverberation sample; Method 2: Based on the reverberation intensity of the first reverberation sample and the second reverberation sample respectively, obtain the effective speech confidence of the first reverberation sample and the second reverberation sample respectively, and use the effective speech confidence of the first reverberation sample and the second reverberation sample respectively as the sample label of the first reverberation sample and the second reverberation sample respectively. The reverberation intensity is negatively correlated with the effective speech confidence.
[0062] In Method 1, if the reverberation intensity exceeds the target intensity, the sample label of the reverberated speech sample is the first label representing invalid speech, meaning the reverberated speech sample is the first reverberation sample. If the reverberation intensity does not exceed the target intensity, the sample label is the second label representing valid speech, meaning the reverberated speech sample is the second reverberation sample. For example, the sample label can be represented by 0 or 1, where 1 represents the second label representing valid speech and 0 represents the first label representing invalid speech.
[0063] In Method 2, reverberation intensity can be expressed numerically. For example, reverberation intensity can be represented as 1, 0.9, 0.2, 0.05, etc. A larger reverberation intensity value indicates a weaker direct sound and a stronger reverberation; correspondingly, the effective speech confidence is lower. Based on the effective speech confidence, different values can be used to numerically represent the sample labels; a larger reverberation intensity results in a lower effective speech confidence and a smaller sample label value, while a smaller reverberation intensity results in a higher effective speech confidence and a larger sample label value. For example, the sample label can be the effective speech confidence, which can range from 0 to 1. A higher confidence value indicates a stronger direct sound and a weaker reverberation in the reverberated speech sample.
[0064] In one possible approach, the electronic device can map reverberation intensity to effective speech confidence according to a pre-configured mapping relationship. For example, the mapping relationship can be represented by a mapping relationship list, which may include multiple reverberation intensity value ranges and the corresponding effective speech confidence for each reverberation intensity value range. Based on this mapping relationship list, the numerical range of reverberation intensity for each reverberated speech sample can be found, and the corresponding effective speech confidence can be obtained. Alternatively, a specific mapping function, such as a linear function, can be used. This application does not specifically limit the mapping method from reverberation intensity to effective speech confidence.
[0065] It should be noted that the reverberation intensity of a reverberant speech sample can be determined by at least one of the following: obtaining the energy attenuation corresponding to the reverberant speech sample, and obtaining the reverberation time corresponding to the reverberant speech sample.
[0066] Among them, energy attenuation represents the degree of attenuation of effective speech during propagation. It can be the ratio of the signal energy of the direct sound and the reverberant sound (sound signal after reflection, refraction, etc.) corresponding to the reverberant speech sample. The greater the energy attenuation, the greater the reverberation intensity.
[0067] Reverberation time, denoted as T60 or RT, is the time required for the sound pressure level to decrease by 60 dB after the sound source of the effective speech data corresponding to the reverberant speech sample stops emitting sound; it is measured in seconds. A longer reverberation time indicates a greater reverberation intensity.
[0068] In one possible implementation, a pre-configured valid speech library and noise library can also be used to construct valid speech samples. For example, the at least one valid speech sample is constructed by performing the following steps: Obtain at least one valid speech data from a pre-configured valid speech library, and obtain at least one noise data from a pre-configured noise library; According to at least one signal-to-noise ratio, the acquired effective speech data and noise data are synthesized to obtain at least one effective speech sample.
[0069] For example, the electronic device can further mix environmental noise data with the valid speech data to more realistically reproduce the actual speech detection scenario. In one possible example, the valid speech data and noise data can be synthesized according to the signal-to-noise ratio using the following formula 2 to obtain the valid speech sample: Formula 2: p = γ(a +βb); Where p represents a valid speech sample; a represents valid speech data; b represents noise data; β represents the signal-to-noise ratio; for example, β ranges from 0 to 1 and can be used to simulate different noise levels. γ represents volume intensity; for example, γ ranges from 0 to 1 and can be used to simulate different volume levels.
[0070] For example, a second label representing the valid speech can be used for a valid speech sample. Alternatively, the sample label can be numerically represented; for example, the signal-to-noise ratio (SNR) can be mapped to the valid speech confidence level, and the valid speech confidence level can be used to represent the sample label of the valid speech sample. For example, the SNR corresponding to a valid speech sample is positively correlated with the valid speech confidence level, that is, the higher the SNR, the stronger the valid speech and the weaker the noise, and correspondingly, the higher the valid speech confidence level. For example, an SNR of 0.9 corresponds to a valid speech confidence level of 0.9; an SNR of 0.7 corresponds to a valid speech confidence level of 0.7, and so on. This application does not limit the specific representation of the sample label.
[0071] In one possible approach, a pre-configured noise library can also be used to construct noise samples. For example, each noise sample is constructed by performing the following steps: obtaining at least one noise data point from the pre-configured noise library as initial data, and processing the obtained initial data point based on at least one volume intensity to obtain each noise sample.
[0072] For example, noise samples can be obtained by processing the data obtained from the noise database according to the volume intensity using the following formula 3: Formula 3: n = βb; Wherein, β represents the volume intensity, and the value of β ranges from 0 to 1, used to simulate noise of different volume levels.
[0073] In this application, formulas 1 to 3 above can be used to construct a large number of diverse samples, including background noise, effective human voices, and environmental noise.
[0074] For example, the sample label of a noise sample can be a first label representing ineffective speech. Additionally, the sample label of a noise sample can further characterize environmental noise belonging to the ineffective speech.
[0075] Step 202: Based on the audio features of each sample in the first and second sample sets, the electronic device uses an initial speech detection model to perform speech detection on each sample, and obtains the sample detection results for each sample.
[0076] The sample detection result indicates whether the corresponding sample belongs to valid speech.
[0077] In this step, the electronic device can extract the audio features of each sample, input the audio features of each sample into the initial speech detection model, and use the initial speech detection model to perform speech detection on the samples to obtain the corresponding speech detection results. The audio features can be FFT (Fast Fourier Transform) features of the samples.
[0078] For example, each sample is a time-domain signal, and the sample data includes the characteristics of each sampling point of the audio in the time domain, such as the amplitude of each sampling point. In this step, each sample can first be segmented and windowed. Then, the FFT transform is performed on each frame signal, and the FFT coefficients of each frame signal in the frequency domain are extracted. The modulus of the FFT coefficients of each frame signal is taken, and the resulting spectral features are used as the FFT features of the sample.
[0079] For example, the speech detection model can employ an LSTM (Long Short-Term Memory) network. For instance, as... Figure 3 As shown, an effective speech detection model can include a multi-layer LSTM network and a fully connected layer (FC). The input audio FFT features can first undergo convolution operations in the FC layer, and then be further processed by the LSTM network to extract the contextual features of the samples. This process is repeated using the FC layer and the LSTM network. Finally, the features obtained after these repeated operations are input into an FC layer for classification to obtain the detection result of the samples. For example, the FC layer can be used to obtain the probability of a sample corresponding to each category, such as the probability of belonging to valid speech and the probability of belonging to invalid speech. The final detection result is obtained based on the probabilities of each category.
[0080] In one possible implementation, the sample label of the noise sample also represents environmental noise in the invalid speech; the sample label of the first reverberant sample also represents background reverberant speech in the invalid speech. In this step, the initial speech detection model can also be used to further detect environmental noise and background reverberant speech in the invalid speech samples; that is, it can detect whether a sample belongs to background reverberant speech in the invalid speech or to environmental noise. Correspondingly, For example, the sample detection result includes at least first result indication information, which indicates whether the sample belongs to valid speech. If the first result indication information indicates that the sample does not belong to valid speech, the sample detection result also includes invalid speech reference information, which indicates that the sample belongs to background reverberant speech or environmental noise in invalid speech.
[0081] Step 203: The electronic device determines the total training loss corresponding to each sample based on the sample detection results and sample labels of each sample.
[0082] In this step, the electronic device can measure the total training loss corresponding to each iteration of training based on the distance between the sample detection result and the sample label of each sample in each iteration of training.
[0083] For example, the electronic device can use the Cross-Entropy Loss (CE Loss) function to measure the degree of difference between the distribution of the sample label of each sample and the sample detection result of each training iteration, in order to obtain the training loss. For example, the electronic device can use the following formula 4 to calculate the similarity between the sample detection result and the sample label of each sample, in order to obtain the total training loss: Formula 4: Loss1= ; Here, Loss1 represents the cross-entropy loss for each sample. The labels represent the actual samples, for example, a first label representing valid speech or a second label representing invalid speech. This represents the probability of the network output belonging to valid or invalid speech, which can be calculated using a normalized exponential function (softmax) corresponding to the i-th category. In this application, iterative processing is performed based on the training loss to make the classification results of the initial speech detection model increasingly closer to the true sample labels.
[0084] In one possible implementation, the distance between samples in the feature space can be further utilized to iteratively optimize the network. For example, step 203 can be implemented by including the following steps 2031-2035: Step 2031: Based on the similarity between the sample detection results and sample labels of each sample, determine the first training loss corresponding to each sample; In one possible example, the first training loss can be determined using the above formula 4, that is, the first training loss can be the cross-entropy loss corresponding to each sample.
[0085] In another possible example, the first training loss can also be determined using the MSE LOSS (Mean Squared Error Loss) function. For example, the first training loss for each sample can be determined using the following formula 5: Formula 5: Loss2= ; Where n is the number of samples. It is the sample label of the i-th sample, for example, The value can be the effective speech confidence level. It is the speech detection result of the i-th sample, such as the probability that it belongs to valid speech.
[0086] Step 2032: Based on the sample labels of each sample, determine at least one triplet, each triplet including a valid speech sample, a first reverberation sample and a second reverberation sample; In this step, the electronic device can construct multiple triples by using valid speech samples as anchor examples, second reverberation samples belonging to valid speech as positive examples, and first reverberation samples belonging to invalid speech as negative examples. Alternatively, noise samples and the first reverberation sample can be used as negative examples. That is, a triple can include one valid speech sample, one negative example, and one second reverberation sample; the negative example can be either the first reverberation sample or a noise sample.
[0087] Step 2033: For each triplet, determine the first feature distance between the audio features of the valid speech sample in the triplet and the first reverberation sample; and determine the second feature distance between the audio features of the valid speech sample in the triplet and the second reverberation sample. The audio features of each sample in the triplet can be obtained by further feature extraction from the initial speech detection model. For example, the audio FFT features of each sample can be input into the initial speech detection model and then... Figure 3 After the FC layer and LSTM network in the initial speech detection model further convolutionally operate on the input audio FFT features, the features obtained from the FC layer and LSTM network in the initial speech detection model are used as the audio features in step 2033. For example, the features obtained from the FC layer and LSTM network in the initial speech detection model can be used as the audio features in step 2033. Figure 3 The feature matrix output by the first or second LSTM is used as the audio feature of each sample in the triplet, so as to calculate the feature distance using the audio features.
[0088] Step 2034: Based on the first feature distance and the second feature distance corresponding to each triplet, determine the second training loss corresponding to each triplet; For example, the first feature distance and the second feature distance corresponding to the triple can be determined using the following formula 6, and the second training loss can be obtained based on the first feature distance and the second feature distance: Formula 6: Loss3= max(d(A, P) - d(A, N) + margin, 0); Where A represents a valid speech sample, also known as an anchor; P represents the second reverberation sample, also known as a positive sample. In one example, N can represent the first reverberation sample, also known as a negative sample; in another example, the negative sample can also include noise samples, meaning N can represent both the first reverberation sample and the noise sample. d(A, P) represents the second feature distance; d(A, N) represents the first feature distance; that is, d(A, P) and d(A, N) are the distances in the feature space between the anchor and the positive sample, and between the anchor and the negative sample, respectively.
[0089] Here, margin is a pre-set distance threshold used to control the difference between positive and negative examples.
[0090] Ideally, the distance between the anchor example and the negative example should be at least greater than the distance between the anchor example and the positive example. For example... Figure 4 As shown, in the early stages of training, the distance between the anchor example and the positive example is greater than the distance between the anchor example and the negative example. As the initial speech detection model is continuously trained iteratively, in the feature space, the distance between the anchor example and the positive example narrows, while the distance between the anchor example and the negative example widens, making the distance between the anchor example and the positive example smaller than the distance between the anchor example and the negative example.
[0091] Step 2035: Based on the first training loss corresponding to each sample and the second training loss corresponding to each triplet, obtain the total training loss.
[0092] For example, the sum of the first training loss and the second training loss can be used as the total training loss. For instance, the total loss Loss = Loss1 + Loss3; or, the total loss Loss = Loss2 + Loss3.
[0093] Step 204: The electronic device iteratively trains the initial speech detection model based on the total training loss corresponding to each sample to obtain the speech detection model.
[0094] In this step, the electronic device can use the directional propagation algorithm to calculate the gradient of the total training loss with respect to the network parameters in the initial speech detection model, and then use an optimization algorithm to update the network parameters. For example, the optimization algorithm can use stochastic gradient descent. Based on this, the network detection results can be made closer to the sample labels, thereby reducing the value of the training loss. When the initial speech detection model reaches the iteration stopping condition, the iterative training stops, and the speech detection model is obtained. For example, the iteration stopping condition may include, but is not limited to: the number of iterations exceeding the target number, the total training loss being less than the target loss threshold multiple times consecutively, etc.
[0095] like Figure 5 As shown, this application mainly includes two processes: audio FFT feature extraction and effective speech detection model training. For audio FFT feature extraction, operations such as framing and windowing of the audio time-domain signal can be performed, and FFT features are extracted from multiple frames. For the network training process, steps 201-204 described above can be used to iteratively train the initial speech detection model to obtain the final speech detection model.
[0096] In related technologies, speech detection is performed using methods such as energy-based detection and statistical model-based detection. Energy-based detection specifically compares the short-time energy of the audio signal with an energy threshold; if the short-time energy exceeds the energy threshold, the audio is considered speech; otherwise, it is considered non-speech. Statistical model-based detection primarily utilizes modeling techniques for speech and non-speech components. For example, a speech GMM (Gaussian Mixture Model) statistical model and a non-speech GMM statistical model can be modeled separately for detection.
[0097] However, none of the methods in the relevant technologies can detect background noise. In some application scenarios, if these unfiltered background noises are sent to a speech recognition device, they will be recognized as meaningless sentences, affecting the accuracy of speech recognition.
[0098] To address the technical problems in related technologies, this application embodiment designs a model training method. The trained speech detection model can effectively detect background noise, thereby effectively filtering out background noise and ensuring the accuracy and precision of valid speech sources. The technical framework adopted in this application embodiment is based on a deep neural network method, rather than a statistical model based on GMM. This allows for the differentiation of background noise beyond just distinguishing between human speech and noise, thus preventing background noise from being mistaken for valid speech. Furthermore, this application embodiment constructs a rich sample containing a large number of background noise samples by simulating valid speech and using reverberant speech to simulate background noise. During the training phase, valid human speech is treated as one category, while background noise and noise are treated as another category to train the deep neural network model. This enables the trained deep neural network to accurately distinguish between valid speech and background noise, improving the accuracy and practicality of speech detection.
[0099] The model training method provided in this application obtains a first sample set of ineffective speech and a second sample set of effective speech. The first sample set includes a first reverberation sample and a noise sample. The effective speech sample includes effective speech data from human voice sources. The first reverberation sample is obtained by propagating the effective speech data in sample rooms with different reverberation intensities. The noise sample includes environmental noise data from non-human voice sources. Based on this, the speech of background human voice sources can be simulated using the first reverberation sample. The first and second sample sets are used for training, so that the trained speech detection model can not only accurately identify effective human voices and ineffective environmental noise, but also effectively identify background noisy human voices that belong to ineffective speech, thereby improving the accuracy and practicality of speech detection.
[0100] Furthermore, the trained speech detection model can effectively assist in speech recognition or voiceprint recognition tasks. For speech recognition tasks, the speech detection model trained using the method described in this application can further improve the robustness of the speech detection process, accurately filtering out background noise in addition to effectively filtering out noise. For voiceprint recognition tasks, the speech detection model described in this application, in addition to filtering out noise, can also finely distinguish between valid human voices and background noise, thereby improving the accuracy of voiceprint recognition. This further enhances the practicality of the model training method and the speech detection method.
[0101] Figure 6 This is a flowchart illustrating a speech detection method provided in an embodiment of this application. The method can be executed by an electronic device. Figure 6 As shown, the method includes the following steps 601-602.
[0102] Step 601: The electronic device acquires at least one frame of audio signal from the target room where the target object is located; In some possible scenario examples, for instance, in a meeting setting, during the speaker's presentation in the meeting room, audio acquisition devices can continuously capture audio from the target object to obtain multiple frames of audio signals.
[0103] For example, in a teaching setting, audio can be continuously collected from the teacher during a lecture in the classroom, resulting in multiple frames of audio signals.
[0104] For example, in online multi-user video or voice scenarios, audio can be captured from each user's room or environment to obtain multiple frames of audio signals.
[0105] Step 602: The electronic device detects each frame of audio signal based on the trained speech detection model to obtain the detection result of each frame of audio signal. The detection result indicates whether the corresponding frame of audio signal belongs to valid speech. The speech detection model was trained using the model training method described above.
[0106] For example, the electronic device can pre-configure the speech detection model locally and use the local speech detection model to detect each frame of audio signal to determine whether each frame of audio signal belongs to valid speech or invalid speech; it can also detect whether the audio signal of each frame of invalid speech belongs to environmental noise or background reverberation speech.
[0107] In one possible implementation, the detection result includes audio indication information and target reference information. The audio indication information indicates whether the corresponding frame audio signal belongs to valid speech, and the target reference information indicates whether the corresponding frame audio signal belongs to background reverberant speech or environmental noise in invalid speech. For example, background reverberant speech may include background noise. Valid and invalid speech are distinguished based on the audio indication information. In addition, based on the reference information, it is possible to further distinguish whether the invalid speech is environmental noise or background voices from people other than the speaker.
[0108] The detection results in this application can also be further applied to other audio processing tasks. For example, the detection results can be combined to further perform speech recognition tasks, voiceprint recognition tasks, etc. Accordingly, after step 602, any one of the following steps 603 and 604 may also be included: Step 603: Based on the audio indication information of each frame of audio signal, remove the first audio signal corresponding to the background reverberation speech or environmental noise that belongs to non-effective speech from each frame of audio signal; In this step, after filtering using audio indication information, valid human voices in the audio are accurately preserved. For example, in speech recognition tasks, not only can environmental noise in an audio clip be effectively filtered out, but background noise can also be further filtered out with greater precision.
[0109] Step 604: Based on the audio indication information and target reference information of each frame of audio signal, determine the second audio signal that belongs to valid speech and the third audio signal that belongs to background reverberation speech in each frame of audio signal; and perform voiceprint recognition on each frame of audio signal based on the audio features of the second audio signal and the third audio signal to obtain the audio signals from the target object and other objects outside the target object in each frame of audio signal.
[0110] In this step, for invalid audio segments in the audio, the target reference information can be used to further distinguish between background voices and noise in the invalid audio segments. For example, in a voiceprint recognition task, the target reference information can be used to extract valid speech from valid voice sources and background reverberant speech from background voice sources. Based on this, the audio features of different voice sources can be combined to distinguish the audio of different voices in an audio segment.
[0111] For example, in a teaching scenario, if an audio clip contains questions asked by student A during a lecture in classroom A, and answers given by student B and student C, the speech detection model of this application can be used to obtain target reference information for each frame of invalid audio, determine the audio frames corresponding to student B and student C, and further obtain the content of the answers given by student B and student C in the subsequent answering process audio.
[0112] For example, in a meeting scenario, when it is necessary to generate the content spoken by different users during the discussion, the target reference information of each frame of invalid audio obtained by the speech detection model can be combined to identify the speech of different users in an audio segment and accurately generate the speech content of each user.
[0113] like Figure 7 As shown, Figure 7 The left and middle channels contain the original speech, while the right channel contains the valid speech retained based on network detection results. For example... Figure 7 As shown, the speech detection model provided in this application can effectively detect background noise and human voices as invalid speech. Table 1 below shows the metrics for the detection results of different techniques: Table 1
[0114] Precision is an indicator used to determine the accuracy of effective speech classification; a higher value indicates higher accuracy in effective speech detection. As shown in Table 1, the speech detection method of this application has higher accuracy than related technologies. The method of this application significantly improves speech detection accuracy.
[0115] The model training method provided in this application obtains a first sample set of ineffective speech and a second sample set of effective speech. The first sample set includes a first reverberation sample and a noise sample. The effective speech sample includes effective speech data from human voice sources. The first reverberation sample is obtained by propagating the effective speech data in sample rooms with different reverberation intensities. The noise sample includes environmental noise data from non-human voice sources. Based on this, the speech of background human voice sources can be simulated using the first reverberation sample. The first and second sample sets are used for training, so that the trained speech detection model can not only accurately identify effective human voices and ineffective environmental noise, but also effectively identify background noisy human voices that belong to ineffective speech, thereby improving the accuracy and practicality of speech detection.
[0116] Furthermore, the trained speech detection model can effectively assist in speech recognition or voiceprint recognition tasks. For speech recognition tasks, the speech detection model trained using the method described in this application can further improve the robustness of the speech detection process, accurately filtering out background noise in addition to effectively filtering out noise. For voiceprint recognition tasks, the speech detection model described in this application, in addition to filtering out noise, can also finely distinguish between valid human voices and background noise, thereby improving the accuracy of voiceprint recognition. This further enhances the practicality of the model training method and the speech detection method.
[0117] The model training method and speech detection method provided in this application involve the aforementioned artificial intelligence technology, machine learning, speech technology, and other technologies, and can be applied to various technical fields such as online conferencing, intelligent transportation, and autonomous driving.
[0118] In essence, Artificial Intelligence (AI) refers to the theories, methods, technologies, and application systems that utilize digital computers or computers-controlled machines to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce new intelligent machines capable of reacting in a manner similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.
[0119] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics. Among these, pre-trained models, also known as large-scale models or foundational models, can be widely applied to downstream tasks across various AI fields after fine-tuning. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0120] Understandably, key technologies in speech technology include Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and Voiceprint Recognition. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, with speech emerging as one of the most promising methods. Large-scale modeling has revolutionized speech technology; pre-trained models such as WavLM and UniSpeech, which utilize the Transformer architecture, possess strong generalization and versatility, enabling them to excel in various speech processing tasks.
[0121] Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and pre-trained learning. Pre-trained models are the latest development in deep learning, integrating all of these techniques.
[0122] Autonomous driving technology refers to vehicles driving themselves without driver intervention. It typically includes technologies such as high-precision mapping, environmental perception, computer vision, behavioral decision-making, path planning, and motion control. Autonomous driving encompasses various development paths, including single-vehicle intelligence, vehicle-to-infrastructure (V2I) communication, and networked cloud control. Autonomous driving technology has broad application prospects, currently focusing on logistics, public transportation, taxis, and intelligent transportation systems, and is expected to see further development in the future.
[0123] With the research and advancement of artificial intelligence (AI) technology, AI is being studied and applied in various fields, such as smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, digital twins, virtual humans, robots, AI-generated content (AIGC), conversational interaction, smart healthcare, smart customer service, and game AI. It is believed that with the development of technology, AI will be applied in more fields and play an increasingly important role.
[0124] Figure 8 This is a schematic diagram of a model training device provided in an embodiment of this application. Figure 8 As shown, the device includes: The sample acquisition module 801 is used to acquire a first sample set and a second sample set, wherein the sample labels corresponding to the first sample set and the second sample set respectively represent invalid speech and valid speech; the first sample set includes at least one first reverberation sample and a noise sample, and the second sample set includes at least one valid speech sample. The effective speech sample includes effective speech data from human voice sources; the first reverberation sample is obtained by propagating the effective speech data in sample rooms with different reverberation intensities; and the noise sample includes environmental noise data from non-human voice sources. The first detection module 802 is used to perform speech detection on each sample based on the sample audio features of each sample in the first sample set and the second sample set, using an initial speech detection model to obtain the sample detection result of each sample. The sample detection result represents whether the corresponding sample belongs to valid speech. The total loss determination module 803 is used to determine the total training loss corresponding to each sample based on the sample detection results and sample labels of each sample. Training module 804 is used to iteratively train the initial speech detection model based on the total training loss corresponding to each sample to obtain the speech detection model.
[0125] In one possible implementation, the second sample set further includes at least one second reverberation sample; the sample acquisition module, when determining at least one first reverberation sample and the second reverberation sample, includes any one of the following: The first determining unit is configured to perform reverberation processing on valid speech data based on at least one first impulse response to obtain at least one first reverberation sample, and to perform reverberation processing on valid speech data based on at least one second impulse response to obtain at least one second reverberation sample. The first impulse response or the second impulse response characterizes the propagation characteristics of effective speech in the corresponding sample room, and the reverberation intensity corresponding to the first impulse response is greater than that of the second impulse response. The second determining unit is configured to perform reverberation processing on at least one valid speech data from the target human voice source based on the target human voice source corresponding to each valid speech sample in the second sample set, to obtain at least one second reverberation sample, and to perform reverberation processing on at least one valid speech data from a non-target human voice source, to obtain at least one first reverberation sample.
[0126] In one possible implementation, the first determining unit is used for: Convolution processing is performed on each first impulse response and effective speech data to generate initial reverberation data corresponding to each first impulse response; Noise data is obtained from a pre-configured noise library, and the initial reverberation data corresponding to each first impulse response and the obtained noise data are synthesized according to at least one signal-to-noise ratio to obtain the at least one first reverberation sample.
[0127] In one possible implementation, the sample acquisition module, when determining at least one first reverberation sample and a second reverberation sample, further includes at least one of the following: The third determining unit is used to determine the sample label of the first reverberation sample and the sample label of the second reverberation sample based on the reverberation intensity corresponding to the first reverberation sample and the second reverberation sample, respectively. The fourth determining unit is used to determine the sample label of the first reverberation sample and the sample label of the second reverberation sample based on the signal-to-noise ratio of the first reverberation sample and the second reverberation sample, respectively.
[0128] In one possible implementation, the third determining unit is used for at least one of the following: The first category label representing non-effective speech is used as the sample label for the first reverberation sample, and the second category label representing effective speech is used as the sample label for the second reverberation sample. Based on the reverberation intensity of the first and second reverberation samples respectively, the effective speech confidence scores of the first and second reverberation samples are obtained respectively, and the effective speech confidence scores of the first and second reverberation samples are used as the sample labels of the first and second reverberation samples respectively. The reverberation intensity is negatively correlated with the effective speech confidence score.
[0129] In one possible implementation, the second sample set further includes at least one second reverberation sample; the total loss determination module is used for: Based on the similarity between the sample detection results and sample labels of each sample, the first training loss corresponding to each sample is determined. Based on the sample labels of each sample, at least one triplet is determined, and each triplet includes a valid speech sample, a first reverberation sample and a second reverberation sample. For each triplet, a first feature distance is determined between the audio features of the valid speech sample in the triplet and the audio features of the first reverberation sample; and a second feature distance is determined between the audio features of the valid speech sample in the triplet and the audio features of the second reverberation sample. Based on the first feature distance and the second feature distance corresponding to each triplet, the second training loss corresponding to each triplet is determined. The total training loss is obtained based on the first training loss corresponding to each sample and the second training loss corresponding to each triplet.
[0130] In one possible implementation, the sample label of the noise sample also characterizes ambient noise in the invalid speech; the sample label of the first reverberant sample also characterizes background reverberant speech in the invalid speech. The sample detection result includes at least a first result indication information, which indicates whether the sample belongs to valid speech; If the first result indication information indicates that the sample does not belong to valid speech, then the sample detection result also includes invalid speech reference information, which indicates that the sample belongs to background reverberant speech or environmental noise in invalid speech.
[0131] In one possible implementation, the sample acquisition module, upon acquiring at least one valid speech sample, is used to: Obtain at least one valid speech data from a pre-configured valid speech library, and obtain at least one noise data from a pre-configured noise library; According to at least one signal-to-noise ratio, the acquired effective speech data and noise data are synthesized to obtain at least one effective speech sample.
[0132] The model training method provided in this application obtains a first sample set of ineffective speech and a second sample set of effective speech. The first sample set includes a first reverberation sample and a noise sample. The effective speech sample includes effective speech data from human voice sources. The first reverberation sample is obtained by propagating the effective speech data in sample rooms with different reverberation intensities. The noise sample includes environmental noise data from non-human voice sources. Based on this, the speech of background human voice sources can be simulated using the first reverberation sample. The first and second sample sets are used for training, so that the trained speech detection model can not only accurately identify effective human voices and ineffective environmental noise, but also effectively identify background noisy human voices that belong to ineffective speech, thereby improving the accuracy and practicality of speech detection.
[0133] Furthermore, the trained speech detection model can effectively assist in speech recognition or voiceprint recognition tasks. For speech recognition tasks, the speech detection model trained using the method described in this application can further improve the robustness of the speech detection process, accurately filtering out background noise in addition to effectively filtering out noise. For voiceprint recognition tasks, the speech detection model described in this application, in addition to filtering out noise, can also finely distinguish between valid human voices and background noise, thereby improving the accuracy of voiceprint recognition. This further enhances the practicality of the model training method and the speech detection method.
[0134] Figure 9 This is a schematic diagram of a voice detection device provided in an embodiment of this application. Figure 9 As shown, the device includes: The audio signal acquisition module 901 is used to acquire at least one frame of audio signal from the target room where the target object is located. The second detection module 902 is used to detect each frame of audio signal based on the trained speech detection model, and obtain the detection result of each frame of audio signal. The detection result indicates whether the corresponding frame of audio signal belongs to valid speech. The speech detection model was trained using the aforementioned model training method.
[0135] In one possible implementation, the detection result includes audio indication information and target reference information, wherein the audio indication information indicates whether the corresponding frame audio signal belongs to valid speech, and the target reference information indicates whether the corresponding frame audio signal belongs to background reverberant speech or environmental noise in invalid speech; the device further includes any one of the following: The cut-off module is used to cut off the first audio signal corresponding to the background reverberant speech or environmental noise that belongs to non-effective speech from each frame of audio signal based on the audio indication information of each frame of audio signal; The recognition module is used to determine the second audio signal that belongs to valid speech and the third audio signal that belongs to background reverberation speech in each frame of audio signal based on the audio indication information and target reference information of each frame of audio signal respectively; and to perform voiceprint recognition on each frame of audio signal based on the audio features of the second audio signal and the third audio signal to obtain the audio signals from the target object and other objects outside the target object in each frame of audio signal respectively.
[0136] The method provided in this application, through the speech detection model trained using the aforementioned model training method, can effectively assist in speech recognition or voiceprint recognition tasks. For speech recognition tasks, the speech detection model trained using the method of this application can further improve the robustness of the speech detection process, accurately filtering out background noise in addition to effectively filtering out noise. For voiceprint recognition tasks, the speech detection model of this application, in addition to filtering out noise, can also finely distinguish between valid human voices and background noise, thereby improving the accuracy of voiceprint recognition. This further enhances the practicality of the model training method and the speech detection method.
[0137] The apparatus in this application embodiment can execute the method provided in this application embodiment, and the implementation principle is similar. The actions performed by each module in the apparatus of each embodiment of this application correspond to the steps in the method of each embodiment of this application. For detailed functional descriptions of each module of the apparatus, please refer to the descriptions in the corresponding methods shown above, which will not be repeated here.
[0138] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. For example... Figure 10 As shown, the electronic device includes: a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the model training method and the speech detection method described above.
[0139] In one alternative embodiment, an electronic device is provided, such as Figure 10 As shown, Figure 10The illustrated electronic device 1000 includes a processor 1001 and a memory 1003. The processor 1001 and the memory 1003 are connected, for example, via a bus 1002. Optionally, the electronic device 1000 may further include a transceiver 1004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 1004 is not limited to one type, and the structure of the electronic device 1000 does not constitute a limitation on the embodiments of this application.
[0140] Processor 1001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 1001 may also be a combination that implements computing functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0141] Bus 1002 may include a pathway for transmitting information between the aforementioned components. Bus 1002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 1002 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 10 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0142] The memory 1003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), disk storage media, other magnetic storage devices, or any other medium capable of carrying or storing computer programs and capable of being read by a computer, without limitation herein.
[0143] The memory 1003 is used to store computer programs that execute the embodiments of this application, and the execution is controlled by the processor 1001. The processor 1001 is used to execute the computer programs stored in the memory 1003 to implement the steps shown in the foregoing method embodiments.
[0144] Electronic devices include, but are not limited to, servers, terminals, or cloud computing center equipment.
[0145] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement the steps and corresponding content of the aforementioned method embodiments.
[0146] This application also provides a computer program product, including a computer program that, when executed by a processor, can implement the steps and corresponding content of the aforementioned method embodiments.
[0147] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0148] Those skilled in the art will understand that, unless otherwise stated, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. The terms “comprising” and “including” as used in the embodiments of this application mean that the corresponding feature can be implemented as the presented feature, information, data, step, or operation, but do not exclude implementation as other features, information, data, steps, or operations supported by this art.
[0149] The terms "first," "second," "third," "fourth," "1," "2," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in a sequence other than that shown in the illustrations or text descriptions.
[0150] It should be understood that although arrows indicate various operation steps in the flowcharts of this application's embodiments, the order in which these steps are implemented is not limited to the order indicated by the arrows. Unless explicitly stated herein, in some implementation scenarios of this application's embodiments, the implementation steps in each flowchart can be executed in other orders as required. Furthermore, some or all steps in each flowchart, based on the actual implementation scenario, may include multiple sub-steps or multiple stages. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage can also be executed at different times. In scenarios where execution times differ, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and this application's embodiments do not limit this.
[0151] The above description is only an optional implementation method for some implementation scenarios of this application. It should be noted that for those skilled in the art, other similar implementation methods based on the technical concept of this application without departing from the technical concept of this application also fall within the protection scope of the embodiments of this application.
Claims
1. A method for training a speech detection model, characterized in that, The method includes: A first sample set and a second sample set are obtained, and the sample labels corresponding to the first sample set and the second sample set respectively represent invalid speech and valid speech; the first sample set includes at least one first reverberation sample and a noise sample, and the second sample set includes at least one valid speech sample. The effective speech samples include effective speech data from human voice sources; the first reverberation sample is obtained by propagating the effective speech data in sample rooms with different reverberation intensities; and the noise samples include environmental noise data from non-human voice sources. Based on the audio features of each sample in the first and second sample sets, an initial speech detection model is used to perform speech detection on each sample to obtain the sample detection results of each sample. The sample detection results represent whether the corresponding sample belongs to valid speech. Based on the sample detection results and sample labels of each sample, the total training loss corresponding to each sample is determined, and the initial speech detection model is iteratively trained based on the total training loss corresponding to each sample to obtain the speech detection model.
2. The method according to claim 1, characterized in that, The second sample set also includes at least one second reverberation sample, wherein the determination of the at least one first reverberation sample and the second reverberation sample includes at least one of the following: The effective speech data is subjected to reverberation processing based on at least one first impulse response to obtain at least one first reverberation sample; and the effective speech data is subjected to reverberation processing based on at least one second impulse response to obtain at least one second reverberation sample. Wherein, the first impulse response or the second impulse response characterizes the propagation characteristics of effective speech in the corresponding sample room, and the reverberation intensity corresponding to the first impulse response is greater than that of the second impulse response; Based on the target human voice source corresponding to each valid speech sample in the second sample set, at least one valid speech data from the target human voice source is subjected to reverberation processing to obtain at least one second reverberation sample, and at least one valid speech data from a non-target human voice source is subjected to reverberation processing to obtain at least one first reverberation sample.
3. The method according to claim 2, characterized in that, The reverberation processing of at least one valid speech data based on the first impulse response to obtain at least one first reverberation sample includes: Convolution processing is performed on each first impulse response and effective speech data to generate initial reverberation data corresponding to each first impulse response; Noise data is obtained from a pre-configured noise library, and the initial reverberation data corresponding to each first impulse response and the obtained noise data are synthesized according to at least one signal-to-noise ratio to obtain the at least one first reverberation sample.
4. The method according to claim 2, characterized in that, The method for determining the at least one first reverberation sample and the second reverberation sample further includes at least one of the following: Based on the reverberation intensities corresponding to the first and second reverberation samples respectively, the sample labels of the first and second reverberation samples are determined. Based on the signal-to-noise ratios of the first and second reverberation samples respectively, the sample labels of the first and second reverberation samples are determined.
5. The method according to claim 4, characterized in that, The determination of the sample label of the first reverberation sample and the sample label of the second reverberation sample based on the reverberation intensity corresponding to the first reverberation sample and the second reverberation sample, respectively, includes at least one of the following: The first category label representing non-effective speech is used as the sample label for the first reverberation sample, and the second category label representing effective speech is used as the sample label for the second reverberation sample. Based on the reverberation intensity corresponding to the first reverberation sample and the second reverberation sample respectively, the effective speech confidence scores corresponding to the first reverberation sample and the second reverberation sample are obtained respectively, and the effective speech confidence scores corresponding to the first reverberation sample and the second reverberation sample are used as the sample labels of the first reverberation sample and the second reverberation sample respectively. The reverberation intensity is negatively correlated with the effective speech confidence score.
6. The method according to any one of claims 1-5, characterized in that, The second sample set also includes at least one second reverberation sample; The determination of the total training loss for each sample based on the sample detection results and sample labels includes: Based on the similarity between the sample detection results and sample labels of each sample, the first training loss corresponding to each sample is determined; Based on the sample labels of each sample, at least one triplet is determined, and each triplet includes a valid speech sample, a first reverberation sample and a second reverberation sample. For each triplet, a first feature distance is determined between the audio features of the valid speech sample in the triplet and the audio features of the first reverberation sample; and a second feature distance is determined between the audio features of the valid speech sample in the triplet and the audio features of the second reverberation sample. Based on the first feature distance and the second feature distance corresponding to each triplet, the second training loss corresponding to each triplet is determined; The total training loss is obtained based on the first training loss corresponding to each sample and the second training loss corresponding to each triplet.
7. The method according to claim 6, characterized in that, The sample label of the noise sample also represents the ambient noise in the invalid speech; the sample label of the first reverberation sample also represents the background reverberation speech in the invalid speech. The sample detection result includes at least a first result indication information, which indicates whether the sample belongs to valid speech. If the first result indication information indicates that the sample does not belong to valid speech, then the sample detection result also includes invalid speech reference information, which indicates that the sample belongs to background reverberant speech or environmental noise in invalid speech.
8. The method according to any one of claims 1-5, characterized in that, The at least one valid speech sample is obtained by performing the following steps: Obtain at least one valid speech data from a pre-configured valid speech library, and obtain at least one noise data from a pre-configured noise library; According to at least one signal-to-noise ratio, the acquired effective speech data and noise data are synthesized to obtain at least one effective speech sample.
9. A speech detection method, characterized in that, The method includes: Acquire at least one frame of audio signal from the target room where the target object is located; The trained speech detection model is used to detect each frame of audio signal to obtain the detection result of each frame of audio signal. The detection result indicates whether the corresponding frame of audio signal belongs to valid speech. The speech detection model is trained using the model training method described in any one of claims 1-8.
10. The method according to claim 9, characterized in that, The detection result includes audio indication information and target reference information. The audio indication information indicates whether the corresponding frame audio signal belongs to valid speech, and the target reference information indicates whether the corresponding frame audio signal belongs to background reverberant speech or environmental noise in invalid speech. The method further includes any one of the following: Based on the audio indication information of each frame of audio signal, the first audio signal corresponding to the background reverberation speech or environmental noise that belongs to non-effective speech is removed from each frame of audio signal; Based on the audio indication information and target reference information of each frame of audio signal, the second audio signal belonging to valid speech and the third audio signal belonging to background reverberant speech in each frame of audio signal are determined respectively; and based on the audio features of the second audio signal and the third audio signal, voiceprint recognition is performed on each frame of audio signal to obtain the audio signals from the target object and other objects outside the target object in each frame of audio signal respectively.
11. A model training device, characterized in that, The device includes: The sample acquisition module is used to acquire a first sample set and a second sample set, wherein the sample labels corresponding to the first sample set and the second sample set respectively represent invalid speech and valid speech; the first sample set includes at least one first reverberation sample and a noise sample, and the second sample set includes at least one valid speech sample. The effective speech samples include effective speech data from human voice sources; the first reverberation sample is obtained by propagating the effective speech data in sample rooms with different reverberation intensities; and the noise samples include environmental noise data from non-human voice sources. The first detection module is used to perform speech detection on each sample based on the sample audio features of each sample in the first sample set and the second sample set, using an initial speech detection model to obtain the sample detection result of each sample. The sample detection result represents whether the corresponding sample belongs to valid speech. The total loss determination module is used to determine the total training loss for each sample based on the sample detection results and sample labels of each sample. The training module is used to iteratively train the initial speech detection model based on the total training loss corresponding to each sample to obtain the speech detection model.
12. A voice detection device, characterized in that, The device includes: The audio signal acquisition module is used to acquire at least one frame of audio signal from the target room where the target object is located. The second detection module is used to detect each frame of audio signal based on the trained speech detection model, and obtain the detection result of each frame of audio signal. The detection result indicates whether the corresponding frame of audio signal belongs to valid speech. The speech detection model is trained using the model training method described in any one of claims 1-8.
13. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the method of any one of claims 1 to 10.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 10.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 10.