Method and apparatus for voice detection, and device, medium and program product

By acquiring target speech features and utilizing a personalized speech activity detection model, the problem of identifying the speech of a specific target user in noisy environments using traditional speech activity detection is solved. This enables accurate identification of the speech information of the target object in noisy environments, thus improving the user experience.

WO2026007737A1PCT designated stage Publication Date: 2026-01-08BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/102808
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-03
Filing Date
2025-06-23
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Traditional speech activity detection technology struggles to accurately identify the speech of a specific target user in noisy environments, resulting in the inability to perform speech recognition correctly.

Method used

By acquiring the target speech features of the audio to be detected and the target object, a personalized speech activity detection model is used to determine the audio segments in the audio to be detected that correspond to the target object based on the audio features and the target speech features.

Benefits of technology

It enables accurate recognition of target speech information in noisy environments, improving user experience and increasing the accuracy of speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025102808_08012026_PF_FP_ABST
    Figure CN2025102808_08012026_PF_FP_ABST
Patent Text Reader

Abstract

A method and apparatus for voice detection, and a device, a medium and a program product. The method comprises: acquiring audio to be subjected to detection and a target voice feature of a target object, wherein said audio comprises voice information of the target object (302); on the basis of said audio, determining an audio feature of said audio (304); and on the basis of the audio feature and the target voice feature, determining from said audio an audio clip corresponding to the target object (306).
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, equipment, medium and program product for detecting speech

[0001] Cross-reference to related applications

[0002] The present application claims priority to the Chinese patent application No. 202410889401.0, filed on July 3, 2024, and entitled “Method, device, equipment, medium and program product for detecting speech”, the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0003] Embodiments of the present disclosure generally relate to the field of audio processing, and in particular, to a method, device, equipment, medium and program product for detecting speech. BACKGROUND

[0004] At present, machine learning is becoming more and more important in people's daily life, and gradually becomes an important tool that people rely on. More and more work begins to use machine learning models to process. For example, word processing work, picture processing work, audio processing work, etc. begin to use machine learning models to process. The development of machine learning technology brings great improvement to the efficiency of multi-modal data processing.

[0005] With the rapid development of machine learning technology, when processing various multi-modal data, the processing process has become faster and more accurate. For example, when performing audio processing work, a pre-deployed audio processing model can be used to assist users in performing audio processing work, such as automatic tuning, intercepting audio segments, etc. In addition, in order to meet the development needs of speech recognition technology, machine learning models are also increasingly used in speech detection. SUMMARY

[0006] Embodiments of the present disclosure provide a method, device, equipment, medium and program product for detecting speech.

[0007] According to a first aspect of the present disclosure, a method for detecting speech is provided. The method includes obtaining a to-be-detected audio and a target speech feature of a target object, the to-be-detected audio including speech information of the target object. The method further includes determining an audio feature of the to-be-detected audio based on the to-be-detected audio. The method further includes determining an audio segment corresponding to the target object in the to-be-detected audio based on the audio feature and the target speech feature.

[0008] In a second aspect of the present disclosure, an apparatus for detecting speech is provided. The apparatus includes a to-be-detected audio and target speech feature acquisition module configured to acquire to-be-detected audio and a target speech feature of a target object, the to-be-detected audio including speech information of the target object; an audio feature determination module configured to determine an audio feature of the to-be-detected audio based on the to-be-detected audio; and an audio segment determination module configured to determine an audio segment corresponding to the target object in the to-be-detected audio based on the audio feature and the target speech feature.

[0009] In a third aspect of the present disclosure, an electronic device is provided, including at least one processor; and a storage device for storing at least one program, when the at least one program is executed by the at least one processor, causing the at least one processor to implement the method according to the first aspect of the present disclosure.

[0010] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, having stored thereon a computer program, which, when executed by a processor, implements the method according to the first aspect of the present disclosure.

[0011] In a fifth aspect of the present disclosure, a computer program product is provided. The computer program product includes a computer program which, when executed by a processor, implements the method according to the first aspect of the present disclosure.

[0012] It should be understood that the content described in this section is not intended to limit the key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS

[0013] The above and other objects, features and advantages of the present disclosure will become more apparent from the following description when taken in conjunction with the accompanying drawings, in which like reference characters designate the same components in several views.

[0014] FIG. 1 illustrates a schematic diagram of an example environment in which devices and / or methods of embodiments of the present disclosure can be implemented;

[0015] FIG. 2 illustrates a schematic diagram of an example for speech audio according to embodiments of the present disclosure;

[0016] FIG. 3 illustrates a schematic diagram of an example method for detecting speech according to embodiments of the present disclosure;

[0017] FIG. 4 illustrates a schematic diagram of a model architecture of a personalized voice activity detection model for detecting speech according to embodiments of the present disclosure;

[0018] FIG. 5 illustrates a schematic diagram of a 6-layer MarbleNet encoder in a personalized voice activity detection model for detecting speech, according to an embodiment of the present disclosure

[0019] FIG. 6 illustrates a schematic block diagram of an apparatus for detecting speech, according to an embodiment of the present disclosure;

[0020] FIG. 7 illustrates a schematic block diagram of an example device suitable for implementing embodiments of the present disclosure.

[0021] In the various drawings, like or corresponding elements are denoted by like or corresponding reference numerals. DETAILED DESCRIPTION

[0022] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and relevant provisions. In response to receiving the active request of the user, the user is sent a prompt message to explicitly prompt the user that the operation requested to be performed will require the acquisition and use of the user's personal information. Thus, the user can voluntarily choose whether to provide personal information to the software or hardware such as electronic devices, application programs, servers or storage media that perform the operation of the technical solution of the present disclosure according to the prompt message.

[0023] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein, but rather these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for exemplary purposes only, and are not intended to limit the scope of protection of the present disclosure.

[0024] In the description of embodiments of the present disclosure, the term "comprising" and its conjugations should be understood as open-ended, i.e., "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "an embodiment" should be understood as "at least one embodiment". The terms "first", "second" and the like can refer to different or identical objects. Other explicit and implicit definitions can also be included below.

[0025] There are still many problems to be solved in the process of detecting speech. For example, taking the user interacting with the computing device as an example, the user can issue a wake-up instruction to the computing device, and then have a conversation with the computing device. The user can issue various instructions such as asking and talking to the computing device according to his own needs, and the computing device can then answer or reply to the content of the instruction issued by the user. For example, after the user wakes up the computing device, the user asks the computing device "do you know someone?", and then the computing device gives a detailed reply to the question raised by the user.

[0026] In the traditional scheme, in the scenario of the user interacting with the computing device, the computing device or the server at the back end relies on the voice activity detection (VAD) technology to identify whether there is a user speaking in the environment. These operations are usually used for automatic speech recognition (ASR) pre-processing. When it is detected that there is someone speaking in the current environment, the corresponding human voice is determined and the remaining non-speech elements (such as non-human voice noise) are distinguished, and the audio segment in which the human voice exists is further determined. However, in this traditional scheme, only the human voice and the non-human voice noise can be distinguished. However, when the surrounding environment is too noisy or there are too many interference factors, this scheme cannot respond to the speech of a specific target user. For example, when the surrounding environment contains too much human voice noise and human voice interference, such as in a noisy environment such as a coffee shop, a shopping mall, or when other colleagues are having a meeting in the office, the speech of other irrelevant people will also be recognized and processed, and it is impossible to process only the speech of the target object, resulting in the failure to correctly complete the speech recognition of the target object.

[0027] At least to solve the above and other potential problems, embodiments of the present disclosure propose a method for detecting speech. In this method, the computing device can first obtain the to-be-detected audio and the target speech feature of the target object at the computing device. The to-be-detected audio includes speech information of the target object. Then, the computing device processes the to-be-detected audio to determine the audio feature of the to-be-detected audio. Next, the computing device uses the determined audio feature of the to-be-detected audio and the target speech feature to identify the audio segment corresponding to the target object from the to-be-detected audio. Through this method, since the target speech feature of the target object is added in the speech detection, the speech information of the target object can be distinguished from other human voice interference, non-human voice noise, etc., so that the audio segment of the target object can be accurately determined, and the user experience is improved.

[0028] Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. FIG. 1 shows an example environment in which an apparatus and / or method of embodiments of the present disclosure can be implemented. In the environment 100, a computing device 112 first obtains a to-be-detected audio 102 and a target speech feature 108. The to-be-detected audio 102 contains speech information 104 of a target object 110. Then, the computing device 112 determines an audio feature 106 for the to-be-detected audio 102 based on the to-be-detected audio 102. After determining the audio feature 106, the computing device 112 can determine a specific audio segment in the to-be-detected audio 102, i.e., an audio segment 114 corresponding to the target object 110, based on the determined audio feature 106 and the target speech feature 108.

[0029] Examples of the computing device 112 include, but are not limited to, a personal computer, a server computer, a handheld or laptop device, a mobile device (such as a mobile phone, a personal digital assistant (PDA), a media player, etc.), a multi-processor system, a consumer electronic product, a minicomputer, a mainframe computer, a distributed computing environment including any of the above systems or devices, etc.

[0030] As shown in FIG. 1, the computing device 112 can obtain the to-be-detected audio 102 and the target speech feature 108 related to the target object 110. In one example, the target speech feature 108 is obtained directly by the computing device 112 or received from other computing devices. In another example, the target speech feature 108 is calculated by the computing device 112. The computing device 112 can determine the target speech feature 108 of the target object 110 by using speech data corresponding to the target object 110.

[0031] In this case, the computing device 112 can perform feature extraction on the speech data to determine the target speech feature 108 corresponding to the target object 110. For example, by pre-processing and acoustic feature extraction (such as mel-frequency cepstral coefficients, etc.) on the speech data, and performing feature selection or optimization, or by a pre-trained model, the target speech feature corresponding to the target object 110 can be determined.

[0032] The computing device 112 can also use the to-be-detected audio 102 to further determine the audio feature 106 for the to-be-detected audio 102. The audio feature 106 is used to represent the to-be-detected audio 102, which includes audio features corresponding to the speech of the target object and audio features corresponding to the speech of other objects or non-human noise. For example, the to-be-detected audio 102 is an audio segment with a length of 1 second, and the computing device 112 can divide the audio segment into a plurality of sub-audios according to a predetermined time length, for example, divide the 1-second audio segment into 100 sub-audios according to 10 ms per sub-audio.

[0033] In some embodiments, the computing device 112 can process each sub-audio in the audio to be detected 102 to determine a spectrogram for each sub-audio. For example, by performing windowing processing on each sub-audio, and then performing Fourier transform, etc. to generate a corresponding spectrogram. The computing device can further determine an audio feature corresponding to each sub-audio according to the spectrogram. For example, by a peak frequency detection method or by a harmonic structure analysis method, etc. to extract the corresponding audio feature from the spectrogram.

[0034] Finally, the computing device 112 will determine the audio segment 114 corresponding to the target object 110 in the audio to be detected 102 according to the determined audio feature 106 and the target voice feature 108. In some embodiments, the audio feature 106 and the target voice feature 108 can be processed by a personalized voice activity detection model to determine the audio segment 114 corresponding to the target object 110 in the audio to be detected 102. Additionally, the personalized voice activity detection model can be enabled or stopped. When the personalized voice activity detection model is stopped, the computing device 112 will determine the audio segments corresponding to all objects in the audio to be detected respectively. Only when the personalized voice activity detection model is enabled, the computing device 112 will only determine the audio segment 114 corresponding to the target object 110.

[0035] The determination of the audio segment of a target object in FIG. 1 is only an example and is not a specific limitation of the present disclosure. The computing device can process the audio segment including multiple target objects. At this time, the computing device can obtain or determine multiple target voice features corresponding to multiple target objects respectively, so as to determine multiple audio segments corresponding to multiple target objects. Additionally, the process of identifying the audio segment corresponding to the target object from the audio to be detected described above is implemented by an application installed in the computing device 112.

[0036] By this method, since the target voice feature of the target object is added in the voice detection, the voice information of the target object can be distinguished from other human voice interference, non-human voice noise, etc., so that the audio segment of the target object can be accurately determined, and the user experience is improved.

[0037] The above describes the schematic diagram of an example environment in which the device and / or method of the embodiments of the present disclosure can be implemented in conjunction with FIG. 1, and the following describes the schematic diagram of an example for detecting voice according to the embodiments of the present disclosure in conjunction with FIG. 2.

[0038] There can be multiple speech segments within a continuous audio segment, and the human voices in the multiple speech segments can come from multiple different objects respectively. As shown in FIG. 2, in example 200, there are three speech segments in the audio segment from 0 second to 1.4 second, which are speech segment 202, speech segment 204, and speech segment 206 respectively. In this example, the three speech segments respectively contain a speech segment of a friend of the target object talking to the target object, a speech segment of the target object talking to the device, and a speech segment of an advertisement on the TV.

[0039] If the computing device 112 does not enable or use the personalized voice activity detection, all the three pieces of human voice speech in the audio segment from 0 second to 1.4 second will be recognized as valid human voice speech segments and received by the computing device 112. If the computing device 112 enables or uses the personalized voice activity detection, the speech segment of the target object in the three pieces of human voice speech in the audio segment from 0 second to 1.4 second is recognized as a valid human voice speech segment, and the computing device 112 intercepts the audio segment from the time when the human voice of the target object first appears in the audio segment to the time when the human voice of the target object last appears as a valid human voice audio segment and receives the speech information in the valid human voice audio segment.

[0040] Additionally, when there are more than one target object, the human voice features corresponding to the target objects can have priorities. For example, in a scenario of interacting with the computing device or an application installed in the computing device, there are two target objects, the target object and a friend of the target object, and when the personalized voice activity detection is enabled, the conversations of the target object and the friend of the target object with the computing device or the application can be interrupted by each other because the computing device can receive the target voice features of both the target object and the friend of the target object. However, when priorities are set, for example, when the priority of the target voice feature of the target object is set to be higher than the priority of the target voice feature of the friend of the target object, the target object can interrupt the interaction of the friend of the target object with the computing device or the application and the interaction of the target object with the computing device or the application at any time, while the friend of the target object cannot interrupt the interaction of the target object with the computing device or the application. It can be understood that the user can set different priorities for the target voice features corresponding to multiple target objects respectively according to the user's own needs.

[0041] The above describes an example of detecting speech according to an embodiment of the present disclosure in combination with FIG. 2. The following describes an example method 300 of detecting speech according to an embodiment of the present disclosure in combination with FIG. 3. The method can be applied in the example environment in FIG. 1 or any suitable environment, and can be executed by the computing device 112, an application installed in the computing device 112, or any suitable computing device.

[0042] As shown in FIG. 3, in the example method 300, at block 302, the computing device 112 obtains the audio to be detected 102 and the target voice feature 108 for the target object 110, the audio to be detected 102 including voice information 104 of the target object 110.

[0043] The audio to be detected 102 can be obtained directly by the computing device 112. For example, when the computing device is a terminal device, the audio to be detected 102 can be voice data collected by the computing device 112 through a connected microphone. For example, a user can open the microphone through an application implementing the method of the present disclosure in a mobile phone, and then the user has a conversation with the application in the mobile phone. The audio to be detected 102 can also be obtained by the computing device 112 from other devices. For example, when the computing device 112 is a server, the audio to be detected 102 can be audio received from other devices. For example, the audio to be detected 102 can be data received by the server from a client. The above examples are only used to describe the present disclosure, and are not specific limitations of the present disclosure.

[0044] In addition, the computing device 112 can directly obtain the target voice feature 108 for the target object 110 from other devices. The target voice feature 108 reflects the voice characteristics of the target user. For example, the target voice feature 108 can include the timbre, loudness and other features of the sound of the target object 110. The computing device 112 can also determine the target voice feature associated with the target object by pre-processing the voice information related to the target object 110. For example, the computing device 112 can frame, window and Fourier transform the voice information to extract voice features corresponding to the voice of the target user.

[0045] Subsequently, at block 304, the computing device 112 determines the audio feature 106 for the audio to be detected 102 based on the audio to be detected 102. After the computing device 112 obtains the audio to be detected 102, the audio to be detected 102 is processed to obtain spectrogram information corresponding to the audio to be detected, so as to further extract the corresponding audio feature 106.

[0046] In some embodiments, the computing device 112 can divide the audio to be detected 102 into a plurality of sub-audios. For example, when the audio to be detected 102 is a piece of audio with a time length of 1 second, the computing device 112 can divide it into a plurality of sub-audios according to a predetermined time length. When the computing device 112 divides the audio to be detected using a time length of 10 milliseconds, the audio to be detected 102 with a time length of 1 second will be divided into 100 sub-audios. Then, the computing device further performs windowing and Fourier transform operations to further determine a spectrogram corresponding to each sub-audio, and then the computing device 112 further determines audio features for the plurality of sub-audios using the obtained spectrograms. At this time, the audio features 106 corresponding to the audio to be detected 102 contain features corresponding to the entire audio to be detected 102, including audio features corresponding to the target object 110 and audio features corresponding to other objects.

[0047] At block 306, the computing device 112 determines the audio segment 114 corresponding to the target object 110 in the audio to be detected 102 based on the audio features 106 and the target speech features 108. The computing device 112 uses the target speech features 108 reflecting the speech features of the target object to detect the audio segment 114 corresponding to the target object 110 from the audio to be detected 102.

[0048] In some embodiments, the computing device 112 performs compression processing on the audio features 106 corresponding to the audio to be detected 102. For example, the computing device 112 compresses the audio features in a high-dimensional space to a low-dimensional space. For example, the audio features are 512-dimensional, and after compression processing, they become 256-dimensional. In one example, after obtaining the compressed audio features, the computing device 112 further combines the compressed audio features with the target speech features to generate combined features.

[0049] In the combination, the computing device 112 adds the compressed audio features and the target speech features in the time dimension. For example, the audio to be detected with a time length of 1 second is divided into a plurality of sub-audios according to a time interval of 10 milliseconds and output after compression by an encoder, resulting in a feature vector of T=100 and 256 elements for each time dimension. Since the target speech features do not have a time dimension, when the audio features and the target speech features are spliced and combined in the time dimension, the target speech features need to be spliced with the corresponding feature vector at each time dimension. For example, the target speech features contain 192 elements, so after splicing and combining the audio features and the target speech features, the feature vector of the combined features is 100*(256+192). Finally, the computing device 112 determines the audio segment for the target object in the audio to be detected using the spliced and combined combined features.

[0050] In some embodiments, the computing device utilizes a personalized voice activity detection model to process the audio features 106 and the target voice features 108 in determining the audio segment 114. The personalized voice activity detection model can include an encoder and a gated recurrent unit, where the encoder is used to compress the audio features 106 into a low-dimensional space, and then the gated recurrent unit processes the combined features to determine the audio segment for the target object in the audio to be detected.

[0051] In some embodiments, the personalized voice activity detection model can output a prediction result, for example, one prediction result for each sub-audio. If the prediction result is greater than or equal to a threshold value, for example, 0.6, it is determined that the sub-audio is for the target user, and the sub-audio is labeled as 1, and if the prediction result is less than the threshold value, it is determined that the sub-audio does not belong to the target user, and the sub-audio is labeled as 0. Therefore, a plurality of sub-audios labeled as 1 in succession can be determined as the audio segment 114 for the target object. Additionally, post-processing can also be performed on the prediction result, if there are less than a predetermined number of sub-audios labeled as 0 between two sub-audios labeled as 1, the labels of the sub-audios of the predetermined number can also be adjusted to 1, which avoids dividing the speech of the target user into two audio segments due to the pause of the target user speaking.

[0052] In some embodiments, the personalized voice activity detection model can be set to be turned on or off according to the user's needs. For example, in a relatively quiet environment with only a few people, the user can set the personalized voice activity detection to be turned off to save computing resources and save the power consumption of the corresponding computing device.

[0053] As described above, the encoder and the gated recurrent unit are components of the personalized voice activity detection model. In one example, the encoder can be a 6-layer MarbleNet encoder, which can be suitable for devices using fewer parameters. The gated recurrent unit is a 3-layer gated recurrent unit.

[0054] It can be understood that the user can select a suitable encoder and a suitable recurrent neural network according to his own needs instead of the above-mentioned 6-layer MarbleNet encoder and 3-layer gated recurrent unit. For example, the user can use a UNET encoder instead of the 6-layer MarbleNet encoder, and use a Long Short Time Memory (LSTM) or a Temporal Convolutional Networks (TCN) in a convolutional neural network instead of the gated recurrent unit. The present application does not limit this.

[0055] In addition, since the gated recurrent unit is a lightweight neural network, i.e., a neural network with a small number of parameters, it can be deployed in a mobile terminal. If the model has a large number of parameters, it needs to be deployed on a server for use.

[0056] In some embodiments, when the personalized voice activity detection model does not detect voice information of the target object within a certain time, the personalized voice activity detection function is automatically turned off, thereby saving computing resources. For example, if no voice information of the target object is detected in the current environment for 2 consecutive minutes, the personalized voice activity detection function is turned off, and the personalized voice activity detection function is turned on again only when voice information of the target object is detected.

[0057] In some embodiments, before using the personalized voice activity detection model, the personalized voice activity detection model needs to be trained. During training, the computing device 112 can obtain sample audio, target sample voice features of a sample object, and sample audio segments corresponding to the sample object in the sample audio, and apply these data to the personalized voice activity detection model for training. Additionally, the user can also train multiple target sample voice features corresponding to multiple sample objects and set priorities for the multiple target sample voice features, so that the personalized voice activity detection model can prioritize the audio segments corresponding to different priorities in the presence of multiple objects.

[0058] It can be understood that the sample data involved in the present disclosure (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws and regulations and relevant provisions.

[0059] Through the method, since the target voice features of the target object are added in the voice detection, the voice information of the target object can be distinguished from other human voice interference and non-human voice noise, so that the audio segment of the target object can be accurately determined, and the user experience is improved.

[0060] The above describes an example method 300 for detecting voice according to an embodiment of the present disclosure in conjunction with FIG. 3. The following describes a schematic diagram of a model architecture of a personalized voice activity detection model for detecting voice according to an embodiment of the present disclosure in conjunction with FIG. 4.

[0061] As shown in an example 400 of FIG. 4, the personalized voice detection model is composed of a Sigmoid activation function 402, an output layer 404, 3 layers of gated recurrent units 406, and 6 layers of MarbleNet encoders 408.

[0062] After obtaining the audio to be detected 412, the computing device 112 first pre-processes the audio to be detected, and divides the audio to be detected according to a predetermined time length, for example, according to a time interval of 10 milliseconds. If the time length of the audio to be detected is 1 second, it will be divided into 100 sub-audios. After obtaining the 100 sub-audios, the computing device 112 processes the 100 sub-audios to obtain a spectrogram 410 for the 100 sub-audios.

[0063] In the process of generating the spectrogram 410, the computing device 112 performs windowing on the divided sub-audios, and then performs short-time Fourier transform to obtain a spectrum corresponding to each frame of sub-audio. Then the computing device 112 performs rotation and mapping operations on the obtained spectrum of all sub-audios to obtain a plurality of transformed spectra. Finally, the plurality of transformed spectra are combined to form the spectrogram 410.

[0064] After the pre-processing of the audio to be detected is completed, the audio features of the audio to be detected can be obtained, and then the computing device 112 inputs the audio features into the 6-layer MarbleNet encoder 408. The encoder compresses the audio features from a high dimension to a low dimension.

[0065] For example, the audio to be detected with a time length of 1 second is divided into a plurality of sub-audios according to a time interval of 10 milliseconds and compressed by the encoder, and then outputted to obtain a time dimension T = 100 and each time dimension contains a feature of 256 elements. At this time, the feature size of the audio to be detected is 100*256.

[0066] In addition, since the target voice feature 414 does not have a time dimension, when the audio features and the target voice feature are spliced and combined in the time dimension, the target voice feature needs to be spliced with the corresponding audio feature in each time dimension. For example, the feature vector of the target voice feature contains 192 elements, so when the audio features and the target voice feature are spliced and combined, the feature size of the combined feature is 100*(256+192). That is, the above combined feature is obtained by linearly splicing and combining the audio features and the target voice feature in the time dimension.

[0067] Then, the computing device 112 inputs the linearly spliced and combined feature into the recurrent gate unit and further determines the audio segment of the target object in the audio to be detected. After inputting the combined feature into the 3-layer gated recurrent unit 406, the output layer 404 further obtains the result output by the 3-layer gated recurrent unit, and then applies the output result of the combined feature in the output layer to the Sigmoid activation function 402.

[0068] The sigmoid activation function 402 maps the output result to the interval (0, 1) after receiving the output result for the combined features. The sub-audio whose output result is greater than or equal to a threshold value can be marked as 1, and the sub-audio whose output result is less than the threshold value can be marked as 0. Therefore, the continuous sub-audio marked as 1 can be determined as the audio segment corresponding to the target object.

[0069] By this method, since the target voice feature for the target object is added in the voice detection, the voice information of the target object can be distinguished from other human voice interference, non-human voice noise, etc., so that the audio segment of the target object can be accurately determined, and the user experience is improved.

[0070] The above describes the schematic diagram of the model architecture of the personalized voice activity detection model for detecting voice according to an embodiment of the present disclosure in combination with FIG. 4. The schematic diagram of the 6-layer MarbleNet encoder in the personalized voice activity detection model for detecting voice according to an embodiment of the present disclosure is described below in combination with FIG. 5.

[0071] As shown in the example 500 of FIG. 5, the 6-layer MarbleNet encoder is composed of a random inactivation layer 502, a rectified linear unit activation function layer 504, a batch normalization layer 506, a one-dimensional deep convolution layer + pointwise convolution layer (K) 508, a batch normalization layer 510, and a 1x1 pointwise convolution layer 512, etc.

[0072] As can be seen, each layer in the 6-layer MarbleNet encoder contains the above four modules of the random inactivation layer 502, the rectified linear unit activation function layer 504, the batch normalization layer 506, and the one-dimensional deep convolution layer + pointwise convolution layer (K) 508.

[0073] In some embodiments, the above four modules are repeated only once in the 1st layer MarbleNet encoder. In other embodiments, the above four modules are repeated twice in the 2nd-6th layer MarbleNet encoder.

[0074] The above describes the schematic diagram of the 6-layer MarbleNet encoder in the personalized voice activity detection model for detecting voice according to an embodiment of the present disclosure in combination with FIG. 5. The schematic block diagram of the apparatus 600 for detecting voice according to an embodiment of the present disclosure is described below in combination with FIG. 6.

[0075] As shown in FIG. 6, the apparatus 600 includes an audio to be detected and target speech feature acquisition module 610 configured to acquire audio to be detected and a target speech feature for a target object, the audio to be detected including speech information of the target object; an audio feature determination module 620 configured to determine an audio feature for the audio to be detected based on the audio to be detected; and an audio segment determination module 630 configured to determine an audio segment corresponding to the target object in the audio to be detected based on the audio feature and the target speech feature.

[0076] In some embodiments, the audio segment determination module 630 further includes a compression module configured to compress the audio feature to obtain a compressed audio feature; and an audio segment determination module configured to determine the audio segment corresponding to the target object in the audio to be detected based on the compressed audio feature and the target speech feature.

[0077] In some embodiments, the audio segment determination module further includes a combination module configured to determine a combined feature by combining the compressed audio feature and the target speech feature; and an audio segment determination module configured to determine the audio segment having the speech of the target object in the audio to be detected based on the combined feature.

[0078] In some embodiments, the combination module further includes a concatenation module configured to generate the combined feature by concatenating the compressed audio feature and the target speech feature.

[0079] In some embodiments, the audio feature determination module 620 further includes a sub-audio division module configured to divide the audio to be detected into a plurality of sub-audios based on a predetermined time length; and an audio feature determination module configured to determine an audio feature corresponding to each of the plurality of sub-audios.

[0080] In some embodiments, the audio feature determination module further includes a spectrogram generation module configured to generate a spectrogram for a sub-audio by processing the sub-audio; and an audio feature generation module configured to generate the audio feature corresponding to the sub-audio based on the spectrogram.

[0081] In some embodiments, the audio segment determination module 630 further includes a personalized voice activity detection model application module configured to apply the audio feature and the target speech feature to a personalized voice activity detection model to determine the audio segment corresponding to the target object in the audio to be detected.

[0082] In some embodiments, the personalized voice activity detection model comprises an encoder and a multi-layered gated recurrent unit, wherein the personalized voice activity detection model application module further comprises: a compression module configured to compress the audio features by the encoder to generate compressed audio features; and a gated recurrent unit application module configured to determine the audio segment corresponding to the target object in the audio to be detected by applying the compressed audio features and the target voice features to the gated recurrent unit.

[0083] In some embodiments, the encoder is a 6-layer MarbleNet encoder, and the multi-layered gated recurrent unit is a 3-layer gated recurrent unit.

[0084] In some embodiments, the personalized voice activity detection model training module further comprises: a sample audio, target sample voice features and sample audio segment acquisition module configured to acquire sample audio, target sample voice features of a sample object and a sample audio segment corresponding to the sample object in the sample audio; and a personalized voice activity detection model training module configured to train the personalized voice activity detection model based on the sample audio, the target sample voice features and the sample audio segment.

[0085] In some embodiments, the apparatus 600 further comprises: a target voice feature determination module configured to determine the target voice features for the target object based on voice data related to the target object.

[0086] FIG. 7 shows a schematic block diagram of an example device 700 that can be used to implement embodiments of the present disclosure. The computing device 112 in FIG. 1 can be implemented with the device 700. As shown, the device 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 702 or loaded into a random access memory (RAM) 703 from a storage unit 708. Various programs and data required for operation of the device 700 can also be stored in the RAM 703. The CPU 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0087] A plurality of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0088] The various processes and processes described above, such as the method 300, can be performed by the processing unit 701. For example, in some embodiments, the method 300 can be implemented as a computer software program tangibly embodied in a machine readable medium, such as the storage unit 708. In some embodiments, portions or all of the computer program can be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded onto the RAM 703 and executed by the CPU 701, one or more acts of the method 300 described above can be performed.

[0089] The present disclosure can be a method, apparatus, system, and / or computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for performing various aspects of the present disclosure.

[0090] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or punched tape, a

[0091] The computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0092] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0093] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0094] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0095] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0096] The flow diagrams and the block diagrams in the drawings are presented to illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams and the block diagrams can represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logic functions. In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and

[0097] Embodiments of the present disclosure have been described above, and the description is intended to be illustrative of the embodiments and not restrictive. Many modifications and variations of the described embodiments are possible and are within the scope of the disclosure. The selection of terms is intended to best describe the principles of the embodiments, practical application, or technical improvements over the technology found in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1.A method for detecting speech, comprising: obtaining a target speech feature for a target object and a to-be-detected audio, the to-be-detected audio comprising speech information of the target object; determining an audio feature for the to-be-detected audio based on the to-be-detected audio; and determining an audio segment corresponding to the target object in the to-be-detected audio based on the audio feature and the target speech feature. 2.The method of claim 1, wherein determining the audio segment corresponding to the target object in the to-be-detected audio based on the audio feature and the target speech feature comprises: compressing the audio feature to obtain a compressed audio feature; and determining the audio segment corresponding to the target object in the to-be-detected audio based on the compressed audio feature and the target speech feature. 3.The method of claim 2, wherein determining the audio segment corresponding to the target object in the to-be-detected audio based on the compressed audio feature and the target speech feature comprises: determining a combined feature by combining the compressed audio feature and the target speech feature; and determining an audio segment having speech of the target object in the to-be-detected audio based on the combined feature. 4.The method of claim 3, wherein determining the combined feature by combining the compressed audio feature and the target speech feature comprises: generating the combined feature by concatenating the compressed audio feature and the target speech feature. 5.The method of claim 1, wherein determining the audio feature for the to-be-detected audio based on the to-be-detected audio comprises: dividing the to-be-detected audio into a plurality of sub-audios based on a predetermined time length; and determining an audio feature corresponding to each of the plurality of sub-audios. 6.The method of claim 5, wherein determining the audio feature corresponding to each of the plurality of sub-audios comprises: generating a spectrogram for the sub-audio by processing the sub-audio; and generating the audio feature corresponding to the sub-audio based on the spectrogram. 7.The method of claim 1, wherein determining the audio segment corresponding to the target object in the to-be-detected audio based on the audio feature and the target speech feature comprises: applying the audio feature and the target speech feature to a personalized voice activity detection model to determine the audio segment corresponding to the target object in the to-be-detected audio. 8.The method of claim 7, wherein the personalized voice activity detection model comprises an encoder and a multi-layer gated recurrent unit, and wherein applying the audio feature and the target speech feature to the personalized voice activity detection model to determine the audio segment corresponding to the target object in the to-be-detected audio comprises: compressing the audio feature by the encoder to generate a compressed audio feature; and ​ ​ ​ ​ ​ ​ The audio segment corresponding to the target object in the audio to be detected is determined by applying the compressed audio feature and the target voice feature to the gated recurrent unit. 9.The method of claim 8, wherein the encoder is a 6-layer MarbleNet encoder and the multi-layer gated recurrent unit is a 3-layer gated recurrent unit. 10.The method of claim 8, wherein the training of the personalized voice activity detection model comprises: obtaining a sample audio, a target sample voice feature of a sample object, and a sample audio segment corresponding to the sample object in the sample audio; and training the personalized voice activity detection model based on the sample audio, the target sample voice feature, and the sample audio segment. 11.The method of claim 1, further comprising: determining the target voice feature for the target object based on voice data related to the target object. 12.An apparatus for detecting voice, comprising: an audio to be detected and target voice feature obtaining module configured to obtain audio to be detected and a target voice feature for a target object, the audio to be detected comprising voice information of the target object; an audio feature determining module configured to determine an audio feature for the audio to be detected based on the audio to be detected; and an audio segment determining module configured to determine an audio segment corresponding to the target object in the audio to be detected based on the audio feature and the target voice feature. 13.An electronic device, comprising: at least one processor; and a storage device configured to store at least one program, when the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1-11. 14.A computer readable storage medium having stored thereon a computer program, the computer program, when executed by a processor, implements the method according to any one of claims 1-11. 15.A computer program product comprising a computer program, the computer program, when executed by a processor, implements the method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Mixed voice recognition method and device and computer-readable storage medium

    CN108962237A

  • Audio detection method, audio detection device, storage medium and electronic equipment

    CN110232933A

  • Audio processing method and device, terminal and storage medium

    CN112820300A

  • Sound detection model training method, data processing method and related device

    CN113506566A

  • Speaker verification methods and apparatus

    US20170061968A1