Voice processing method and device, equipment, storage medium and program product
By dynamically adjusting the pause duration threshold during voice recording, the problem that fixed thresholds in traditional voice activity detection systems cannot adapt to differences in user speech rate is solved, achieving more accurate voice activity detection and more efficient voice processing.
Patent Information
- Application Number
- CN202410869821.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-28
- Publication Date
- 2025-12-30
AI Technical Summary
In traditional speech activity detection systems, fixed pause duration thresholds cannot adapt to the differences in speech rates among different users, resulting in inaccurate detection during speech recording and affecting the response speed of subsequent processes.
By determining the speech rate attribute information of the subject based on the speech collected during the speech recording process, the pause duration threshold is dynamically adjusted to control the speech activity detection in a personalized and adaptive manner.
It improves the accuracy and adaptability of voice activity detection, ensures timely response to the start and end of the user's voice during voice recording, and enhances the personalization and efficiency of the voice processing system.
Smart Images

Figure CN121237130A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Example embodiments of the present disclosure generally relate to the field of computer technology, and in particular, to a method, an apparatus, an electronic device, a computer readable storage medium and a computer program product for speech processing. BACKGROUND
[0002] With the development of Internet technology, more and more applications or platforms provide voice services, which brings great convenience to the general public. Handheld or portable devices support voice interaction dialogue using voice technology, such as mobile phones, tablet computers, notebook computers, and even some wearable devices support voice interaction. Voice activity detection (VAD) is a technology for speech processing, which aims to detect whether a voice signal exists. Applications or platforms with voice services can use VAD technology to detect the start and end of user speech, avoiding noise or other sounds from being recorded by the device to produce unintended results. SUMMARY
[0003] In a first aspect of the present disclosure, a method for speech processing is provided. The method comprises: determining attribute information related to a subject based on a voice collected during a voice recording process of the subject, the attribute information at least indicating a speech rate of the subject; determining configuration information of a pause duration threshold for the subject during the voice recording process based on the attribute information; and controlling voice activity detection of the voice recording process based on the configuration information of the pause duration threshold.
[0004] In a second aspect of the present disclosure, an apparatus for speech processing is provided. The apparatus comprises: an attribute information determination module configured to determine attribute information related to a subject based on a voice collected during a voice recording process of the subject, the attribute information at least indicating a speech rate of the subject; a configuration information determination module configured to determine configuration information of a pause duration threshold for the subject during the voice recording process based on the attribute information; and a voice detection control module configured to control voice activity detection of the voice recording process based on the configuration information of the pause duration threshold.
[0005] In a third aspect of the present disclosure, an electronic device is provided. The device comprises at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the electronic device to perform the method of the first aspect.
[0006] In a fourth aspect of the present disclosure, a computer readable storage medium is provided. The medium has stored thereon a computer program, which, when executed by a processor, implements the method of the first aspect.
[0007] In a fifth aspect of the disclosure, a computer program product is provided. The product includes a computer program, wherein the computer program, when executed by a processor, implements the method according to the first aspect of the disclosure.
[0008] It should be understood that all statements herein made regarding the examples described in this section are intended to be non-limiting and are intended to be included in the scope of the embodiments of the present disclosure. Moreover, it should be understood that the features of the various aspects of the present disclosure can be combined with each other, unless specifically noted otherwise. BRIEF DESCRIPTION OF DRAWINGS
[0009] The above and other features, aspects, and advantages of embodiments of the present disclosure will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings. In the drawings, like reference numerals refer to like elements, wherein:
[0010] Figure 1A A schematic diagram illustrating an example environment in which embodiments of the present disclosure can be implemented is shown;
[0011] Figure 1B An example of voice activity detection is shown;
[0012] Figure 2 A schematic diagram illustrating an example architecture for speech processing according to some embodiments of the present disclosure is shown;
[0013] Figure 3 An example of a fitted curve of a difference scaling factor according to some embodiments of the present disclosure is shown;
[0014] Figure 4 A schematic diagram illustrating a signaling flow for speech processing according to some embodiments of the present disclosure is shown;
[0015] Figure 5 A flow diagram of a method for speech processing according to some embodiments of the present disclosure is shown;
[0016] Figure 6 An exemplary structural block diagram of an apparatus for speech processing according to some embodiments of the present disclosure is shown; and
[0017] Figure 7 A block diagram of an electronic device that can implement one or more embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be implemented in various forms and should not be interpreted as being limited to the embodiments set forth herein, but rather, these embodiments are provided so as to more completely and thoroughly understand the present disclosure. It is understood that the drawings and embodiments of the present disclosure are for exemplary purposes only and are not intended to limit the scope of protection of the present disclosure.
[0019] In the description of embodiments of the present disclosure, the term "comprising" and its conjugations should be understood to encompass the meaning of "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions can also be included below.
[0020] In this document, unless explicitly stated, performing a step "in response to A" does not mean performing the step immediately after A, but can include one or more intermediate steps.
[0021] It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the obtaining, use, storage or deletion of the data) should comply with the requirements of relevant laws and regulations and relevant provisions.
[0022] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type of information involved in the present disclosure, the scope of use, the use scenario, etc. should be informed to the relevant user and the authorization of the relevant user should be obtained by appropriate means, wherein the relevant user can include any type of right subject, such as an individual, an enterprise or a group.
[0023] For example, in response to receiving the active request of the user, a prompt information is sent to the relevant user to explicitly prompt the relevant user that the operation requested to be performed will require the information of the relevant user to be obtained and used, so that the relevant user can voluntarily choose whether to provide the information to the software or hardware such as electronic device, application program, server or storage medium, etc. performing the operation of the technical solutions of the present disclosure according to the prompt information.
[0024] As an optional but non-limiting implementation manner, in response to receiving the active request of the relevant user, the prompt information can be sent to the relevant user in the form of a pop-up window, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide information to the electronic device.
[0025] It can be understood that the above notification and user authorization obtaining process is only illustrative and does not limit the implementation of the present disclosure, and other ways that meet relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0026] As used herein, the term “model” can learn the relationship between the corresponding input and output from the training data, so that after the training is completed, the corresponding output can be generated for a given input. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes input and provides a corresponding output by using multiple layers of processing units. The neural network model is one example of a model based on deep learning. In this document, “model” can also be referred to as “machine learning model”, “learning model”, “machine learning network” or “learning network”, which are used interchangeably herein.
[0027] A “neural network” is a machine learning network based on deep learning. The neural network is capable of processing input and providing a corresponding output, which generally includes an input layer and an output layer and one or more hidden layers between the input layer and the output layer. The neural network used in deep learning applications generally includes many hidden layers, thereby increasing the depth of the network. The layers of the neural network are connected in sequence, so that the output of the previous layer is provided as the input of the next layer, where the input layer receives the input of the neural network, and the output of the output layer is the final output of the neural network. Each layer of the neural network includes one or more nodes (also known as processing nodes or neurons), each of which processes input from the previous layer.
[0028] Generally, machine learning can include three stages, namely a training stage, a testing stage and an application stage (also known as an inference stage). In the training stage, a given model can be trained using a large amount of training data, iteratively updating the parameter values until the model can obtain consistent inferences from the training data that meet the expected target. Through training, the model can be considered to learn the relationship between input and output (also known as the mapping from input to output) from the training data. The parameter values of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, thereby determining the performance of the model. In the application stage, the model can be used to process the actual input based on the parameter values obtained by training to determine the corresponding output.
[0029] Figure 1AA schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. The example environment 100A involves a client device 110 and a server device 130. The client device 110 can have installed therein, for example, an application with natural language processing functionality. The application can be any suitable type of application, such as a social application, a chat application, a media item application, etc. The application supports at least voice-based services. The server device 130 can provide background services for this application with natural language processing functionality.
[0030] The client device 110 and / or the application in the client device 110 can include an audio collector 112 and a voice activity detection system 114. The audio collector 112 is configured to collect the voice 125 of the user 120 during a voice recording process. The audio collector 112 can be a microphone built-in or attached to the client device 110, for example. The voice activity detection system 114 is configured to control voice activity detection (VAD) of the voice recording process based on a pause duration to determine the start and end of the user voice. For example, the voice activity detection system 114 can determine whether the user speech sound ends based on the pause duration during the voice recording process. The audio collector 112 is instructed whether to collect the voice 125 of the user 120 based on the determination of the end.
[0031] Reference is made to Figure 1B , Figure 1B An example 100B of voice activity detection is shown. As shown in Figure 1B , the start and end of the voice recording are generally referred to as VAD START and VAD END. The client device 110 can determine the VAD START in response to detecting the voice 125 from the user 120, and start recording the voice. Specifically, when the sound is recorded, the human speech characteristics can be detected to match the predetermined characteristic target through spectral analysis, at which time the voice activity detection system 114 determines that the VAD START is detected, and instructs the audio collector 112 to start recording. The client device 110 can start VAD END pause detection in response to detecting no voice 125. Specifically, when the user stops speaking, at which time the voice activity detection system 114 waits for a period of time to monitor whether the user continues to speak. If no user voice is detected, at which time the voice activity detection system 114 sends a VAD END signal. The duration of the VAD END pause detection can be referred to as a pause duration. If the pause duration reaches a pause duration threshold, the voice activity detection system 114 can determine that the voice ends, and thus can stop recording the voice, pass the recorded voice for subsequent processing, etc.
[0032] In some embodiments, Figure 1AClient device 110 in environment 100 communicates with server device 130 to enable provisioning of voice processing services. Client device 110 can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a tablet computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a game device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. In some embodiments, client 110 can also support any type of interface to the user (such as "wearable" circuitry, etc.).
[0033] Server device 130 can be a standalone physical server, a server cluster or distributed system of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and big data and artificial intelligence platforms, etc. basic cloud computing services. Server device 130 may, for example, include a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, etc.
[0034] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only, without implying any limitation on the scope of the present disclosure.
[0035] As can be seen from the above voice activity detection process, the detection of the end of speech (VAD END) in the voice recording plays an important role in whether the voice system can respond to the subsequent process in time. For example, in many voice interaction scenarios, voice recognition (ASR) needs to be performed on the voice, and subsequent responses to the user and text-to-speech (TTS) are determined based on the voice or ASR results, etc. If the user stops speaking cannot be detected in time, the response speed of the subsequent action will be affected, resulting in a long waiting time for the user.
[0036] Conventionally, the voice activity detection system often controls the voice activity detection of the voice recording process based on a fixed pause duration threshold. Since the duration of the pause during recording by different users of different types may be different, using a fixed pause duration threshold to determine the end time will not match all users. For example, for a user with a slower speech rate, if the pause duration threshold is small, the voice activity monitoring system may determine to stop recording before the user has finished speaking. It is desirable to set different pause duration thresholds for different types of users.
[0037] In view of this, according to an embodiment of the present disclosure, an improved solution for speech processing is provided. According to the solution of the present disclosure, based on the speech collected during the speech recording process of the subject, attribute information related to the subject is determined, the attribute information at least indicating the speech speed of the subject. Based on the attribute information, configuration information of a pause duration threshold for the subject during the speech recording process is determined. Based on the configuration information of the pause duration threshold, voice activity detection on the speech recording process is controlled.
[0038] In this way, the speech speed of the subject can be determined based on the collected speech, and the pause duration threshold when performing voice activity detection can be adjusted based on the speech speed of the subject. The pause duration threshold required for voice activity detection can be dynamically adjusted for different subjects in different scenarios, which helps to improve the personalization and adaptability of voice activity detection for different subjects, and improve the accuracy of the pause duration threshold.
[0039] Some example embodiments of the present disclosure will be described below with continuous reference to the drawings.
[0040] Figure 2 A schematic diagram of an example architecture 200 for speech processing according to some embodiments of the present disclosure is shown. The example architecture 200 involves a client device 110 and a server device 130. For ease of discussion, the architecture 200 will be described with reference to the environment 100 of FIG. 1. It should be noted that the operations performed by the aforementioned client device 110 and the operations performed by the client device 110 described later can be performed by a relevant application installed on the client device 110. In some embodiments, the operations performed on the electronic device can be completed with the assistance of the server device 130.
[0041] The client device 110 can include an audio collector 112, a voice activity detection system 114 (also referred to as a VAD system 114), a VAD adjustment system 212, and a speech processing system 214. The client device 110 may, for example, in response to receiving a speech recording request, instruct the audio collector 112 to collect speech 125 from a subject 201. The subject 201 may, for example, include a user (e.g., the user 120), or can include a sound-emitting object other than a human being. The audio collector 112 may, for example, include a microphone.
[0042] The example architecture 200 also involves a VAD configuration system 230. The VAD configuration system 230 can be deployed at the server device 130 or at the client device 110. If the VAD configuration system 230 is configured at the client device 110, the client device 110 can directly send the speech 125 collected during the speech recording process of the subject 201 to the VAD configuration system 230. If the VAD configuration system 230 is configured at the server device 130, the client device 110 may, for example, send the speech 125 collected during the speech recording process of the subject 201 to the speech service 222 of the server device 130, so that the speech service 222 sends the speech 125 to the VAD configuration system 230.
[0043] The VAD configuration system 230 can determine attribute information related to the subject 201 based on the speech 125. In some embodiments, the VAD configuration system 230 can determine the attribute information based on the speech 125 only if authorization is obtained. The VAD configuration system 230 may, for example, provide a user with an authorization request interface when the user first interacts with the digital assistant or each time the digital assistant is woken up, and receive the user’s configuration of the permission via the authorization request interface. The authorization request interface may, for example, include permission prompt information and at least one operation control. The at least one operation control can include an authorization confirmation control. The permission prompt information may, for example, include the text “whether to allow the system to determine relevant information based on the user’s speech”, and the VAD configuration system 230 can determine that the user’s authorization is obtained in response to detecting the triggering of the confirmation control.
[0044] As will be described below, the attribute information is used to determine or adjust the pause duration threshold for the speech activity detection of the subject 201. In some embodiments, the VAD configuration system 230 can determine the attribute information based on the speech 125 using a machine learning model. This machine learning model can be based on any appropriate model structure, including but not limited to a Transformer model, a convolutional neural network (CNN), a recurrent neural network (RNN), a deep neural network (DNN), etc. In some embodiments, the machine learning model can be based on a language model (LM). The language model can have the ability to ask and answer by learning from a large amount of corpus.
[0045] Exemplarily, as Figure 2As shown, the VAD configuration system 230 can include a language model 232. The VAD configuration system 230 can provide the speech 125 to the language model 232. The VAD configuration system 230 may, for example, determine a prompt for the language model 232 based on the speech 125. The prompt includes guidance information to the language model 232, which may, for example, be used to guide how the model uses the speech 125 to generate the attribute information. In some embodiments, the speech processing service can be provided based on a trained at least one machine learning model. The at least one machine learning model can include one or more language models. A language model used for processing speech (e.g., a language model used to perform speech synthesis TTS, a language model used to perform speech recognition ASR, etc.) can be referred to as a primary language model or language model, and the language model 232 used to determine the attribute information can be referred to as a secondary language model.
[0046] The model output of the language model 232 may, for example, indicate the attribute information of the object 201. This attribute information can at least indicate the speech rate of the object 201. This is because the speech rate of the object can affect the pause duration in the speech of the object, and thus affect the setting of the pause duration threshold. Of course, the attribute information can include any suitable information, which can also indicate any suitable information other than the speech rate (e.g., the type of the object), as long as the information can affect the pause duration in the speech of the object. In some embodiments, the attribute information may, for example, also indicate a speaking scenario related to the object 201. In different speaking scenarios, the object can exhibit different speech rates or pause durations. For example, in a lecture or meeting scenario, the normal pause duration in the speech of the object can be longer, and thus the pause duration threshold can be set to be larger. In a noisy scenario, the speech of the object can be more rapid, and thus the pause duration threshold can be set to be smaller.
[0047] The VAD configuration system 230 can also include a VAD service module 234. The VAD service module 234 may, for example, determine the configuration information of the pause duration threshold for the object 201 in the speech recording process based on the determined attribute information. It is noted that in some embodiments, the VAD configuration system 230 can also directly provide the attribute information to the client device 110, so that the client device 110 determines the configuration information of the pause duration threshold based on the received attribute information.
[0048] VAD service module 234 can determine the configuration information for the pause duration threshold in any appropriate manner. For example, VAD service module 234 can use predetermined rules or algorithms to determine the configuration information for the pause duration threshold based on attribute information. Alternatively, VAD service module 234 can also use a trained machine learning model to determine the configuration information for the pause duration threshold based on attribute information related to object 201. This machine learning model and the machine learning model used to determine the attribute information can be the same model or different models. VAD service module 234 can, for example, provide the attribute information to the machine learning model and determine the configuration information for the pause duration threshold based on the model output of this machine learning model.
[0049] The VAD configuration system 230 can directly determine the pause duration threshold based on attribute information. In some embodiments, the configuration information for the pause duration threshold can directly indicate the specific value of the pause duration threshold. In this case, the VAD configuration system 230 can directly determine the pause duration threshold based on the configuration information. Alternatively or additionally, in some embodiments, the configuration information may also indicate adjustment information for the predetermined pause duration threshold, such as a threshold adjustment factor. The VAD configuration system 230 can, for example, adjust the predetermined pause duration threshold based at least on the threshold adjustment factor. The threshold adjustment factor may, for example, include a difference scaling factor (factor) for scaling, an increment factor (delta) for increasing or decreasing, and so on.
[0050] In some embodiments, taking the threshold adjustment factor as an example of the difference scaling factor, the VAD service module 234 can determine the difference scaling factor based on the object's attribute information and can adjust the predetermined pause duration threshold based on the difference scaling factor. In some embodiments, the difference scaling factor is used to scale the predetermined pause duration threshold to obtain the adjusted pause duration threshold. The fitting curve of the difference scaling factor can be configured as follows: Figure 3 As shown. Figure 3 Example 300 shows a fitted curve for the difference scaling factor according to some embodiments of the present disclosure. The difference scaling factor can, for example, be in the range of (0, 1). The faster the speech rate indicated by the attribute information, the closer the value of the difference scaling factor is to 1, which means a smaller scaling degree to the predetermined pause duration threshold. The slower the speech rate indicated by the attribute information, the closer the value of the difference scaling factor is to 0, which means a larger scaling degree to the predetermined pause duration threshold.
[0051] In some embodiments, the VAD service module 234 can also determine the longest waiting time (T) during the voice recording process based on attribute information. f ) and minimum waiting time (T) sThe waiting time can be, for example, the interval between acquiring two adjacent speech units. The VAD service module 234 can, for example, determine the predetermined pause duration threshold using the following formula based on the difference scaling factor, the longest waiting time, and the shortest waiting time:
[0052] t=(T f -T s )*factor (1)
[0053] V t =T s +t (2)
[0054] VAD service module 234, for example, can use a threshold adjustment factor (factor) to determine the adjustment amount of the predetermined pause duration threshold based on formula (1); and adjust the predetermined pause duration threshold T based on formula (2). s It is understood that the above only provides an example of how to adjust the predetermined pause duration threshold. In other embodiments, the predetermined pause duration threshold can be adjusted in other ways.
[0055] In some embodiments, as mentioned above, the VAD service module 234 can directly determine the pause duration threshold to be used without needing to determine adjustment information for the predetermined pause duration threshold. In this case, the dynamically determined pause duration threshold can be directly used for speech activity detection.
[0056] In some embodiments, after determining the pause duration threshold for a specific object, the VAD configuration system 230 can provide the determined pause duration threshold to the client device 110. It should be noted that in some embodiments, the VAD configuration system 230 can also directly provide the pause duration threshold configuration information to the client device 110, so that the client device 110 can determine the pause duration threshold itself based on the pause duration threshold configuration information. For example, the client device 110 includes a VAD adjustment system 212.
[0057] The VAD adjustment system 212 can determine the pause duration threshold during speech detection based on configuration information of the pause duration threshold. In some embodiments, if the newly determined pause duration threshold differs from a previous pause duration threshold (e.g., a historical pause duration threshold determined based on historical speech), the VAD adjustment system 212 can also update the pause duration threshold, and the updated pause duration threshold is the newly determined pause duration threshold. The VAD adjustment system 212 can, for example, update the pause duration threshold in real time. The VAD adjustment system 212 can also, for example, determine the pause duration threshold to be adjusted in response to the satisfaction of a predetermined condition. The predetermined condition may include, for example, that the difference between the newly determined pause duration threshold and the previously used pause duration threshold is greater than a predetermined difference. The predetermined condition may also include, for example, that the time elapsed since the last update of the pause duration threshold until the time elapsed since the determination of the new pause duration threshold reaches a predetermined duration.
[0058] Client device 110 is configured to perform a voice recording process. Client device 110 can perform voice activity detection during the voice recording process based on a comparison between the pause duration detected during the voice recording process and a pause duration threshold. For example, the voice activity detection system 114 of client device 110 can detect pause durations during the voice recording process where no voice is detected, and stop acquiring voice when the pause duration reaches the pause duration threshold. The voice activity detection system 114 can, for example, provide the acquired voice 125 to the voice processing system 214. The voice processing system 214 can, for example, provide the acquired voice 125 to the voice service service 222 of server device 130. In some embodiments, the voice processing system 214 can also determine the type of voice processing service (e.g., speech recognition, speech synthesis, etc.) that object 201 wants to perform on the voice 125, and inform the voice service service 222 of this type.
[0059] Taking the determination of whether to perform a speech recognition service on speech 125 as an example, the voice service 222 may provide speech 125 to the speech recognition model 224, and may use the speech recognition model 224 to determine the text corresponding to speech 125. The voice service 222 may provide the text corresponding to speech 125 to the client device 110. The client device 110 may present a specific page (e.g., a session page) and provide the text corresponding to speech 125 to the object 201 based on the page.
[0060] In some embodiments, example architecture 200 may further include a VAD feedback system 242. Similar to VAD configuration system 230, VAD feedback system 242 can also be configured at server device 130 or client device 110. VAD feedback system 242 may, for example, acquire feedback information during the process of performing speech activity detection using a pause duration threshold. The feedback information may, for example, indicate to object 201 its level of satisfaction (e.g., satisfied or dissatisfied) with the performance of speech activity detection using the pause duration threshold, suggestions for adjustments to the performance of speech activity detection using the pause duration threshold (e.g., a desire to decrease or increase the pause duration threshold), and so on.
[0061] The VAD feedback system 242 can, for example, adjust the determination of object-related attribute information based on feedback information. Exemplarily, the VAD feedback system 242 can instruct the VAD configuration system 230 to adjust the method of determining attribute information, adjust the parameters of the language model 232, etc., to determine more accurate attribute information. The VAD feedback system 242 can also, for example, adjust the determination of pause duration threshold configuration information based on feedback information. Exemplarily, the VAD feedback system 242 can instruct the VAD configuration system 230 to adjust the method of determining the pause duration threshold configuration information, adjust the method of determining the threshold adjustment factor, etc., to determine more accurate pause duration threshold configuration information, thereby adaptively determining the pause duration threshold to be used.
[0062] Figure 4 A schematic diagram of a signaling stream 400 for voice processing according to some embodiments of the present disclosure is shown. In the signaling stream 400, a voice activity detection system 114 may initiate a voice recording process in response to receiving (401) a voice recording request from an object 201. During the voice recording process, the voice activity detection system 114 may, for example, provide (402) the voice collected during the voice recording process to the server device 130 in response to a pause duration reaching a predetermined pause duration threshold.
[0063] Server device 130 can provide (403) speech to language model 232 so that language model 232 can determine attribute information related to object 201 based on the speech collected during the speech recording of object 201, the attribute information indicating at least the speech rate of object 201. Server device 130 can also perform corresponding speech processing on the speech based on the type of speech processing service to be performed on the speech of object 201. For example, if it is determined that ASR / TTS should be performed on the speech of object 201, server device 130 can perform ASR / TTS on the speech using ASR model / TTS model (404). Server device 130 can provide (405) the text obtained from performing ASR / TTS on the speech to speech activity detection system 114. Speech activity detection system 114 can be configured at client device 110, for example, and client device 110 can provide the text via a specific page.
[0064] Language model 232 can, for example, provide (406) determined attribute information to speech activity detection system 114. Client device 110 can, for example, determine configuration information for pause duration thresholds for an object during speech recording based on the determined attribute information, and can determine the pause duration threshold based on the configuration information. If the newly determined pause duration threshold is different from the previous pause duration threshold, client device 110 can also adjust the pause duration threshold, and the adjusted pause duration threshold is also the newly determined pause duration threshold.
[0065] The speech activity detection system 114 can, for example, control the speech activity detection during the speech recording process based on an adjusted pause duration threshold. The speech activity detection system 114 can, for example, provide (408) newly acquired speech (i.e., adjusted speech) during the speech recording process to the server device 130 in response to a pause duration reaching the adjusted pause duration threshold. Similarly, the server device 130 can provide (409) the text obtained by performing speech processing on the speech to the speech activity detection system 114.
[0066] Client device 110 may, for example, receive feedback information during the process of performing speech activity detection using an adjusted pause duration threshold. Speech activity detection system 114 in client device 110 may, for example, provide the feedback information (410) to VAD feedback system 242. VAD feedback system 242 may, for example, adjust the determination of configuration information for object attribute information / pause duration threshold based on the feedback information. VAD feedback system 242 may, for example, determine a specific adjustment method / adjustment suggestion for the configuration information based on the feedback information, and may provide the adjustment method / adjustment suggestion (411) to speech activity detection system 114. Speech activity detection system 114 may, for example, adjust the determination of configuration information based on the received adjustment method / adjustment suggestion. VAD feedback system 242 may also, for example, determine a specific adjustment method / adjustment suggestion for attribute information based on the feedback information, and may provide the adjustment method / adjustment suggestion (412) to language model 232. Language model 232 may, for example, adjust the determination of attribute information based on the received adjustment method / adjustment suggestion.
[0067] In summary, according to the embodiments of this disclosure, the speech rate of an object can be determined based on the acquired speech, and the pause duration threshold during speech activity detection can be automatically adjusted based on the object's speech rate. Different pause duration thresholds can be determined for different objects, which helps to improve the personalization and adaptability of speech activity detection for different objects and improve the accuracy of the pause duration thresholds.
[0068] Figure 5 A flowchart of a method 500 for voice processing according to some embodiments of the present disclosure is shown. Method 500 may be implemented at a client device 110 and / or a server device 130.
[0069] In box 510, client device 110 and / or server device 130 determine object-related attribute information based on the speech collected during the speech recording of the object, the attribute information indicating at least the speech rate of the object.
[0070] In box 520, client device 110 and / or server device 130 determine configuration information for a pause duration threshold for an object during voice recording based on attribute information.
[0071] In box 530, client device 110 and / or server device 130 control the detection of voice activity during the voice recording process based on configuration information of pause duration threshold.
[0072] In some embodiments, method 500 is implemented at client device 110, which is configured to perform a voice recording process, and wherein controlling voice activity detection of the voice recording process based on configuration information of a pause duration threshold includes: determining a pause duration threshold for a target based on the configuration information; and performing voice activity detection of the voice recording process based on a comparison between a pause duration detected during the voice recording process and the pause duration threshold.
[0073] In some embodiments, method 500 is implemented at server device 130, and wherein controlling voice activity detection of the voice recording process based on configuration information of a pause duration threshold includes: sending the configuration information of the pause duration threshold to a client device, so that the client device determines the pause duration threshold based on the configuration information, the client device being configured to perform the voice recording process.
[0074] In some embodiments, determining attribute information related to an object includes: using a machine learning model to determine attribute information based on collected speech.
[0075] In some embodiments, the configuration information indicates a threshold adjustment factor, and wherein the pause duration threshold for an object is determined by adjusting a predetermined pause duration threshold based at least on the threshold adjustment factor.
[0076] In some embodiments, method 500 further includes: obtaining feedback information during the process of performing speech activity detection using a pause duration threshold; and adjusting at least one of the following based on the feedback information: determining object-related attribute information or determining configuration information for the pause duration threshold.
[0077] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 6 An exemplary structural block diagram of an apparatus 600 for voice processing according to some embodiments of the present disclosure is shown. The apparatus 600 may be implemented as or included in a client device 110 and / or a server device 130. The various modules / components in the apparatus 600 may be implemented by hardware, software, firmware, or any combination thereof.
[0078] like Figure 6 As shown, device 600 includes an attribute information determination module 610, configured to determine attribute information related to the object based on speech acquired during the object's speech recording process. The attribute information at least indicates the object's speech rate. Device 600 also includes a configuration information determination module 620, configured to determine configuration information for a pause duration threshold for the object during speech recording based on the attribute information. Device 600 further includes a speech detection control module 630, configured to control speech activity detection during the speech recording process based on the pause duration threshold configuration information.
[0079] In some embodiments, the apparatus 600 is implemented at a client device 110, which is configured to perform a voice recording process, and the voice detection control module 630 includes: a pause duration threshold determination module, configured to determine a pause duration threshold for a target based on configuration information; and a voice activity detection execution module, configured to perform voice activity detection of the voice recording process based on a comparison between the pause duration detected during the voice recording process and the pause duration threshold.
[0080] In some embodiments, the apparatus 600 is implemented at the server device 130, and the voice detection control module 630 includes: a configuration information sending module configured to send configuration information of a pause duration threshold to the client device, so that the client device can determine the pause duration threshold based on the configuration information, and the client device is configured to perform a voice recording process.
[0081] In some embodiments, the attribute information determination module 610 is further configured to: determine attribute information based on the collected speech using a machine learning model.
[0082] In some embodiments, the configuration information indicates a threshold adjustment factor, and wherein the pause duration threshold for an object is determined by adjusting a predetermined pause duration threshold based at least on the threshold adjustment factor.
[0083] In some embodiments, the apparatus 600 further includes: a feedback information acquisition module configured to acquire feedback information during the process of performing speech activity detection using a pause duration threshold; and an adjustment module configured to adjust at least one of the following based on the feedback information: the determination of object-related attribute information, or the determination of configuration information for the pause duration threshold.
[0084] The units and / or modules included in device 600 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units and / or modules in device 600 can be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0085] It should be understood that one or more steps in the above methods can be performed by appropriate electronic devices or combinations of electronic devices. Such electronic devices or combinations of electronic devices may, for example, include the client device 110 and / or the server device 130 in FIG1.
[0086] Figure 7 A block diagram of an electronic device 700 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 7 The electronic device 700 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 7 The electronic device 700 shown can be used to implement the client device 110 and / or server device 130 of FIG1.
[0087] like Figure 7 As shown, electronic device 700 is in the form of a general-purpose electronic device. Components of electronic device 700 may include, but are not limited to, one or more processors or processing units 710, memory 720, storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. Processing unit 710 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 720. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 700.
[0088] Electronic device 700 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 720 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 730 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 700.
[0089] Electronic device 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 7As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 720 may include computer program product 725 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.
[0090] The communication unit 740 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 700 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0091] Input device 750 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 760 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 700 can also communicate with one or more external devices (not shown) via communication unit 740 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 700, or with any device that enables electronic device 700 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0092] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0093] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0094] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0095] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0096] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some, as newer, implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0097] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for speech processing, comprising: determining, based on speech collected during a speech recording process of a subject, attribute information related to the subject, the attribute information indicating at least a speaking rate of the subject; determining, based on the attribute information, configuration information of a pause duration threshold for the subject during the speech recording process; and controlling, based on the configuration information of the pause duration threshold, voice activity detection of the speech recording process. 2.The method of claim 1, wherein the method is implemented at a client device configured to perform the speech recording process, and wherein controlling, based on the configuration information of the pause duration threshold, voice activity detection of the speech recording process comprises: determining, based on the configuration information, the pause duration threshold for the subject; and performing, based on a comparison between a pause duration detected during the speech recording process and the pause duration threshold, voice activity detection of the speech recording process. 3.The method of claim 1, wherein the method is implemented at a server device, and wherein controlling, based on the configuration information of the pause duration threshold, voice activity detection of the speech recording process comprises: sending, to a client device configured to perform the speech recording process, the configuration information of the pause duration threshold for determining, by the client device, the pause duration threshold based on the configuration information. 4.The method of claim 1, wherein determining attribute information related to the subject comprises: determining, based on the collected speech, the attribute information using a machine learning model. 5.The method of claim 1, wherein the configuration information indicates a threshold adjustment factor, and wherein the pause duration threshold for the subject is determined by adjusting a predetermined pause duration threshold based at least on the threshold adjustment factor. 6.The method of claim 1, further comprising: obtaining feedback information during the voice activity detection performed using the pause duration threshold; and adjusting, based on the feedback information, at least one of: determination of attribute information related to the subject, or determination of configuration information of a pause duration threshold. 7.An apparatus for speech processing, comprising: an attribute information determination module configured to determine, based on speech collected during a speech recording process of a subject, attribute information related to the subject, the attribute information indicating at least a speaking rate of the subject; a configuration information determination module configured to determine, based on the attribute information, configuration information of a pause duration threshold for the subject during the speech recording process; and a voice detection control module configured to control, based on the configuration information of the pause duration threshold, voice activity detection of the speech recording process. 8.An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit cause the electronic device to perform the method according to any one of claims 1 to 6.
9. A computer readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 6.
10. A computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1 to 6.