Audio processing method and apparatus, device, medium and program product

By dynamically adjusting sensitivity parameters and semantic detection strategies according to environment type, the problem of poor flexibility and adaptability in existing technologies is solved, improving the accuracy and efficiency of audio processing and enhancing the user experience.

WO2026002047A1PCT designated stage Publication Date: 2026-01-02BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/103460
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-28
Filing Date
2025-06-25
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

In existing audio processing technologies, fixed sensitivity parameters and semantic detection strategies result in poor flexibility and adaptability, affecting audio processing efficiency and user experience.

Method used

The sensitivity parameters and semantic detection strategy are dynamically adjusted based on the environment of the client device to control the audio processing.

Benefits of technology

It improves the accuracy of voice rejection and audio processing efficiency, enhancing the user's audio processing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025103460_02012026_PF_FP_ABST
    Figure CN2025103460_02012026_PF_FP_ABST
Patent Text Reader

Abstract

An audio processing method and apparatus, a device, a medium and a program product. The method comprises: determining from first audio collected by a client device an environment type in which the client device is located (310); and, at least on the basis of the environment type, controlling a sensitivity parameter and / or a semantic detection policy applied to second audio collected by the client device (320), the sensitivity parameter being configured to control whether to perform speech recognition on the second audio, and the semantic detection policy being applied to a semantic detection process executed on text corresponding to the second audio.
Need to check novelty before this filing date? Find Prior Art

Description

Method, apparatus, device, medium and program product for audio processing

[0001] The present application claims priority to the Chinese patent application No. 202410870237.9, filed on June 28, 2024, entitled “Method, apparatus, device, medium and program product for audio processing”, the whole content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] The example embodiments of the present disclosure generally relate to the field of computer, and in particular, to a method, apparatus, electronic device, computer readable storage medium and computer program product for audio processing. BACKGROUND

[0003] With the development of Internet technology, more and more applications or platforms, etc. provide audio processing function / voice processing function, which brings great convenience to the general public. Common voice processing technologies may include, for example, speech synthesis (TTS) technology (also known as text-to-speech technology), speech recognition (ASR) technology (also known as speech-to-text technology), back channel response technology, reject recognition technology (also known as reject technology), etc. Reject technology plays an important role in voice interaction, which mainly functions to determine whether the audio needs to be processed further, and whether the audio is noise or voice without obvious semantics. SUMMARY

[0004] In a first aspect of the present disclosure, a method for audio processing is provided. The method comprises: determining, from a first audio collected by a client device, an environment type in which the client device is located; and controlling, based at least on the environment type, a sensitivity parameter and / or a semantic detection strategy applied to a second audio collected by the client device, the sensitivity parameter being configured to control whether to perform speech recognition on the second audio, and the semantic detection strategy being applied to a semantic detection process performed on a text corresponding to the second audio.

[0005] In a second aspect of the present disclosure, an apparatus for audio processing is provided. The apparatus comprises: an environment type determination module configured to determine, from a first audio collected by a client device, an environment type in which the client device is located; and a parameter strategy control module configured to control, based at least on the environment type, a sensitivity parameter and / or a semantic detection strategy applied to a second audio collected by the client device, the sensitivity parameter being configured to control whether to perform speech recognition on the second audio, and the semantic detection strategy being applied to a semantic detection process performed on a text corresponding to the second audio.

[0006] In a third aspect of the disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the electronic device to perform the method of the first aspect.

[0007] In a fourth aspect of the disclosure, a computer-readable storage medium is provided. The medium has stored thereon a computer program which, when executed by a processor, implements the method of the first aspect.

[0008] In a fifth aspect of the disclosure, a computer program product is provided. The product includes a computer program, wherein the computer program, when executed by a processor, implements the method according to the first aspect of the disclosure.

[0009] It is to be understood that the details set forth in this section are not intended to limit key or critical features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0010] The above and other features, advantages and aspects of embodiments of the present disclosure will become more apparent upon reading the following detailed description in conjunction with the accompanying drawings, in which like references refer to like elements. In the drawings:

[0011] FIG. 1A shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;

[0012] FIG. 1B shows a schematic diagram of an example of a voice interaction;

[0013] FIG. 2A shows a schematic diagram of an example architecture for audio processing according to some embodiments of the present disclosure;

[0014] FIG. 2B shows a schematic diagram of an example architecture for audio processing according to some other embodiments of the present disclosure;

[0015] FIG. 2C shows example spectrograms of ambient sound and user speech according to the present disclosure;

[0016] FIG. 3 shows a flowchart of a method for audio processing according to some embodiments of the present disclosure;

[0017] FIG. 4 shows an exemplary structural block diagram of an apparatus for audio processing according to some embodiments of the present disclosure; and

[0018] FIG. 5 shows a block diagram of an electronic device in which one or more embodiments of the present disclosure can be implemented. DETAILED DESCRIPTION

[0019] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be implemented in various forms and should not be interpreted as being limited to the embodiments set forth herein, but rather, these embodiments are provided so as to more completely and thoroughly understand the present disclosure. It is understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0020] In the description of embodiments of the present disclosure, the term "comprising" and similar terms are to be interpreted as open-ended, i.e., "including but not limited to". The term "based on" is to be interpreted as "based, at least in part, on". The term "one embodiment" or "the embodiment" is to be interpreted as "at least one embodiment". The term "some embodiments" is to be interpreted as "at least some embodiments". Other explicit or implicit definitions can also be included below.

[0021] In this document, unless explicitly stated otherwise, performing a step "in response to A" does not mean that the step is performed immediately after A, but can include one or more intermediate steps.

[0022] It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the obtaining, use, storage or deletion of the data) should comply with the requirements of relevant laws and regulations and relevant provisions.

[0023] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type of information involved in the present disclosure, the scope of use, the use scenario, etc. should be informed to the relevant user and the authorization of the relevant user should be obtained through appropriate means, wherein the relevant user can include any type of right subject, such as an individual, an enterprise or a group.

[0024] For example, in response to receiving a proactive request of a user, a prompt information is sent to the relevant user to explicitly prompt the relevant user that the operation requested to be performed will require obtaining and using the information of the relevant user, so that the relevant user can voluntarily choose whether to provide the information to the software or hardware such as an electronic device, an application program, a server or a storage medium performing the operation of the technical solutions of the present disclosure according to the prompt information.

[0025] As an optional but non-limiting implementation manner, in response to receiving a proactive request of a relevant user, the manner of sending a prompt information to the relevant user can be, for example, a pop-up window manner, in which the prompt information can be presented in the form of text. In addition, the pop-up window can also carry a selection control for the user to select "agree" or "disagree" to provide information to the electronic device.

[0026] It can be understood that the above notification and user authorization obtaining process is only illustrative and does not limit the implementation of the present disclosure, and other ways that meet relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0027] As used herein, the term “model” can learn the relationship between the corresponding input and output from the training data, so that after the training is completed, the corresponding output can be generated for a given input. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes input and provides a corresponding output by using multiple layers of processing units. The neural network model is one example of a model based on deep learning. In this document, “model” can also be referred to as “machine learning model”, “learning model”, “machine learning network” or “learning network”, which are used interchangeably herein.

[0028] A “neural network” is a machine learning network based on deep learning. The neural network is capable of processing input and providing a corresponding output, which generally includes an input layer and an output layer and one or more hidden layers between the input layer and the output layer. The neural network used in deep learning applications usually includes many hidden layers, thereby increasing the depth of the network. The layers of the neural network are connected in sequence, so that the output of the previous layer is provided as the input of the next layer, where the input layer receives the input of the neural network, and the output of the output layer is the final output of the neural network. Each layer of the neural network includes one or more nodes (also known as processing nodes or neurons), each of which processes input from the previous layer.

[0029] Generally, machine learning can include three stages, namely a training stage, a testing stage and an application stage (also known as an inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values are updated iteratively until the model can obtain consistent inferences from the training data that meet the expected target. Through training, the model can be considered to learn the relationship between input and output (also known as the mapping from input to output) from the training data. The parameter values of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, thereby determining the performance of the model. In the application stage, the model can be used to process the actual input based on the parameter values obtained by training to determine the corresponding output.

[0030] FIG. 1A illustrates a schematic diagram of an example environment 100A in which embodiments of the present disclosure can be implemented. In this example environment 100A, an application 120 is installed in a client device 110. A user 140 can interact with the application 120 via the client device 110 and / or an attached device of the client device 110. For example, the application 120 can capture speech 145 of the user 140 via an audio capturing device (e.g., a microphone) of the client device 110.

[0031] In embodiments of the present disclosure, the application 120 can be any suitable application with audio processing functionality. The application 120 may, for example, be a social application, a chat application, a media item application, or any suitable type of application. The application 120 may, for example, also support voice-based services. For example, the application 120 can provide a digital assistant for human-to-computer dialog. The digital assistant supports text dialog services, voice dialog services, and content dialog in other modalities with the user 140. In some embodiments, the application 120 or the digital assistant therein can utilize a machine learning model 160 (which can include one or more machine learning models, e.g., can include a machine learning model 160-1, a machine learning model 160-2, …, a machine learning model 160-N, etc., where N is a positive integer) to support interactions with the user 140. For example, the application 120 or the digital assistant therein can utilize one or more machine learning models 160 to provide question-and-answer services to the user 140.

[0032] In the environment 100, the client device 110 can present a user interface 150 of the application 120 if the application 120 is in an active state. The user interface 150 can include various pages that the application 120 is capable of providing, such as a dialog page of the user with the digital assistant (in which current and historical dialog, including text dialog content, can be presented), etc. In some embodiments, the client device 110 can play speech 152 and present text 154 in the user interface 150. The speech 152 may, for example, include the speech 145 from the user 140 or speech in response to the speech 145.

[0033] The machine learning models 160 can be different types of models. In some embodiments, one or more machine learning models 160 can be built based on a language model (LM). The machine learning models used are content generative models, capable of generating a corresponding output based on a model input. In some embodiments, the machine learning models based on a language model are capable of processing a model input in a text modality (e.g., natural language and / or machine language) and / or a non-text modality (e.g., image, speech, video, etc.), and are capable of generating a desired output from the model input and a prompt word. The prompt word here is used to guide the machine learning model to generate an output that is capable of addressing a user need indicated by the model input. In an application scenario for supporting user conversation, an input from the user 140 can be provided as at least a portion of the model input (other portions can include the prompt word) to the machine learning model 160. The input from the user 140 is considered as a question. Based on the model output, a corresponding answer can be generated to be provided to the user 140.

[0034] In some embodiments, one or more machine learning models 160 can be voice related models, including an automatic speech recognition (ASR) model and a text-to-speech (TTS) model. The input to the ASR model is speech, and the output is text. The input to the TTS model is text, and the output is corresponding speech.

[0035] In some embodiments, the client device 110 communicates with the server device 130 to implement provisioning of services of the application 120. As shown in FIG. 1, the server device 130 can invoke the machine learning models 160 to support a human-to-computer dialog function between the application 120 and the user 140 based on the output of the machine learning models 160. The client device 110 can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media player, a

[0036] It should be understood that the structures and functions of the various elements in the environment 100A are described for illustrative purposes only, and do not imply any limitation on the scope of the present disclosure.

[0037] FIG. IB illustrates a schematic diagram of an example 100B of voice interaction. The example 100B includes an architecture 170, which can be implemented at the client device 110 and / or the server device 130. The architecture 170 designs a speech recognition module 171, a speech discrimination module 172, a semantic discrimination module 174, a question-answering model 175, and a speech synthesis module 176. In some embodiments, the architecture 170 can further include a pacification module 173.

[0038] The architecture 170 can obtain audio collected by the client device 110, which can include the speech 145 of the user 140, for example. A sensitivity parameter can be used to control a speech recognition process on the collected audio, for example. Illustratively, the sensitivity parameter can affect a content richness of the audio on which the speech recognition process is performed, for example. The higher the sensitivity parameter is, the higher sensitivity the speech recognition process is performed on the audio. For example, if the sensitivity parameter is high, background sound in the audio can also be recognized as text; if the sensitivity parameter is low, the background sound in the audio can be filtered out and the clearer human voice in the audio can be recognized as text.

[0039] The collected audio is provided to the speech discrimination module 172. The speech discrimination module 172 can determine an acceptance confidence for acoustic rejection based on the sensitivity parameter, for example. The sensitivity parameter and the acceptance confidence are inversely proportional. That is, the greater the sensitivity parameter is, the lower the acceptance confidence is. The speech discrimination module 172 can determine whether and how to perform the speech recognition on the audio based on the determined acceptance confidence. The speech discrimination module 172 can determine the audio in which the intensity reaches the acceptance confidence as the audio on which the speech recognition is performed and the audio in which the intensity does not reach the acceptance confidence as the audio on which the speech recognition is not performed, for example. That is, the speech discrimination module 172 can perform the speech recognition on only at least a portion of the audio.

[0040] The speech recognition module 171 can perform the speech recognition on the speech 145 to determine text corresponding to the speech 145. In some embodiments, if it is determined that the speech recognition is to be performed on the audio, the pacification module 173 can generate a prompt text for prompting that the user 140 is performing the processing on the audio. The prompt text can be a text “processing the audio”, a text “generating an answer for you”, and the like, for example. The prompt text can prompt the user 140 that the processing on the audio is currently being performed, can avoid the user from being anxious due to not seeing the execution result during the processing on the audio, and can improve the voice interaction experience of the user.

[0041] The speech recognition module 171 may, for example, provide text corresponding to the speech 145 to the semantic discrimination module 174. The semantic discrimination module 174 may, for example, perform semantic detection on the text to determine whether the text is semantically coherent and complete. The speech recognition module 171 may, for example, determine whether the speech 145 contains an explicit question demand or has an explicit semantic based on a semantic detection policy. The semantic detection policy may affect the strictness of semantic detection, and thus the passing rate of the text in semantic detection.

[0042] If it is determined that the recognized text passes semantic detection, the semantic discrimination module 174 may provide the text corresponding to the speech 145 to the question-answering model 175. The question-answering model 175 may determine an answer text based at least on the text. The architecture 170 may, for example, provide the answer text directly to the user 140. In some embodiments, the architecture 170 may also provide the answer text to the speech synthesis module 176. The speech synthesis module 176 may generate speech corresponding to the answer text based on the answer text. The architecture 170 may, for example, also provide the speech corresponding to the answer text to the user 140. Illustratively, the architecture 170 may present the answer text to the user 140 and play the answer speech to the user 140 via the user interface 150.

[0043] FIG. IB illustrates an example speech processing flow. In some other scenarios, for the collected audio, it can only be necessary to perform speech recognition and return the speech recognition result to the client device. In this case, the control parameter required for the entire speech processing flow is the sensitivity parameter.

[0044] Conventionally, in the entire speech processing flow, a fixed sensitivity parameter and a fixed semantic detection policy are often used to process the audio. The fixed sensitivity parameter and the fixed semantic detection policy can result in poor flexibility and adaptability of audio processing, fail to flexibly adjust the manner of audio processing based on the external environment, and can affect the audio processing efficiency and the user’s audio processing experience.

[0045] In view of this, according to an embodiment of the present disclosure, an improved scheme for audio processing is provided. According to the scheme of the present embodiment, the type of the environment in which the client device is located is determined based on first audio collected from the client device. At least based on the type of the environment, a sensitivity parameter and / or a semantic detection policy applied to second audio collected from the client device are controlled. The sensitivity parameter is configured to control whether speech recognition is performed on the second audio, and the semantic detection policy is applied to a semantic detection process performed on text corresponding to the second audio.

[0046] In this way, the sensitivity parameter and / or the semantic detection policy applied to the processing of the audio can be dynamically determined based on the type of the environment in which the client device is located. This can improve the accuracy of speech recognition, and thus improve the efficiency of audio processing.

[0047] Some example embodiments of the present disclosure will be hereinafter described with continuous reference to the drawings.

[0048] It is to be noted that the audio processing method of the present disclosure can be implemented completely at the client device 110, completely at the server device 130, partially at the client device 110 and partially at the server device 130. The client device 110 can collect audio, locally execute the audio processing method, or send the audio to the server device 130 for implementing the audio processing method.

[0049] FIG. 2A illustrates a schematic diagram of an example architecture 200A for audio processing, according to some embodiments of the present disclosure. The example architecture 200A can be implemented at the client device 110. For ease of discussion, the example architecture 200A will be described with reference to the environment 100 of FIG. 1. It is to be noted that the operations performed by the aforementioned client device 110 and the operations performed by the client device 110 as described hereinafter can be performed by a relevant application installed on the client device 110.

[0050] Referring to FIG. 2A, the client device 110 can collect the ambient sound 202 of the environment in which the client device 110 is located and the speech 145 from the user 140 by means of an audio collection device 210 (e.g., a microphone). The client device 110 can perform any appropriate processing on the collected audio (including the ambient sound 202 and / or the speech 145).

[0051] The client device 110 can further determine the type of the environment in which the client device 110 is located from the collected first audio. The first audio can include at least the ambient sound 202 of the environment in which the client device 110 is located, for example. In some embodiments, the first audio can be collected periodically by the client device 110. For example, the client device 110 can collect the ambient sound 202 of the environment in which it is located by means of the audio collection device 210 every predetermined time period (which can be determined by the user). In some embodiments, the collection of the ambient sound by the client device 110 is performed periodically only during the user’s use of the digital assistant (i.e., during the user’s interaction with the digital assistant).

[0052] In some embodiments, the client device 110 can periodically collect audio only if authorization is obtained. The client device 110, for example, can provide a user with an authorization request interface at the first interaction of the user with the digital assistant or each time the digital assistant is woken up, and receive a user’s configuration of permissions via the authorization request interface. The authorization request interface, for example, can include permission prompt information and at least one operation control. The at least one operation control can include an authorization confirmation control. The permission prompt information, for example, can include text “Allow to periodically collect audio?” and the client device 110 can determine that authorization of the user is obtained in response to detecting a trigger of the authorization confirmation control.

[0053] The client device 110 can determine a sound intensity of the ambient sound 202. It can be appreciated that the client device 110 can employ any suitable manner to determine the sound intensity. The sound intensity, for example, can be associated with a volume, a frequency, a pitch, etc. of the audio. The client device 110, for example, can determine the sound intensity by means of a machine learning model. The client device 110, for example, can convert the ambient sound 202 into a format matching the machine learning model (e.g., encode the ambient sound 202 into an acoustic feature representation), and provide the converted ambient sound 202 to the machine learning model to determine the sound intensity by means of the machine learning model.

[0054] In some embodiments, the client device 110 can determine a type of environment in which the client device 110 is located based on at least the sound intensity of the ambient sound 202. The client device 110, for example, can also determine the type of environment in which the client device 110 is located based on any suitable factor such as a time period in which the ambient sound is collected (e.g., whether it is daytime or nighttime), a location in which the ambient sound is collected, etc.

[0055] In some embodiments, the client device 110 can pre-obtain a correspondence between sound intensities of different ranges and different types of environments. The different types of environments, for example, can be pre-determined by the user and can include, but are not limited to, a quiet type, a meeting type, a noisy type, etc. The correspondence between the sound intensities and the types of environments, for example, can be as shown in Table 1:

[0056] Table 1

[0057] The client device 110, in turn, can determine the type of environment in which the client device 110 is located based on the correspondence between the sound intensities and the types of environments as shown in Table 1 and the determined sound intensity of the ambient sound 202. For example, if the sound intensity of the ambient sound 202 is a, a ∈ [A1, A2), the client device 110 can determine that the type of environment in which the client device 110 is located is Type A based on the correspondence shown in Table 1. In addition to or alternatively to the sound intensity of the ambient sound, the ambient sound can be classified based on other acoustic properties of the ambient sound.

[0058] In some embodiments, the client device 110 can also determine the environment type directly with the help of a machine learning model. For example, the client device 110 can directly provide the environment sound 202 to a target machine learning model (e.g., the machine learning model 220). For another example, the client device 110 can also provide the converted environment sound 202 (e.g., the environment sound converted to a format matching the model) to the machine learning model 220. The machine learning model 220 can be a machine learning model local to the client device 110, or a machine learning model deployed at another electronic device (e.g., the server device 130). In some embodiments, the machine learning model 220 can determine an environment type matching the environment sound collected by the client device 110 from a plurality of predetermined environment types 225 (e.g., can include Type A, Type B, …, Type N, N being any positive integer), and output a model output that can indicate the environment type.

[0059] The client device 110 can obtain a model output of a target machine learning model. The target machine learning model can be based on any appropriate model structure, including but not limited to a Transformer model, a convolutional neural network (CNN), a recurrent neural network (RNN), a deep neural network (DNN), etc. In some embodiments, the machine learning model can be based on a language model (LM). A language model can have a question-answering capability by learning from a large amount of corpus. In some embodiments, a language model used for processing speech (e.g., a language model used for performing speech synthesis TTS, a language model used for performing speech recognition ASR, etc.) can be referred to as a primary language model or a language model, and a language model used for determining an environment type can be referred to as a secondary language model. In this way, the client device 110 can determine the environment type locally, without uploading the environment sound to the server device 130, which can protect user privacy and improve the security of audio processing.

[0060] The client device 110 can control a sensitivity parameter and / or a semantic detection strategy applied to second audio collected by the client device 110 based at least on the environment type. The sensitivity parameter is configured to control whether to perform speech recognition on the second audio, and the semantic detection strategy is applied to a semantic detection process performed on text corresponding to the second audio. The second audio can be, for example, audio collected by the client device 110 in real time in response to a voice recording request. The voice recording request can be, for example, a request generated by a user requesting to record a voice for ASR processing. The voice recording request can also be, for example, a request generated by a user requesting to interact with a digital assistant via voice. The second audio can include the environment sound 202 of the environment in which the client device 110 is located and the voice 145 of the user 140.

[0061] The second audio and the first audio may, for example, be the same audio. In this case, the client device 110 can perform the determination of the environment type while the audio is being captured, and control the subsequent processing of the audio using the determined environment type. The second audio and the first audio may, for example, also be the same audio. In this case, the client device 110 can determine the environment type based on a pre-captured ambient sound, and control the audio processing of the currently captured audio directly based on the environment type. Since determining the environment type and determining the sensitivity parameter and / or the semantic detection strategy based on the environment type consume certain time, pre-capturing the ambient sound can avoid the delay problem caused by determining the environment type based on the real-time captured audio. In addition, the environment in which the client device is located usually does not change frequently over time, and compared to determining the environment type each time the audio is captured, pre-determining the environment type consumes fewer resources and can save the cost required for audio processing.

[0062] In some embodiments, the environment information 232 can be determined based on the environment type. That is, the environment information 232 can include an indication of the environment type. The indication of the environment type may, for example, include a name, a label, a description, etc. of the environment type. In some embodiments, the client device 110 can determine the sensitivity parameter and / or the semantic detection strategy locally based on the environment information 232 and the second audio (e.g. the audio 234).

[0063] Alternatively or additionally, in some embodiments, the client device 110 can also send the environment information 232 and the second audio to the server device 130 for determining the sensitivity parameter and / or the semantic detection strategy by the server device 130. For example, the client device 110 can provide the environment information 232 and the audio 234 to the server voice module 240. The server voice module 240 can be a module in the server device 130 for providing audio processing related services.

[0064] In some embodiments, if the sensitivity parameter and / or the semantic detection strategy are determined by the server device 130, to improve the accuracy of determining the environment type and subsequently determining the sensitivity parameter and / or the semantic detection strategy, the environment information 232 can include the ambient sound 202 of the environment in which the client device 110 is located. In this case, the server device 130 can determine the environment type by itself, and determine the sensitivity parameter and / or the semantic detection strategy based on the environment type. The specific manner in which the server device 130 determines the environment type is similar to the specific manner in which the client device 110 determines the environment type, and will not be described here.

[0065] An example embodiment of determining the sensitivity parameter and / or the semantic detection strategy by the server device 130 is described below with reference to FIG. 2B. FIG. 2B shows a schematic diagram of an example architecture 200B for audio processing, which is implemented at the server device 130, according to some other embodiments of the present disclosure. For ease of discussion, the example architecture 200B is described with reference to the environment 100 of FIG. 1. As shown in FIG. 2B, the example architecture 200B illustrates an example of a server-side speech module 240. The example architecture 200B can involve, for example, an ambient sound information preprocessing module 250, a speech discrimination module 260, an ASR model 265, a semantic discrimination selector 270, and a question-answering model 280. The server-side speech module 240 can obtain the ambient information 232 and the audio 234 (i.e., the second audio). If the ambient information 232 includes an indication of the ambient type, the server-side speech module 240 can determine the ambient type in which the client device 110 is located based on the ambient information 232 directly. If the ambient information 232 includes the first audio, the server-side speech module 240 can determine the ambient type in which the client device 110 is located based on the first audio. The ambient sound information preprocessing module 250 in the server device 120 can determine the corresponding sensitivity parameter and / or the semantic detection strategy based on the ambient type. The ambient sound information preprocessing module 250 can determine the sensitivity parameter and / or the semantic detection strategy in any suitable manner. For example, the ambient sound information preprocessing module 250 can determine the sensitivity parameter and / or the semantic detection strategy based on a rejection adjustment algorithm 252.

[0066] In some embodiments, the ambient types are distinguished based on the noisy level of the ambient. The more noisy the ambient indicated by the ambient type, the higher the sensitivity value of the sensitivity parameter determined by the ambient sound information preprocessing module. That is, if the ambient type is a first ambient type, the ambient sound information preprocessing module 250 can determine the sensitivity parameter as a first sensitivity value. If the ambient type is a second ambient type, the ambient sound information preprocessing module 250 can determine the sensitivity parameter as a second sensitivity value. If the second ambient type indicates a more noisy ambient than the first ambient type, the second sensitivity value is lower than the first sensitivity value.

[0067] In some embodiments, the ambient sound information preprocessing module 250 can determine the sensitivity parameter corresponding to the ambient type directly. The sensitivity parameter can generally be in the interval [0, 1]. It can be appreciated that in a noisy ambient, a lower sensitivity parameter is needed to avoid collecting a large amount of noise (e.g., noise, ambient sound, etc.), and in a quiet ambient, a higher sensitivity parameter is needed to improve the subsequent speech processing instruction.

[0068] Alternatively or additionally, in some embodiments, the ambient sound information preprocessing module 250 can also determine the adjustment information to the predetermined sensitivity parameter based on the environment type. Similarly, the ambient sound information preprocessing module 250 can employ any suitable manner to determine the adjustment information to the predetermined sensitivity parameter. For example, the ambient sound information preprocessing module 250 can determine the adjustment information to the predetermined sensitivity parameter based on a rejection adjustment algorithm 252. The predetermined sensitivity parameter can generally be within the interval [0.5, 0.6]. The adjustment information can be, for example, a delta to the predetermined sensitivity parameter or a scaling factor to the predetermined sensitivity parameter.

[0069] In some embodiments, the noisier the environment indicated by the environment type, the smaller the adjustment information to the predetermined sensitivity parameter determined by the ambient sound information preprocessing module 250. Taking the delta as an example of the adjustment information, the noisier the environment indicated by the environment type, the smaller the value of the adjustment information determined by the ambient sound information preprocessing module. The ambient sound information preprocessing module 250 can adjust the predetermined sensitivity parameter based on the adjustment information to obtain a sensitivity parameter applied to the second audio. For example, if the environment type includes a quiet type and a noisy type, the adjustment information to the predetermined sensitivity parameter corresponding to the quiet type is 0.25, the adjustment information to the predetermined sensitivity parameter corresponding to the noisy type is -0.1, and the sensitivity value of the predetermined sensitivity parameter is 0.5, then the sensitivity value of the sensitivity parameter corresponding to the quiet type is 0.75, and the sensitivity value of the sensitivity parameter corresponding to the noisy type is 0.4.

[0070] In some embodiments, the noisier the environment indicated by the environment type, the lower the semantic pass rate corresponding to the semantic detection strategy. That is, if the environment type is a first environment type, the ambient sound information preprocessing module 250 can determine the semantic detection strategy as a first semantic detection strategy. If the environment type is a second environment type, the ambient sound information preprocessing module 250 can determine the semantic detection strategy as a second semantic detection strategy. If the second environment type indicates a noisier environment than the first environment type, the second semantic pass rate corresponding to the second semantic detection strategy is lower than the first semantic pass rate corresponding to the first semantic detection strategy.

[0071] In some embodiments, the server-side voice module 240 further includes a parameter feedback module 242. The parameter feedback module 242 can obtain feedback information for the sensitivity parameter and / or the semantic detection strategy. The feedback information may, for example, be received by the client device 110 and uploaded to the server device 110. The feedback information may, for example, indicate a user’s 140 satisfaction level (e.g., satisfied or not satisfied) for the sensitivity parameter and / or the semantic detection strategy, a user’s 140 adjustment opinion (e.g., expecting to lower the sensitivity parameter, expecting to select a lenient semantic discrimination model, etc.) for the sensitivity parameter and / or the semantic detection strategy, etc. The parameter feedback module 242 can determine the sensitivity parameter and / or the semantic detection strategy to be applied to the third audio collected by the client device based on the feedback information and the environment type. For example, in a case where the environment type does not change, the parameter feedback module 242 can adjust the sensitivity parameter and / or the semantic detection strategy corresponding to the environment type based on the feedback information, and process the third audio using the adjusted sensitivity parameter and / or the semantic detection strategy. The third audio may, for example, be audio collected by the client device 110 in response to a voice recording request, and be different from the second audio.

[0072] The environment sound information preprocessing module 250 can provide the determined sensitivity parameter to a voice discrimination module 260, and provide the determined semantic detection strategy to a semantic discrimination selector 270. The voice discrimination module 260 can determine an acceptance confidence for voice discrimination (may also be referred to as for acoustic rejection) based on the determined sensitivity parameter. In some embodiments, the voice discrimination module 260 can also determine adjustment information for a predetermined acceptance confidence based on the sensitivity parameter, and adjust the predetermined acceptance confidence based on the adjustment information for the predetermined acceptance confidence. It can be understood that the higher the sensitivity parameter, the smaller the adjustment information for the predetermined acceptance confidence.

[0073] As mentioned above, the acceptance confidence is inversely proportional to the sensitivity parameter, and the larger the sensitivity parameter, the smaller the acceptance confidence. For example, if the environment types include a quiet type and a noisy type, the sensitivity value of the sensitivity parameter corresponding to the quiet type is 0.75, and the sensitivity value of the sensitivity parameter corresponding to the noisy type is 0.4, then the acceptance confidence corresponding to the quiet type can be 0.3, and the acceptance confidence corresponding to the noisy type can be 0.7. Of course, these are only some examples, and the specific values can be adjusted according to actual applications.

[0074] The speech discrimination module 260 can process the audio 234 based on the determined acceptance confidence. Referring to FIG. 2C, FIG. 2C illustrates an example spectrogram 200C of ambient sound and user speech according to the environment of the present disclosure. The example spectrogram 200C includes a spectrogram 292 of ambient sound and a spectrogram 294 of user speech. It can be found that the intensity of the user speech is higher than that of the ambient sound. Thus, the speech discrimination module 260 can screen the audio 234 based on the acceptance confidence. Illustratively, the speech discrimination module 260 can determine a portion corresponding to an intensity higher than the acceptance confidence from the audio 234, and determine the portion as the speech 145 of the user 140. The speech discrimination module 260 can provide the speech 145 to the ASR model 265. The ASR model 265 may, for example, perform speech recognition on the speech 145 to output text corresponding to the speech 145.

[0075] The text corresponding to the speech 145 is provided to the semantic discrimination selector 270. In some embodiments, the server speech module 240 can also provide the text corresponding to the speech 145 to the client device 110 to be shown to the user 140 by the client device 110. In this way, the user 140 can check the text corresponding to the speech to determine whether an error of recognition is made.

[0076] The semantic discrimination selector 270 can determine a semantic pass rate matching the semantic detection strategy based on the determined semantic detection strategy, and determine a semantic discrimination model matching the semantic pass rate from a plurality of semantic discrimination models (e.g., which can include a strict semantic discrimination model 272, a normal semantic discrimination model 274, a lenient semantic discrimination model 276, etc., and it can be understood that the plurality of semantic discrimination models can include more or less semantic discrimination models) based on the semantic pass rate. Alternatively or additionally, in some embodiments, the semantic discrimination selector 270 can also directly determine a semantic discrimination model matching the semantic detection strategy from the plurality of semantic discrimination models based on the semantic detection strategy.

[0077] In some embodiments, the semantic discrimination selector 270 can pre-acquire a correspondence between different semantic detection strategies and different semantic discrimination models. The different semantic detection strategies can be pre-determined by the user, and can include, but are not limited to, a semantic detection strategy corresponding to a quiet type of environment, a semantic detection strategy corresponding to a conference type of environment, a semantic detection strategy corresponding to a noisy type of environment, etc. The correspondence between the semantic detection strategy and the semantic discrimination model may, for example, be as shown in Table 2:

[0078] Table 2

[0079] Referring to Table 2, the semantic discrimination selector 270 may, for example, determine that the speech is less interfered in response to determining that the semantic detection strategy corresponds to a semantic detection strategy for a quiet type of environment, and may select a loose semantic discrimination model in response to determining that loose detection is possible. The semantic discrimination selector 270 may, for example, determine that the speech is more interfered in response to determining that the semantic detection strategy corresponds to a semantic detection strategy for a conference type of environment or a semantic detection strategy for a noisy type of environment, and may select a strict semantic discrimination model in response to determining that strict detection is required to process the speech only after a clear semantic is detected.

[0080] The semantic discrimination selector 270 may provide the text corresponding to the speech 145 to the determined semantic discrimination model, and obtain a discrimination result of the semantic discrimination model on the speech 145. The discrimination result may, for example, indicate whether the speech 145 has an explicit question demand. If the speech 145 has an explicit question demand, it may be determined that the speech 145 passes the discrimination.

[0081] If the speech 145 passes the discrimination, the server-side speech module 240 may provide the text of the speech 145 to the question-answering model 280 to determine, with the help of the question-answering model 280, a response text corresponding to the text of the speech 145. The server-side speech module 240 may provide the response text to the client device 110 to display the response text to the user 140 by the client device 110.

[0082] In some embodiments, the server-side speech module 240 may also provide the response text to a TTS model to obtain, with the help of the TTS model, a response speech corresponding to the response text. The server-side speech module 240 may provide the response speech to the client device 110 to play the response speech to the user 140 by the client device 110.

[0083] It should be noted that the ASR model 265 / question-answering model 280 / TTS model may include an ASR model 265 / question-answering model 280 / TTS model deployed on the server-side device 130 and an ASR model 265 / question-answering model 280 / TTS model deployed on the client device 110. The server-side device 130 may, for example, by default, determine the text corresponding to the speech 145 / response text / response speech by using the ASR model 265 / question-answering model 280 / TTS model deployed at the server-side device 130 in the manner described above. If an exception occurs (for example, the network state is poor, the model at the server-side device 130 is abnormal, etc.), the server-side device 130 may also send the determined sensitivity parameter and / or semantic detection strategy to the client device 110 to determine the text corresponding to the speech 145 / response text / response speech by the client device 110 with the help of the local ASR model 265 / question-answering model 280 / TTS model.

[0084] As mentioned previously, in some embodiments, the sensitivity parameter and / or the semantic detection strategy can also be determined by the client device 110 itself. In this case, the client device 110 may, for example, determine the sensitivity parameter and / or the semantic detection strategy in a similar manner. If a response to the speech 145 is to be determined, the client device 110 may, for example, provide the determined sensitivity parameter and / or the semantic detection strategy to the server device 110. The server device 110 can process the audio based on the sensitivity parameter and / or the semantic detection strategy, and provide the processed audio to the question-answering model and the TTS model to determine the response text to the speech 145 and the audio corresponding to the response text.

[0085] In summary, according to embodiments of the present disclosure, the sensitivity parameter and / or the semantic detection strategy applied to processing of the audio can be determined based on the type of environment in which the client device is located. The accuracy of speech recognition can be improved, and in turn, the efficiency of audio processing can be improved.

[0086] FIG. 3 illustrates a flowchart of a method 300 for audio processing, according to some embodiments of the present disclosure. The method 300 can be implemented at the client device 110 and / or the server device 130. The method 300 will be described with reference to the environment 100 of FIG. 1.

[0087] At block 310, the client device 110 and / or the server device 130 determines the type of environment in which the client device 110 is located from first audio collected by the client device 110.

[0088] At block 320, the client device 110 and / or the server device 130 controls a sensitivity parameter and / or a semantic detection strategy applied to second audio collected by the client device 110 based at least on the type of environment, the sensitivity parameter being configured to control whether speech recognition is performed on the second audio, the semantic detection strategy being applied to a semantic detection process performed on text corresponding to the second audio.

[0089] In some embodiments, the second audio is collected by the client device 110 in response to a speech recording request.

[0090] In some embodiments, the method 300 is implemented at the client device 110, and wherein controlling the sensitivity parameter and / or the semantic detection strategy applied to the second audio collected by the client device 110 comprises sending an indication of the type of environment and the second audio to the server device 130 for determination of the sensitivity parameter and / or the semantic detection strategy by the server device 130.

[0091] In some embodiments, determining the type of environment in which the client device 110 is located comprises extracting an environmental sound from the first audio, and determining the type of environment in which the client device 110 is located based at least on a sound intensity of the environmental sound.

[0092] In some embodiments, determining the environment type in which the client device 110 is located includes: providing the first audio to a target machine learning model; and obtaining a model output of the target machine learning model, the model output indicating the environment type.

[0093] In some embodiments, controlling the sensitivity parameter applied to the second audio collected by the client device 110 includes: determining the sensitivity parameter as a first sensitivity value if the environment type is a first environment type; and determining the sensitivity parameter as a second sensitivity value if the environment type is a second environment type, the second environment type indicating a noisier environment than the first environment type, and the second sensitivity value being lower than the first sensitivity value.

[0094] In some embodiments, controlling the semantic detection strategy applied to the second audio collected by the client device 110 includes: determining the semantic detection strategy as a first semantic detection strategy if the environment type is a first environment type; and determining the semantic detection strategy as a second semantic detection strategy if the environment type is a second environment type, the second environment type indicating a noisier environment than the first environment type, and the second semantic detection strategy corresponding to a second semantic pass rate lower than a first semantic pass rate corresponding to the first semantic detection strategy.

[0095] In some embodiments, controlling the sensitivity parameter applied to the second audio collected by the client device 110 includes: determining adjustment information to a predetermined sensitivity parameter based on the environment type; and adjusting the predetermined sensitivity parameter based on the adjustment information to obtain the sensitivity parameter applied to the second audio.

[0096] In some embodiments, the method 300 further includes: obtaining feedback information for the sensitivity parameter and / or the semantic detection strategy; and determining the sensitivity parameter and / or the semantic detection strategy applied to the third audio collected by the client device 110 based on the feedback information and the environment type.

[0097] Embodiments of the present disclosure also provide corresponding apparatuses for implementing the above-described methods or processes. FIG. 4 shows an exemplary structural block diagram of an apparatus 400 for audio processing according to some embodiments of the present disclosure. The apparatus 400 can be implemented as or included in the client device 110 and / or the server device 130. Various modules / components in the apparatus 400 can be implemented by hardware, software, firmware, or any combination thereof.

[0098] As shown in FIG. 4, the apparatus 400 includes an environment type determination module 410 configured to determine, from first audio collected by the client device, an environment type in which the client device is located. The apparatus 400 further includes a parameter policy control module 420 configured to control, based at least on the environment type, a sensitivity parameter and / or a semantic detection policy applied to second audio collected by the client device, the sensitivity parameter configured to control whether to perform speech recognition on the second audio, and the semantic detection policy applied to a semantic detection process performed on text corresponding to the second audio.

[0099] In some embodiments, the second audio is collected by the client device 110 in response to a voice recording request.

[0100] In some embodiments, the apparatus 400 is implemented at the client device 110, and the parameter policy control module 420 is further configured to send, to the server device 130, an indication of the environment type and the second audio, for determining, by the server device 130, the sensitivity parameter and / or the semantic detection policy.

[0101] In some embodiments, the environment type determination module 410 is further configured to extract, from the first audio, an environmental sound; and determine, based at least on a sound intensity of the environmental sound, the environment type in which the client device 110 is located.

[0102] In some embodiments, the environment type determination module 410 includes a providing module configured to provide the first audio to a target machine learning model; and a model output obtaining module configured to obtain a model output of the target machine learning model, the model output indicating the environment type.

[0103] In some embodiments, the parameter policy control module 420 includes a first determination module configured to determine, if the environment type is a first environment type, the sensitivity parameter as a first sensitivity value; and a second determination module configured to determine, if the environment type is a second environment type, the sensitivity parameter as a second sensitivity value, the second environment type indicating a noisier environment than the first environment type, and the second sensitivity value being lower than the first sensitivity value.

[0104] In some embodiments, the parameter policy control module 420 includes a third determination module configured to determine, if the environment type is a first environment type, the semantic detection policy as a first semantic detection policy; and a fourth determination module configured to determine, if the environment type is a second environment type, the semantic detection policy as a second semantic detection policy, the second environment type indicating a noisier environment than the first environment type, and a second semantic pass rate corresponding to the second semantic detection policy being lower than a first semantic pass rate corresponding to the first semantic detection policy.

[0105] In some embodiments, the parameter strategy control module 420 comprises an adjustment information determination module configured to determine adjustment information for the predetermined sensitivity parameter based on the environment type; and an adjustment module configured to adjust the predetermined sensitivity parameter based on the adjustment information to obtain the sensitivity parameter applied to the second audio.

[0106] In some embodiments, the apparatus 400 further comprises a feedback information obtaining module configured to obtain feedback information for the sensitivity parameter and / or the semantic detection strategy; and a fifth determination module configured to determine, based on the feedback information and the environment type, the sensitivity parameter and / or the semantic detection strategy applied to the third audio collected by the client device 110.

[0107] The units and / or modules included in the apparatus 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, e.g., machine executable instructions stored on a storage medium. In addition to or alternatively, some or all of the units and / or modules in the apparatus 400 can be implemented at least partially by one or more hardware logic components. As an example and not by way of limitation, example types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SOCs), complex programmable logic devices (CPLDs), etc.

[0108] It should be understood that one or more steps in the above methods can be performed by an appropriate electronic device or combination of electronic devices. Such an electronic device or combination of electronic devices can include, for example, the client device 110 and / or the server device 130 in FIG. 1.

[0109] FIG. 5 shows a block diagram of an electronic device 500 in which one or more embodiments of the present disclosure can be implemented. It should be understood that the electronic device 500 illustrated in FIG. 5 is merely an example and should not be construed to limit the functionality and scope of the embodiments described herein. The electronic device 500 illustrated in FIG. 5 can be used to implement the client device 110 and / or the server device 130 of FIG. 1.

[0110] As shown in FIG. 5, electronic device 500 is in the form of a general-purpose electronic device. Components of electronic device 500 can include, but are not limited to, one or more processors or processing units 510, memory 520, storage 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processing unit(s) 510 can be actual or virtual processors and capable of executing various processing in accordance with programs stored in memory 520. In a multi-processing system, multiple processing units execute computer-executable instructions in parallel to improve the processing power of electronic device 500.

[0111] Electronic device 500 typically includes a plurality of computer storage media. Such media can be removable and / or non-removable, and can include volatile and / or nonvolatile media. Memory 520 can be volatile (such as, for example, registers, cache, random access memory (RAM)), non-volatile (such as, for example, read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage 530 can be removable or non-removable and can include machine-readable media, such as, for example, flash drives, disks, or any other media capable of storing information and / or data and accessible by electronic device 500.

[0112] Electronic device 500 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 5, a disk drive or other computer-readable media drive can be provided for reading from or writing to a removable, non- volatile magnetic disk (e.g., a "hard drive"), and a disk drive or other computer-readable media drive can be provided for reading from or writing to a removable, non-volatile optical disk (such as a CD-ROM or other optical medium). In these instances, each drive can be connected to the bus (not shown) by one or more data media interfaces. Memory 520 can include a computer program product 525 having one or more program modules configured to carry out the various methods or actions of the various embodiments of the present disclosure.

[0113] Communication unit(s) 540 enable communication with other electronic devices via communication media. Additionally, functionality of components of electronic device 500 can be implemented in a single computing cluster or a plurality of computer machines capable of communication through a communication connection. Accordingly, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes.

[0114] The input device 550 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 560 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 500 can also communicate with one or more external devices (not shown), such as storage devices, display devices, etc., one or more devices that enable a user to interact with the electronic device 500, or any devices (e.g., a network card, a modem, etc.) that enable the electronic device 500 to communicate with one or more other electronic devices, as desired via the communication unit 540. Such communication can be carried out via an input / output (I / O) interface (not shown).

[0115] According to an example implementation of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above.

[0116] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0117] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0118] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0119] The computer program product of the present disclosure can be a computer program product, which is a machine-readable medium (media) having instances of the software embodied thereon, such as computer software, firmware, wireless application protocol (WAP), middleware or microcode. For example, a computer program product can be a floppy disk, a CD-ROM, a DVD, a Blu-ray Disc™, a flash drive, a memory stick, a magnetic tape, or a hard disk drive. The machine-readable medium can be a single medium, or multiple media, of the same or different type. The computer program product can be one or more computer program components embodied in medium and / or transmission signals. The computer program product can have one or more computer program components embodied in medium and / or transmission signals.

[0120] The implementations of the disclosure have been described above with the intent to be illustrative rather than limiting. Although the implementations of the disclosure have been described with regard to one or more implementations, it will be recognized that a variety of modifications and changes can be made to these implementations without departing from the broader spirit and scope of the implementations as set forth in the preceding disclosure. For example, certain aspects of the implementations can be performed using hardware, software, and / or firmware, or any combination thereof. The above-described implementations should therefore be regarded as merely illustrative, and not as narrowing the scope of the disclosure, which is defined by the appended claims and their equivalents.

Claims

1. A method for audio processing, comprising: determining, from a first audio captured by a client device, a type of environment in which the client device is located; and controlling, based at least on the type of environment, a sensitivity parameter and / or a semantic detection strategy applied to a second audio captured by the client device, the sensitivity parameter being configured to control whether to perform speech recognition on the second audio, the semantic detection strategy being applied to a semantic detection process performed on text corresponding to the second audio.

2. The method of claim 1, wherein the second audio is captured by the client device in response to a voice recording request.

3. The method of claim 1, wherein the method is implemented at the client device, and wherein controlling the sensitivity parameter and / or the semantic detection strategy applied to the second audio captured by the client device comprises: sending, to a server device, an indication of the type of environment and the second audio, to determine, by the server device, the sensitivity parameter and / or the semantic detection strategy.

4. The method of claim 1, wherein determining the type of environment in which the client device is located comprises: extracting, from the first audio, an environmental sound; and determining, based at least on a sound intensity of the environmental sound, the type of environment in which the client device is located.

5. The method of claim 1, wherein determining the type of environment in which the client device is located comprises: providing the first audio to a target machine learning model; and obtaining a model output of the target machine learning model, the model output indicating the type of environment.

6. The method of claim 1, wherein controlling the sensitivity parameter applied to the second audio captured by the client device comprises: determining the sensitivity parameter as a first sensitivity value if the type of environment is a first type of environment; and determining the sensitivity parameter as a second sensitivity value if the type of environment is a second type of environment, the second type of environment indicating a noisier environment than the first type of environment, and the second sensitivity value being lower than the first sensitivity value.

7. The method of claim 1, wherein controlling the semantic detection strategy applied to the second audio captured by the client device comprises: determining the semantic detection strategy as a first semantic detection strategy if the type of environment is a first type of environment; and determining the semantic detection strategy as a second semantic detection strategy if the type of environment is a second type of environment, the second type of environment indicating a noisier environment than the first type of environment, and a second semantic pass rate corresponding to the second semantic detection strategy being lower than a first semantic pass rate corresponding to the first semantic detection strategy.

8. The method of claim 1, wherein controlling the sensitivity parameter applied to the second audio captured by the client device comprises: determining, based on the type of environment, adjustment information to a predetermined sensitivity parameter; and adjusting the predetermined sensitivity parameter based on the adjustment information to obtain the sensitivity parameter applied to the second audio. ​ ​ ​ ​ ​ 9. The method of claim 1, further comprising: obtaining feedback information for the sensitivity parameter and / or the semantic detection strategy; and determining, based on the feedback information and the environment type, a sensitivity parameter and / or a semantic detection strategy to be applied to third audio captured by the client device.

10. An apparatus for audio processing, comprising: an environment type determination module configured to determine, from first audio captured by a client device, an environment type in which the client device is located; and a parameter strategy control module configured to control, based on at least the environment type, a sensitivity parameter and / or a semantic detection strategy to be applied to second audio captured by the client device, the sensitivity parameter configured to control whether speech recognition is performed on the second audio, the semantic detection strategy applied to a semantic detection process performed on text corresponding to the second audio.

11. An electronic device, comprising: at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, cause the electronic device to perform the method according to any one of claims 1-9.

12. A computer-readable storage medium having computer-executable instructions stored thereon that are executable by a processor to implement the method according to any one of claims 1-9.

13. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1-9. ​

Citation Information

Patent Citations

  • Control method and device and electronic device

    CN110197663A

  • Electronic apparatus and control method thereof

    CN111433737A

  • Speech recognition method and device, storage medium and electronic device

    CN116389179A

  • Speech recognition environment judging method

    JP2004219918A

  • Systems and methods for dynamically improving user intelligibility of synthesized speech in a work environment

    US20120296654A1