Improve the audio quality of speech in a sound system

By performing voice recognition on the input and output audio data of the sound system, comparing differences and taking corrective measures, the problem of poor voice quality in the sound system is solved, and the effect of automatically improving voice quality is achieved.

CN114667568BActive Publication Date: 2025-05-27INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202080071083.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-10
Filing Date
2020-10-05
Publication Date
2025-05-27
Estimated Expiration
2040-10-05

AI Technical Summary

Technical Problem

When playing voice, existing sound systems often lead to poor voice quality, such as unclear or distortion, which makes it difficult for listeners to hear or understand.

Method used

By performing speech recognition on the input audio data and the output audio data, the difference between the two is compared, if the difference is greater than or equal to the threshold, the reproduced speech quality is determined to be unqualified, and correction actions are taken, such as adjusting parameters of the sound system or issuing instructions to the user to improve speech quality.

Benefits of technology

Automatic and real-time detection of when the voice quality reproduced by the sound system fails, and measures are taken to improve the voice quality, thereby improving the listener's experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114667568B_ABST
    Figure CN114667568B_ABST
Patent Text Reader

Abstract

A computer-implemented method, apparatus, and computer program product for a sound system. Speech recognition is performed on input audio data including a speech input to the sound system. Additionally, speech recognition is performed on at least one instance of output audio data including speech reproduced by one or more audio speakers of the sound system. A difference between a result of the speech recognition performed on the input audio data and a result of the speech recognition performed on the corresponding instance of the output audio data is determined. When the difference is greater than or equal to a threshold, the quality of the reproduced speech is determined to be unacceptable. If it is determined that the speech quality of the reproduced sound is unacceptable, a corrective action may be performed to improve the quality of the speech reproduced by the sound system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to techniques for improving the quality of audio output from a sound system, and more particularly to improving the speech quality of a listener with respect to the audio output of a sound system. Background Art

[0002] Sound systems are often used to play speech to a listener via an audio speaker, e.g., attendees at a conference, lecture, and performance in a theater or auditorium, or participants in a conference call and web conference at distributed geographical locations over a communication network. In such a system, input speech to a microphone is received and optionally recorded by a host system, audio data is transmitted by the host system to one or more audio speakers, and the (multiple) audio speakers output (i.e., "play") the reproduced speech to the listener. In many cases, the reproduced speech played via the audio speaker is not a perfect reproduction of the input speech (e.g., the speech may be unclear). For example, if the settings of the audio speaker are not optimized, the reproduced sound and thus the reproduced speech may be distorted, making it difficult for the listener to hear and / or understand. In other cases, the input speech itself may be imperfect, e.g., due to the position of the microphone relative to the speech source or its suboptimal settings. This again makes it difficult for the listener to hear or understand the speech being played by the audio speaker. Generally, these problems associated with the clarity of the speech of the audio output can be solved by adjusting the sound system. For example, if the listener notifies the host that the speech is difficult to hear or understand, then the host can adjust the configurable settings of the sound system or ask the human speaker to move relative to the microphone. However, when making adjustments, this results in interruptions and delays. In addition, since the adjustments are manual, they may not fully solve the listener's difficulties. Summary of the Invention

[0003] According to one aspect of the present invention, a computer-implemented method is provided. The computer-implemented method includes performing speech recognition on input audio data, the input audio data including a speech input to a sound system. The computer-implemented method further includes performing speech recognition on at least one instance of output audio data, the output audio data including speech reproduced by one or more audio speakers of the sound system. The computer-implemented method further includes determining a difference between a result of the speech recognition of the input audio data and a result of the speech recognition of the at least one instance of the output audio data. The computer-implemented method further includes determining that the quality of the reproduced speech is unacceptable when the difference is greater than or equal to a threshold.

[0004] According to another aspect of the present invention, there is provided a device. The device includes a processor and a data storage device. The processor is configured to perform speech recognition on input audio data, the input audio data including a voice input to a sound system. The processor is further configured to perform speech recognition on at least one instance of output audio data, the output audio data including speech reproduced by one or more audio speakers of the sound system. The processor is further configured to determine a difference between a result of the speech recognition on the input audio data and a result of the speech recognition on at least one instance of the output audio data. The processor is further configured to determine that the quality of the reproduced speech is unacceptable when the difference is greater than or equal to a threshold.

[0005] According to yet another aspect of the present invention, there is provided a computer program product. The computer program product includes a computer-readable storage medium embodied with program instructions. The program instructions are executable by a processor to cause the processor to: perform speech recognition on input audio data, the input audio data including a voice input to a sound system; perform speech recognition on at least one instance of output audio data, the output audio data including speech reproduced by one or more audio speakers of the sound system; determine a difference between a result of the speech recognition on the input audio data and a result of the speech recognition on at least one instance of the output audio data; and determine that the quality of the reproduced speech is unacceptable when the difference is greater than or equal to a threshold. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Exemplary embodiments of the present disclosure will be described below with reference to the following drawings.

[0007] Figure 1 is a schematic diagram showing a sound system according to an embodiment of the present invention.

[0008] Figure 2 is a flowchart of a method for detecting and correcting unacceptable speech quality according to an embodiment of the present invention.

[0009] Figure 3 is a flowchart of a method for adjusting a sound system to improve speech quality according to an embodiment of the present invention.

[0010] Figure 4 is a block diagram of a sound system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0011] The present invention provides systems and methods for detecting when the quality of speech reproduced by a sound system is unacceptable (e.g., difficult to hear or unclear / incoherent for a listener) and for reconfiguring the sound system to improve the quality of the reproduced speech. The techniques of the present disclosure can be performed automatically and in real time to limit interruptions and improve the listener's experience.

[0012] Specifically, according to the present disclosure, one or more microphones located at positions within a listening environment are used to detect reproduced speech played by one or more audio speakers of a sound system. Output audio data including the reproduced speech received by each of the one or more microphones is recorded. Speech recognition is performed on the output audio data associated with each of the one or more microphones to determine the quality of the speech being played at the corresponding microphone position. In addition, speech recognition is performed on input audio data including a voice input to the sound system to determine the quality of the speech from the source. A comparison is performed between the result of the speech recognition performed on the input audio data and the result of the speech recognition performed on the corresponding output audio data for each microphone. The result of the comparison is used to determine whether the speech quality is unacceptable to a listener, and if so, corrective action is taken, such as adjusting the sound system to improve the quality of the reproduced speech.

[0013] In the present disclosure, the term "speech" is used to refer to sounds or audio containing speech. The term "input audio data" refers to digital audio data of sounds or audio containing speech from a source (e.g., a human speaker) detected by a microphone of the sound system (herein "input microphone"). The term "output audio data" refers to digital audio data of sounds or audio containing speech reproduced by one or more audio speakers of the sound system and detected by a microphone of the sound system (herein "output microphone"). Thus, the audio data "represents" or "includes" sounds or audio containing speech received by a microphone of the sound system. A reference to "recording" the audio data refers to storing the audio data in a data storage device, which includes transient storage of the audio data for communication over a network and temporary and long-term data storage of the audio data as an audio data file.

[0014] Figure 1 is a schematic diagram showing a sound system according to an embodiment of the present invention. The sound system 100 includes a host processing system 110, a plurality of microphones 120, and a plurality of audio speakers 130 interconnected via a data communication network 140. In Figure 1 the system shown, the sound system 100 is a distributed system including microphones 120 and audio speakers 130 located at different geographical locations (e.g., a conference or negotiation room). At least one location (Location 1) includes the host processing system 110, which in the example shown is the location of the source of the speech input to the sound system 100. Other locations (Location 2 and Location 3) include audio speakers 130 for playing audio to listeners in the corresponding listening environments.

[0015] The host processing system 110 generally includes a user computing system (e.g., a notebook computer), a dedicated sound system controller, etc., which can be operated by a user to manage the sound system 100. The plurality of microphones 120 includes an input microphone 122, which is used to detect and record speech from a source (e.g., a human speaker) for reproduction by the sound system 100. The input microphone 122 can be a dedicated microphone (e.g., a microphone on a podium) for receiving input sound to the sound system 100, or can be the microphone of the user computing system that is "turned on" under the control of the host processing system 110. The plurality of audio speakers 130 reproduce the input speech and are distributed at different positions in one or more positions (Position 2 and Position 3) that form the listening environment. Specifically, each audio speaker 130 receives and plays the audio data corresponding to the recorded speech from the host processing system 110 via the network 140. The audio speakers 130 can include one or more dedicated speakers 134 of the sound system at the position (e.g., speakers at fixed positions in a theater) or the audio speakers 132 of the user computing system, a networked phone, etc. The communication network 140 can include any suitable wired or wireless network for data communication between the host processing system 110, the microphones 120, and the audio speakers 130.

[0016] According to one embodiment, the plurality of microphones 120 further includes output microphones 124 located at positions within the listening environment (Position 2 and Position 3) for receiving and recording the reproduced speech for analysis, as described herein. The output microphones 124 can include, for example, dedicated microphones of the sound system associated with the audio speakers 134 at the position. The output microphones 124 can also include the microphones of the user computing system or other devices present within the listening environment, which can be identified and utilized by the host processing system 110 for the purpose. In Figure 1 the distributed system, the output microphones 124 within one listening environment (Position 3) are each associated with a local processing system 150 (such as a system device) for recording the reproduced speech as output audio data for transmission to the host processing system 110 via the network 140. The output microphones 124 within another listening environment (Position 2) are configured to transmit the audio data to the host processing system 110 that records the output audio data via the network 140. As will be understood by those skilled in the art, the plurality of microphones 120 can include digital microphones that generate digital output signals and / or analog microphones that generate analog output signals, and the analog output signals are received by another component along the audio signal chain and converted into digital audio data.

[0017] The host processing system 110 records the voice received from the source by the input microphone 122 as input audio data. In addition, the host processing system 110 receives output audio data corresponding to the reproduced voice played by the audio speaker 130, which is recorded by each output microphone 124 and transmitted through the network 140. According to an embodiment of the present invention, the host processing system 110 is configured to perform speech recognition on the input audio data associated with the input microphone 122 and output the audio data associated with each output microphone 124. Speech recognition techniques are known in the art, and the host processing system 110 can implement any suitable speech recognition technique. Speech recognition can generate a transcript of the speech and / or a value or level of a confidence metric, etc., indicating the reliability of the speech recognition. Speech recognition can be performed continuously or periodically on the input audio data and the corresponding output audio data. The host processing system 110 also compares the results of the speech recognition determined for the output audio data associated with each output microphone 124 with the results of the speech recognition determined for the corresponding input audio data associated with the input microphone 122. If the comparison determines an unacceptable difference between the results determined for the output audio data of one or more output microphones 124 and the results determined for the input audio data of the input microphone 122, the host processing system 110 determines that the speech quality of the reproduced voice is unacceptable for the listener and takes corrective action. For example, the corrective action can include adjusting the parameters of the components of the sound system (e.g., the gain or channel equalization settings of the sound card controlling the audio speaker), or sending a message to the user to take a certain action (e.g., an instruction for a human speaker to move closer to or further away from the input microphone). For example, if the difference in the comparison results obtained from the speech recognition (such as the difference in the confidence level or the measured difference in the speech transcript) is less than or equal to a threshold, an unacceptable difference can be determined, as further described below.

[0018] Accordingly, the host processing system 110 is capable of detecting when the quality of the voice reproduced by the sound system 100 is unacceptable for the listener (i.e., unclear, distorted, or too quiet, etc.) and taking action to improve the voice quality. Since the disclosed technique can be performed automatically and in real time, the listener experience is improved. In Figure 1 the illustrated embodiment, the disclosed technique is performed in the host processing system 110. As will be understood by those skilled in the art, the present invention can be implemented in any other processing device or system that communicates with the sound system 100, such as the local processing system 150 or a combination of processing devices.

[0019] Figure 2 is a flowchart of a method 200 for detecting and correcting unacceptable voice quality according to an embodiment of the present invention. For example, the method 200 can be performed by Figure 1is executed by the host processing system 110 of the audio system 100.

[0020] Method 200 begins at step 205. For example, step 205 can be initiated in response to the start of a sound check of the audio system, in response to the start of a conversation, or otherwise.

[0021] At step 210, the audio system receives input audio data of a voice input to the audio system from a source. For example, in response to the speech of a human speaker talking into the input microphone 122, Figure 1 input audio data can be received from the input microphone 122 of the audio system 100. For example, Figure 1 the host processing system 110 of the audio system 100 generally receives and records (e.g., stores in an input audio data file) the input audio data substantially in real time (i.e., with minimal latency).

[0022] At step 220, the audio system performs speech recognition on the input audio data to determine an input audio speech recognition result indicating the speech quality. Any suitable speech recognition technology or algorithm can be used to perform the speech recognition. The result of the speech recognition typically includes a transcript of the speech contained in the relevant audio data. The quality of the transcript can indicate the quality of the speech. In addition, the result of the speech recognition can include a value or level of a confidence metric indicating the reliability of the speech recognition, etc. Such confidence metrics are known in the field of speech recognition. Thus, the confidence level can also indicate the quality of the speech. The speech recognition can provide other results indicating the quality of the speech.

[0023] At step 230, the audio system receives output audio data for the reproduced speech to be played by the audio speakers of the audio system at one or more locations within the listening environment. For example, an instance of the output audio data can be received from each of one or more output microphones 124 that detect the reproduced speech played by the audio speakers 130 of the audio system 100 as shown. The output audio data is generally received substantially in real time, but must be delayed relative to the input audio data. For example, the delay is caused by the communication of the input audio data to the audio speakers and the communication of the corresponding output audio data associated with the output microphones 124 via Figure 1 the network 140 as shown. In some embodiments, the output audio data is recorded, for example, by Figure 1 the local processing system 150 and / or the host processing system 110 as shown. Figure 1 in the audio system 100.

[0024] In step 240, the sound system performs speech recognition on the output audio data to determine an output audio speech recognition result indicative of the speech quality. Specifically, in step 240, the sound system performs speech recognition on each received instance of the output audio data. In step 240, the sound system utilizes the speech recognition technology used in step 220 such that the results in each of step 220 and step 240 are comparable.

[0025] Accordingly, in steps 210 and 220, the sound system derives (a) speech recognition result(s) for the input audio data, and in steps 230 and 240, the sound system performs to derive (a) speech recognition result(s) for each instance of the output audio. In each case, the speech recognition result includes a speech transcript and / or a confidence level, etc. As will be understood by those skilled in the art, in practice, steps 210 to 240 may be performed simultaneously, particularly in applications where input audio data and output audio data are received and processed continuously, such as in real time.

[0026] In Figure 2 an example embodiment of, a plurality of instances of the output audio data are received. Specifically, each instance of the output audio data is associated with a particular output microphone located within the listening environment. In step 250, the sound system selects the speech recognition result for the first instance of the output audio data.

[0027] In step 260, the sound system compares the selected speech recognition result for an instance of the output audio data with the speech recognition result for the corresponding input audio data and determines the difference in speech quality. The difference is a quantitative value representing the difference in speech quality calculated using the (multiple) speech recognition results. In one implementation, in step 260, the sound system may compare the text of the transcript of the speech recognition determined for an instance of the output audio data with the text of the transcript of the corresponding input audio data and may determine the difference as, for example, the raw numerical difference or percentage difference in the transcribed text (e.g., words). The difference in the text of the transcribed speech indicates a degradation in the speech quality at the audio speaker compared to the speech quality at the source, and the amount of the difference indicates the amount of degradation in quality. In another embodiment, at step 260, the sound system may compare the confidence level of the speech recognition determined for an instance of the output audio data with the confidence level determined for the corresponding input audio data and may determine the difference. As described above, the confidence level indicates the reliability of the transcribed speech (e.g., the value of a confidence metric expressed as a percentage). Since the reliability of the transcribed speech depends on the quality of the reproduced speech for the listener, the difference in the confidence levels indicates a degradation in the quality of the reproduced speech at the audio speaker compared to the speech quality at the source, and the amount of the difference indicates the amount of degradation in quality. In some other implementations, the sound system may compare other metrics derived from the (multiple) speech recognition results to detect a degradation in speech quality. As will be understood by those skilled in the art, in step 260, the sound system may compare the speech recognition result for a sample of the output audio data with the speech recognition result for the corresponding sample of the input audio data. In some scenarios, steps 210 to 240 may be performed continuously (e.g., substantially in real time) on the input speech and the reproduced speech. In such cases, any suitable technique may be used to identify the corresponding samples of the input audio data and the output audio data that include the same input speech and reproduced speech, such as one or more of time synchronization and audio matching (identifying the same portion of the audio in the audio data) or speech recognition transcript matching (identifying the same portion of the text by matching words and phrases in the transcribed text). In some other cases, steps 210 to 240 may be performed by periodically sampling the input speech and the reproduced speech (e.g., sampling the input speech and the reproduced speech on synchronized, temporally separated time windows) such that the (multiple) speech recognition results are associated with the corresponding samples of the input audio data and the output audio data.

[0028] In step 270, the sound system determines whether the difference is greater than or equal to a threshold. The threshold is the value of a difference metric (e.g., the number / percentage of single instances in the text or a confidence metric) that represents an unacceptable degradation in the speech quality of the reproduced speech compared to the quality of the input speech. The value of the threshold can be selected according to application requirements or can be changed by the user. For example, in some applications, a difference in speech quality of up to 5% may be acceptable and thus the threshold is set to 5%, while in some other applications, a difference in speech quality of up to 10% may be acceptable and thus the threshold is set to 10%. In some example embodiments, the value of the threshold can be adjusted based on the quality of the input speech, as described below.

[0029] If the difference is less than the threshold (the "no" branch of step 270), then the quality of the speech reproduced in the selected instance of the output audio data is acceptable, and the sound system performs step 280. In step 280, the sound system determines whether there are more instances of the output audio data to consider. If there are more instances of the output audio data to consider (the "yes" branch of step 280), then the sound system begins to perform step 250 and then continues steps 260 and 270 in a loop until the sound system determines in step 280 that there are no more instances of the output audio data to consider. After determining that there are no more instances of the output audio data to consider, the sound system ends execution in step 295.

[0030] If the difference is greater than or equal to the threshold (the "yes" branch of step 270), then the quality of the speech reproduced in the selected instance of the output audio data is unacceptable, and the sound system performs step 290. In step 290, the sound system performs a corrective action to improve the quality of the speech reproduced by the sound system. For example, the corrective action can involve changing configurable parameters of the sound system and / or sending a message to the user, as described below with reference to Figure 3 described.

[0031] As will be understood by those skilled in the art, in Figure 2Many variations of the example embodiments shown are possible. For example, speech recognition may be performed on the output audio data associated with each output microphone at the corresponding local processing device or the user device associated therewith. Thus, at step 240, the sound system may instead receive, via network 140, the output audio speech recognition results for the output audio data associated with each output microphone, and step 230 may be omitted. Thereby, the processing burden of speech recognition is distributed across multiple processing devices. Further, prior to step 210, the sound system may identify available microphones at locations within the listening environment and select a set of microphones to be used as output audio microphones. For example, the microphones of the user device may be identified based on the connection established to network 140 from the listening environment or via global positioning system coordinates (or equivalents) within the listening environment. In such a case, a message may be sent to the user seeking to use the identified microphones of the user device for listening to the reproduced speech, and the user may choose to allow or deny permission. If permission is allowed, the sound system may then turn on the microphones and any other features necessary to cause the user device to send the output audio data at step 230. In another example, stand-alone microphones connected to the network in the listening environment may be identified and used to send the output audio data (e.g., if the configuration of the microphones permits). Further, the sound system performs a corrective action in response to determining that the quality of the reproduced speech in a single selection instance of the output audio data is unacceptable. In some other implementations, the corrective action may be performed based on other criteria. For example, the sound system performs a corrective action in response to determining that the quality of the reproduced speech in multiple instances of the output audio data is unacceptable. In another instance, the corrective action may be performed based on the location of the output microphone(s) within the listening environment associated with the unacceptable reproduced speech.

[0032] Figure 3 is a flowchart of a method 300 for adjusting a sound system to improve speech quality according to an embodiment of the present invention. For example, method 300 may be performed as Figure 2 the corrective action of step 290 shown. Method 300 may be performed by Figure 1 the host processing system 110 of the sound system 100 shown, or by another processing device of the sound system.

[0033] Method 300 begins at step 305. For example, method 300 may begin in response to determining that the difference between the output audio data and the speech recognition result(s) of the input audio data is greater than or equal to a threshold. A difference greater than or equal to the threshold indicates that the quality of the reproduced speech is unacceptable.

[0034] In step 310, the sound system tests the voice quality of the input voice from the source. For example, in step 310, the sound system can compare the result of voice recognition with a threshold of voice quality. The threshold can be a predefined confidence level (e.g., 60%). The threshold can be configured by the user. Being below the threshold indicates that the quality of the input voice is unacceptable. In some other examples, in step 310, the sound system can use one or more techniques that identify problems that adversely affect the quality of the input voice, such as a high volume of background noise relative to the input voice (indicated by a low signal-to-noise ratio), the proximity of the human speaker to the input microphone (indicated by a "pop" effect), the settings of the input microphone (e.g., audio sensitivity or gain / volume level), etc. Thus, in step 310, the sound system can perform a series of tests to identify potential problems associated with the audio input.

[0035] In step 320, the sound system determines whether the quality of the input voice from the source is acceptable. For example, in step 320, the sound system can determine whether the sound system has identified a problem with the input voice in step 310, which indicates that the voice quality is unacceptable. In response to determining that the quality of the input voice is acceptable (the "yes" branch of step 320), the sound system proceeds to step 340. However, in response to determining that the quality of the input voice is unacceptable (the "no" branch of step 204), the sound system enters step 330.

[0036] In step 330, the sound system sends an alert message to the user at the source. Specifically, the alert message sent to the user can include instructions for making adjustments based on the result of the (multiple) tests in step 310. For example, if the result of voice recognition performed on the input audio data is below the threshold, the alert message can instruct the human speaker to speak more clearly. In another example, if the test identifies a problem with the proximity of the human speaker to the input microphone, then the message can contain instructions to move closer to or further away from the input microphone. In yet another example, if the test identifies a problem with the input microphone, then the alert message can contain instructions to adjust the microphone settings (e.g., audio sensitivity or gain / volume level). In some other implementations, in the case where the sound system in step 310 identifies a problem with the input microphone, the automatic adjustment of the input microphone can be performed, for example, using steps 340 to 370 described below.

[0037] In step 340, the sound system performs a first parameter adjustment. The parameters adjusted can include any independently configurable parameter or setting of individual components of the sound system (e.g., a sound card, an audio speaker, or a microphone). Suitable parameters can include gain and equalization settings for audio components, etc. As will be understood by those skilled in the art, equalization settings include adjustable settings for multiple frequency ranges (also referred to as bands or channels) of an audio signal. Thus, in terms of equalization, each adjustable band of a component corresponds to an adjustable parameter. Accordingly, the adjustable parameters of the sound system include the configurable settings of each configurable component of the sound system, e.g., gain and equalization settings. The parameter adjustment can include a positive or negative increment to the value of a parameter of a particular audio component (e.g., an audio speaker). The parameter adjustment can be defined as an increment to the existing value of the parameter or the new (target) value of the parameter of the component. In step 340, the sound system can arbitrarily select the first parameter adjustment. Alternatively, the sound system can use an intelligent selection scheme to select the first parameter adjustment, which can be predefined or learned, as described below. Thus, in step 340, the sound system can include sending a configuration instruction to a remote component of the sound system (e.g., the sound card of an audio speaker) to adjust the identified parameter of the component of the sound system by a defined amount or increment. In some embodiments, the sound system in step 340 can include sending a configuration instruction to adjust the identified parameter to a target value.

[0038] In step 350, the sound system determines the impact of the first parameter adjustment in step 340 and stores information about the determined relationship. Specifically, in step 350, the sound system can perform an iteration of steps 210 to 260 of method 200 as shown in Figure 2 after the first parameter adjustment and determine the impact of the adjustment. For example, before and after the parameter adjustment, in step 350, the system can determine the impact of the adjustment by comparing the difference in the speech quality of the reproduced speech determined in step 260 during the iteration of steps 210 to 260 of method 200 as shown in Figure 2 with the speech quality of the input speech. At step 350, the sound system can determine a positive or negative impact (e.g., a percentage improvement or deterioration in speech quality) on the quality of the reproduced speech caused by the first parameter adjustment. The sound system can store the parameters and increments corresponding to the parameter adjustment and the determined impact on speech quality, which together provide information about the relationship between the first parameter and the quality of the reproduced speech for the sound system.

[0039] After the first parameter adjustment in step 340, in step 360, the sound system determines whether the quality of the reproduced speech is acceptable. For example, the sound system can correspond to Figure 2Step 270 of method 200 shown in []. In some embodiments, the value of the threshold used in step 360 to determine whether the quality of the reproduced speech is acceptable may be adjusted based on the quality of the input speech. For example, the confidence level (or equivalent) of the speech recognition result of the reproduced speech must be less than that of the input speech. Thus, the threshold confidence level (or equivalent) may be adjusted or determined based on the quality of the reproduced speech. For example, the threshold may vary according to the value of the confidence level (or equivalent) of the input speech, such as a fixed or variable percentage (e.g., 90% to 95%). In response to determining that the quality of the reproduced speech is acceptable (the "yes" branch of step 360), the method ends at step 375. However, in response to determining that the quality of the reproduced speech is still unacceptable (the "no" branch of step 360), the method continues to step 370.

[0040] At step 370, the sound system determines whether there are more configurable parameter adjustments to be made. Specifically, in some embodiments, the sound system may cycle through a predefined set of parameter adjustments only once, as Figure 2 part of the calibration action of step 290 of method 200 shown in []. In response to determining that there are more parameter adjustments to be made (the "yes" branch of step 370), the method returns to step 340, which makes the next parameter adjustment. The sound system continues from step 350 to step 370 in a loop until the sound system determines that there are no more parameter adjustments to be made. In response to determining that there are no more parameter adjustments to be made (the "no" branch of step 370), the method ends at step 375. In some other embodiments, the sound system may repeat the loop through a set of parameter adjustments of the sound system until a predefined condition is met. For example, the condition may be that the quality of the reproduced speech is acceptable, the quality of the reproduced speech has not improved significantly through successive parameter adjustments (i.e., the quality of the speech is maximized), or a timer expires. In this case, step 370 may be omitted, and the sound system determines at step 360 whether one or more of the conditions are met, and if not, method 300 returns to step 340, which makes the next parameter adjustment. Then, method 300 continues cyclically from step 350 to step 360 until the sound system determines that the quality of the audio output is acceptable (or another condition is met), and the method ends at step 375.

[0041] Thus, with method 300, when the input speech has acceptable quality, the sound system improves the quality of the reproduced speech by automatically adjusting the configuration of the sound system. Specifically, the sound system automatically adjusts the configurable parameters of the sound system to improve the quality of the reproduced speech.

[0042] In addition, the sound system determines and stores information regarding the relationship between one or more configurable parameters of the components of the sound system and the voice quality. Over time, this information can be used for more intelligent adjustment of the configurable parameters of the sound system. For example, in response to detecting unacceptable voice quality at one or more specific locations in the listening environment, the information can be used to predict the specific adjustment(s) required for a specific parameter or group of parameters to provide a minimum expected difference between the input voice and the reproduced voice quality.

[0043] As will be understood by those skilled in the art, there may be interdependencies among the configurable parameters of the sound system related to their impact or effect on voice quality. For example, a first parameter adjustment that includes a positive increment of a first parameter results in improved but unacceptable voice quality, a second parameter adjustment that includes a positive increment of a second parameter results in decreased voice quality, but a subsequent third parameter adjustment that includes a negative increment of the first parameter (to below its original level) results in acceptable voice quality. In this example, the first and second parameters are interdependent - a negative adjustment of the first parameter should be combined with a positive adjustment of the second parameter in order to improve voice quality. Such patterns of interdependency between the configurable parameters of the sound system and voice quality can be determined from the stored information, collected over a period of time, and used to develop an intelligent scheme for parameter adjustment of the sound system from step 340 to step 370.

[0044] In some implementations, machine learning can be used to develop an intelligent scheme for adjusting the sound system. Specifically, in response to one or a series of incremental parameter adjustments, the information recorded at step 350 can be stored in a centralized database for one or more sound systems and used as training data for a machine learning model. This training data can additionally include information about the input audio data (e.g., input microphone type / quality, gain / amplitude / volume, background noise, etc.) and / or information about the input voice (e.g., pitch, language, accent, etc.) and information about the type and arrangement of the associated sound system. In this way, a machine learning model can be developed to accurately predict the optimal configuration of a specific sound system for a specific type of input voice (e.g., a specific type of human speaker). The model can then be used to intelligently and / or simultaneously adjust multiple configurable parameters of the sound system (e.g., related to the same and / or different audio components) for optimizing the output voice quality. Concurrent parameter adjustments to achieve the predicted optimal configuration can reduce or eliminate the need for multiple incremental parameter adjustments and iterations from step 340 to step 370. After the development of the model, the information recorded at step 350 can be used as feedback to improve model performance.

[0045] Figure 4It is a block diagram of a system 400 according to an embodiment of the present invention. Specifically, the system 400 includes a processing component for a sound system as described herein.

[0046] The system 400 includes a host processing system 410, a database 470, and a processing device 450 (e.g., a local processing device and a user device) at a listening location that communicate with a host processing system 410 via a network 440. The network 440 may include any suitable wired or wireless data communication network, such as a mobile communication network, a local area network (LAN), a wide area network (WAN), or the Internet. The host processing system 410 includes a processing unit 412, a memory unit 414, and an input / output (I / O) unit 416. The host processing system 410 may include a user interface device 460 connected to the I / O unit 416. The user interface device 460 may include one or more of a display (e.g., a screen or a touch screen), a printer, a keyboard, a pointing device (e.g., a mouse, a joystick, a touchpad), an audio device (e.g., a microphone and / or a speaker), and any other type of user interface device.

[0047] The memory unit 414 includes an audio data file 420 and one or more processing modules 430 for performing the methods according to the present disclosure. The audio data file 420 includes input audio data 420A associated with an input microphone of the sound system. In addition, the audio data file 420 includes output audio data 420B associated with output microphones at distributed locations within the listening environment received via the I / O unit 416 through the network 440. Each processing module 430 includes instructions for being executed by the processing unit 412 to process data and / or instructions received from the I / O unit 416 and / or stored in the memory unit 414, such as the audio data file 420.

[0048] According to an exemplary implementation of the present disclosure, the processing module 430 includes a voice evaluation module 432, a configuration module 434, and a feedback module 436.

[0049] The voice evaluation module 432 is configured to evaluate the quality of the voice corresponding to the reproduction of the audio data played by the audio speaker of the sound system in the output audio data 420A. Specifically, the voice evaluation module 432 includes a voice recognition module 432A and a detection module 432B. The voice recognition module 432A is configured to perform voice recognition on the input audio data 420A and the output audio data 420B from the audio data file 420, e.g., at steps 220 and 240 of method 200 as shown in Figure 2 The detection module 432B is configured to detect when the quality of the reproduced voice in the output audio data 420B is unacceptable to the listener, e.g., from Figure 2Steps 250 to 270 of the method 200 shown. Accordingly, the speech evaluation module 432 retrieves and processes the audio data file 420 to perform Figure 2 the method 200 shown therein. It should be noted that, as described herein, when the speech evaluation module 432 receives the input audio data 420A and the output audio data 420B from the sound system, the input audio data 420A and the output audio data 420B can be used to perform the processing of the speech evaluation module 432 in real time.

[0050] The configuration module 434 is configured to adjust the configurable parameters of the sound system to optimize the quality of the reproduced speech. The configuration module 434 includes a calibration module 434A, a parameter adjustment module 434B, and an adjustment evaluation module 434C. The calibration module 434A is configured to calibrate the sound system, for example, at a set time and subsequently when needed. Specifically, the calibration module 434A can use the pre-recorded input audio data file 420A including speech that is considered "perfect" for the purpose of speech recognition in combination with the speech evaluation module 432 to calibrate the sound system. If the detection module 432B detects that the quality of the reproduced speech is unacceptable to the listener, the parameter adjustment module 434B and the adjustment evaluation module 434C are used to perform parameter adjustment and evaluation, as described below, until the quality of the reproduced speech is maximized. As will be understood by those skilled in the art, the calibration of the sound system using the "perfect" speech sample determines the difference in the determined quality of the reproduced speech compared to the input speech that can be achieved using the sound system in the best case. For example, this can be used to set the initial threshold for the acceptable quality of the reproduced speech used in Figure 2 step 270 of the method 200 shown. As described above, the threshold can be adjusted based on the quality of the actual input speech during use. The parameter adjustment module 434B is configured to adjust the configurable parameters of the sound system. For example, the parameter adjustment module 434B can use any or intelligent scheme to iteratively perform parameter adjustment, such as, as shown in Figure 3 step 340 of the method 300 shown. The adjustment evaluation module 434C is configured to evaluate the impact of the parameter adjustment performed by the parameter adjustment module 434B. Specifically, the adjustment evaluation module 434C is configured to determine whether the quality of the reproduced speech is acceptable after the parameter adjustment, as from Figure 3 steps 350 and 360 of the method 300 shown. As described above, the parameter adjustment module 434B and the adjustment evaluation module 434C can be called by the calibration module 434A and the detection module 432B to respectively configure and reconfigure the parameters of the sound system to optimize the quality of the reproduced speech.

[0051] The feedback module 436 is configured to provide information obtained from the operation of the speech evaluation module 432 and / or the configuration module 434 as feedback. For example, the feedback module 436 may provide feedback (e.g., an alert message) indicating that the quality of the input speech determined by the speech recognition module 432A or other analysis of the input audio data is unacceptable to the human speaker, e.g., as at steps 310 and 320 of method 300 as shown in Figure 3 . Additionally or alternatively, the feedback module 436 may provide information regarding the impact of parameter adjustment on the quality of the reproduced speech to a system or model for developing or improving an intelligent parameter adjustment scheme to optimize performance, for use at step 340 of method 300 as shown in Figure 3 . For example, the feedback module 436 may send feedback including information regarding the relationship between the sound system parameters stored at step 350 of method 300 as shown in Figure 3 and the speech quality to the centralized database 470 or to another data storage device via the network 440. The stored data may be used as training data for a machine learning model for optimizing sound system performance or as feedback for refining an existing machine learning model. Further, in the case where the configuration module 434 cannot optimize the parameters of the sound system to provide reproduced speech with acceptable speech quality, the feedback module 436 may provide feedback to the user of the host processing system 410. For example, when the sound system determines at step 280 of method 200 as shown in Figure 2 that no further output audio data instances will be considered before method 200 ends at step 295, an alert message may be sent. The alert message may provide recommendations to the owner of the sound system, such as recommending actions to improve the performance of the sound system. For example, the alert message may instruct the owner to perform a manual inspection of the sound system components (e.g., sound card) to change the number or location of the components and / or change the total power of the sound system. Techniques for making recommendations for manual inspection and changes to the sound system are known in the art and any suitable technique, whether currently known or developed in the future, may be used.

[0052] Referring to Figure 4 , a computer program product 480 is provided. The computer program product 480 includes a computer-readable medium 482 having a storage medium 484 and program instructions 486 (i.e., program code) embodied therewith. The program instructions 486 are configured to be loaded onto the memory unit 414 of the host processing system 410 via the I / O unit 416 (e.g., by a user interface device 460 or one of the other devices 450 connected to the network 440). In an exemplary embodiment, the program instructions 486 are configured to perform one or more of the steps of the methods disclosed herein, such as those described above in Figure 2 or Figure 3Steps of the method shown in

[0053] Although the present disclosure has been described and illustrated with reference to example implementations, those skilled in the art will recognize that the present disclosure itself is applicable to many different variations and modifications that are not specifically shown herein.

[0054] The present invention can be a system, method, and / or computer program product. The computer program product can include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to execute aspects of the present invention.

[0055] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium can be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device (such as a punched card or protruding structures in a groove having instructions recorded thereon), and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable) or an electrical signal transmitted through a wire.

[0056] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device via a network (e.g., the Internet, a local area network (LAN), a wide area network (WAN), and / or a wireless network), or to an external computer or external storage device. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.

[0057] The computer-readable program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages (such as Smalltalk, C++) and conventional procedural programming languages (such as the C programming language) or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions to perform aspects of the present invention, thereby executing the computer-readable program instructions.

[0058] The present invention will now be described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0059] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing apparatus create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, which instructions cause a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, so that the computer-readable storage medium storing the instructions includes a manufacture, including instructions for implementing aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0060] The computer-readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device, such that a series of operation steps are performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0061] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.

[0062] The description of the various exemplary embodiments of the present disclosure has been presented for purposes of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The terms used herein have been chosen to best explain the principles of the exemplary implementations, the practical application, or technical improvements over technologies found in the marketplace, or to enable those of ordinary skill in the art to understand the implementations disclosed herein.

Claims

1. A computer-implemented method for processing audio data, comprising: performing speech recognition on input audio data, the input audio data including a speech input to a sound system; performing speech recognition on at least one instance of output audio data, the output audio data including speech reproduced by one or more audio speakers of the sound system; determining a difference between a result of the speech recognition of the input audio data and a result of the speech recognition of the at least one instance of the output audio data; and when the difference is greater than or equal to a threshold, determining that the quality of the reproduced speech is unacceptable, the method further comprising: in response to determining that the quality of the reproduced speech is unacceptable, performing one or more parameter adjustments of the sound system to improve the quality of the reproduced speech, the method further comprising: based on a difference between a result of the speech recognition of the input audio data and a result of the speech recognition of the at least one instance of the output audio data, determining an impact of the parameter adjustment on the quality of the reproduced speech, the input audio data and the output audio data including the input speech and the reproduced speech after the parameter adjustment, and storing relationship information, the relationship information including parameters and increments corresponding to the parameter adjustment and the impact on the quality of the reproduced speech for use as feedback for an intelligent selection scheme or a machine learning model for optimizing the quality of the reproduced speech.

2. The computer-implemented method according to claim 1, wherein the difference includes a quantitative value calculated from a result of the speech recognition of the input audio data for a sample of the input speech and a result of the speech recognition of the output audio data for a sample of the reproduced speech corresponding to the sample of the input speech.

3. The computer-implemented method according to claim 1, wherein determining the difference between a result of the speech recognition of the input audio data and a result of the speech recognition of the at least one instance of the output audio data comprises: comparing text of a transcript of the speech recognition of the input audio data with text of a transcript of the speech recognition of the at least one instance of the output audio data; and determining a quantitative value of the difference for a group selected from the group consisting of the number of different letters, the number of different words, and the percentage of different letters or words.

4. The computer-implemented method according to claim 1, wherein determining the difference between a result of the speech recognition of the input audio data and a result of the speech recognition of the at least one instance of the output audio data comprises: Compare a first confidence level determined by speech recognition of the input audio data with a second confidence level determined by speech recognition of the at least one instance of the output audio data, the first confidence level including a value of a confidence metric indicating the reliability of the speech recognition of the input audio data, and the second confidence level including a value of a confidence metric indicating the reliability of the speech recognition of the at least one instance of the output audio data; and Determine the difference between the first confidence level and the second confidence level.

5. The computer-implemented method according to claim 1, wherein performing the one or more parameter adjustments of the sound system comprises:[[]] Perform a first parameter adjustment, the first parameter adjustment including adjusting an increment or adjustment defined by a parameter of the sound system to a target value; and In response to determining that the quality of the reproduced speech is still unacceptable, perform further parameter adjustments until a predefined condition is met.

6. The computer-implemented method according to claim 5, wherein the predefined condition is selected from the group consisting of: the quality of the reproduced speech is acceptable, the quality of the reproduced speech is maximized, a predefined set of parameter adjustments has been performed, the predefined set of parameter adjustments has been performed according to an intelligent selection scheme, the predefined set of parameter adjustments has been adjusted according to the machine learning model for optimizing the quality of the reproduced speech, and a timer expires.

7. The computer-implemented method according to any one of claims 1 to 6, further comprises:[[]] Determine whether the quality of the speech input to the sound system is acceptable, and In response to determining that the quality of the speech input to the sound system is not acceptable, send a message to the user to make changes related to the speech input to the sound system.

8. The computer-implemented method according to any one of claims 1 to 6, wherein the parameters in the one or more parameter adjustments are selected from the group consisting of: audio gain and audio channel equalization for each frequency band of components of the sound system.

9. A computer device,[[]] comprises:[[]] A processor and a data storage device, wherein the processor is configured to:[[]] Perform speech recognition on input audio data, the input audio data including a speech input to a sound system; Perform speech recognition on at least one instance of output audio data, the output audio data including speech reproduced by one or more audio speakers of the sound system; Determine the difference between the result of the speech recognition of the input audio data and the result of the speech recognition of the at least one instance of the output audio data; and When the difference is greater than or equal to a threshold, determine that the quality of the reproduced speech is unacceptable, The processor is further configured to:[[]] In response to determining that the quality of the reproduced speech is unacceptable, perform one or more parameter adjustments of the sound system to improve the quality of the reproduced speech, The processor is further configured to:[[]] Determine the impact of parameter adjustment on the quality of the reproduced speech based on the difference between the result of the speech recognition of the input audio data and the result of the speech recognition of the at least one instance of the output audio data, where the input audio data and the output audio data include the input speech and the reproduced speech after the parameter adjustment, and Store relationship information, where the relationship information includes parameters and increments corresponding to the parameter adjustment and the impact on the quality of the reproduced speech, for use as feedback for an intelligent selection scheme or a machine learning model for optimizing the quality of the reproduced speech.

10. The apparatus according to claim 9, wherein the difference includes a quantitative value calculated from the result of the speech recognition of the input audio data for a sample of the input speech and the result of the speech recognition of the output audio data for a sample of the reproduced speech corresponding to the sample of the input speech.

11. The apparatus according to claim 9, wherein the processor is configured to determine the difference between the result of the speech recognition of the input audio data and the result of the speech recognition of the at least one instance of the output audio data by: Comparing the text of the transcript of the speech recognition of the input audio data with the text of the transcript of the speech recognition of the at least one instance of the output audio data; and Determining a quantitative value for the difference selected from the group including the number of different letters, the number of different words, and the percentage of different letters or words.

12. The apparatus according to claim 9, wherein the processor is configured to determine the difference between the result of the speech recognition of the input audio data and the result of the speech recognition of the at least one instance of the output audio data by: Comparing a first confidence level determined by the speech recognition of the input audio data with a second confidence level determined by the speech recognition of the at least one instance of the output audio data, where the first confidence level includes a value of a confidence metric indicating the reliability of the speech recognition of the input audio data, and the second confidence level includes a value of a confidence metric indicating the reliability of the speech recognition of the at least one instance of the output audio data; and Determining the difference between the first confidence level and the second confidence level.

13. The apparatus according to claim 9, wherein, the processor is configured to perform the one or more parameter adjustments of the sound system by: Performing a first parameter adjustment, where the first parameter adjustment includes adjusting an increment defined by the parameters of the sound system or adjusting to a target value; and In response to determining that the quality of the reproduced speech is still unqualified, performing further parameter adjustments until a predefined condition is met.

14. The apparatus according to claim 13, wherein the predefined condition is selected from the group consisting of: the quality of the reproduced speech is acceptable, the quality of the reproduced speech is maximized, a predefined set of parameter adjustments has been performed, the predefined set of parameter adjustments has been performed according to an intelligent selection scheme, the predefined set of parameter adjustments has been adjusted according to the machine learning model for optimizing the quality of the reproduced speech, and a timer expires.

15. The apparatus according to any one of claims 9 to 14, wherein the processor is further configured to: determine whether the quality of the voice input to the sound system is acceptable, and in response to determining that the quality of the voice input to the sound system is not acceptable, send a message to the user to make a change related to the voice input to the sound system.

16. The apparatus according to any one of claims 9 to 14, wherein the parameters in the one or more parameter adjustments are selected from the group consisting of: audio gain and audio channel equalization for each frequency band of the components of the sound system.

17. A computer-readable storage medium comprising program instructions, wherein the program instructions are executable by a processor to cause the processor to: perform speech recognition on input audio data, the input audio data including a voice input to a sound system; perform speech recognition on at least one instance of output audio data, the output audio data including speech reproduced by one or more audio speakers of the sound system; determine a difference between a result of the speech recognition of the input audio data and a result of the speech recognition of the at least one instance of the output audio data; and when the difference is greater than or equal to a threshold, determine that the quality of the reproduced speech is unacceptable, the program instructions are executable by the processor to further cause the processor to: in response to determining that the quality of the reproduced speech is unacceptable, perform one or more parameter adjustments of the sound system to improve the quality of the reproduced speech, the program instructions are executable by the processor to further cause the processor to: determine an impact of the parameter adjustment on the quality of the reproduced speech based on a difference between a result of the speech recognition of the input audio data and a result of the speech recognition of the at least one instance of the output audio data, the input audio data and the output audio data including input speech and the reproduced speech after the parameter adjustment, and store relationship information, the relationship information including parameters and increments corresponding to the parameter adjustment and the impact on the quality of the reproduced speech for use as feedback to an intelligent selection scheme or a machine learning model for optimizing the quality of the reproduced speech.

18. A computer program product comprising program code, which when run on a computer, is adapted to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method and apparatus for adjusting audio signal, terminal and computer readable storage medium

    CN107484081A

  • Method and system for speech quality perception evaluation based on speech semantic recognition technology

    CN108877839A

  • Speech clarity systems and techniques

    US20170287355A1

  • Speech recognition system having multiple speech recognizers

    US7228275B1