Improving the audio quality of speech in sound systems
A computer-implemented method using speech recognition to detect and correct unsatisfactory speech quality in sound systems by automatically adjusting settings, addressing distortion and clarity issues in real-time.
Patent Information
- Application Number
- JP2022520788
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-10-10
- Filing Date
- 2020-10-05
- Publication Date
- 2025-08-06
- Estimated Expiration
- 2040-10-05
AI Technical Summary
Existing sound systems often reproduce speech that is distorted or difficult to understand due to suboptimal settings or imperfect input speech, leading to interruptions and manual adjustments that may not fully resolve the issues.
Implementing a computer-implemented method that uses speech recognition on input and output audio data to detect unsatisfactory speech quality and automatically adjusts sound system parameters to improve clarity.
The method allows for real-time, automatic correction of speech quality issues, enhancing listener experience by minimizing interruptions and ensuring clearer speech reproduction.
Smart Images

Figure 0007719573000001 
Figure 0007719573000002 
Figure 0007719573000003
Abstract
Description
[Technical Field]
[0001] The present invention relates generally to techniques for improving the quality of audio output from a sound system, and more particularly to improving the quality of speech to a listener to the audio output of a sound system.
[0002] Sound systems are frequently used to play speech through audio speakers to audiences, such as participants in conferences, lectures, and demonstrations in theaters or auditoriums, conference calls, and webinars at distributed geographic locations over a communications network. In such systems, input speech is received into a microphone and optionally recorded at a host system, the audio data is communicated by the host system to one or more audio speakers, and the audio speaker(s) output (i.e., "play") the reproduced speech to the listeners. In many cases, the speech reproduced through the audio speakers is not a perfect reproduction of the input speech (e.g., the speech may be incomprehensible). For example, if the audio speaker settings are not optimized, the reproduced sound and the resulting reproduced speech may be distorted, making it difficult for the listener to hear or understand, or both. In other cases, the input speech itself may be imperfect, for example, due to the microphone's position relative to the speech source or its suboptimal settings. This, in turn, makes it difficult to hear or understand the speech played over the audio speakers. Typically, such problems related to speech in the audio output can be resolved by adjustments to the sound system. For example, if a listener notifies the host that they are having difficulty hearing or understanding the speech, the host can adjust configurable settings on the sound system or ask a human speaker to move relative to the microphone. However, this creates interruptions and delays while the adjustments are made. Additionally, because the adjustments are manual, they may not completely resolve the listener's difficulties. Summary of the Invention
[0003] According to an aspect of the present invention, a computer-implemented method is provided. The computer-implemented method includes performing speech recognition on input audio data including speech input to an audio system. The computer-implemented method further includes performing speech recognition on at least one instance of output audio data including speech reproduced by one or more audio speakers of the audio system. The computer-implemented method further includes determining a difference between a speech recognition result for the input audio data and a speech recognition result for the at least one instance of the output audio data. The computer-implemented method further includes determining that the quality of the reproduced speech is unsatisfactory if the difference is greater than or equal to a threshold value.
[0004] According to another aspect of the present invention, an apparatus is provided. The apparatus includes a processor and storage. The processor is configured to perform speech recognition on input audio data including speech input to an audio system. The processor is further configured to perform speech recognition on at least one instance of output audio data including speech reproduced by one or more audio speakers of the audio system. The processor is further configured to determine a difference between a result of the speech recognition on the input audio data and a result of the speech recognition on the at least one instance of output audio data. The processor is further configured to determine that the quality of the reproduced speech is unsatisfactory if the difference is greater than or equal to a threshold value.
[0005] According to yet another aspect of the present invention, there is provided a computer program product comprising a computer-readable storage medium having program instructions embodied thereon, which, when executed by a processor, causes the processor to: perform speech recognition on input audio data including speech input to an audio system; perform speech recognition on at least one instance of output audio data including speech reproduced by one or more audio speakers of the audio system; determine a difference between a result of the speech recognition on the input audio data and a result of the speech recognition on the at least one instance of the output audio data; and determine that the quality of the reproduced speech is unsatisfactory if the difference is greater than or equal to a threshold value. [Brief explanation of the drawings]
[0006] [Figure 1] FIG. 1 is a schematic diagram illustrating an audio system according to one embodiment of the present invention. [Figure 2] FIG. 2 is a flowchart of a method for detecting and correcting unsatisfactory speech quality according to one embodiment of the present invention. [Figure 3] FIG. 3 is a flowchart of a method for adjusting an acoustic system to improve speech quality according to one embodiment of the present invention. [Figure 4] FIG. 4 is a block diagram of an audio system according to one embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0007] Implementations of embodiments of the present disclosure are described below with reference to the drawings.
[0008] The present disclosure provides systems and methods for detecting when the quality of speech reproduced by an audio system is unsatisfactory (e.g., difficult for a listener to hear or unclear / incoherent) and for reconfiguring the audio system to improve the quality of the reproduced speech. The techniques of the present disclosure are performed automatically, limiting interruptions and improving the listener's experience.
[0009] In particular, according to the present disclosure, one or more microphones are distributed at locations within a listening environment and are used to detect reproduced speech reproduced by one or more audio speakers of a sound system. Output audio data including the reproduced speech received by each of the one or more microphones is recorded. Speech recognition is performed on the output of the audio data associated with each of the one or more microphones to determine the quality of the speech reproduced at the corresponding microphone location. Additionally, speech recognition is performed on input audio data including the speech input to the sound system to determine the quality of the speech from the source. A comparison is performed between the results of the speech recognition performed on the input audio data and the results of the speech recognition performed on the corresponding output audio data for each microphone. The results of the comparison are used to determine whether the speech quality is unsatisfactory to the listener and, if so, to take corrective action, such as adjusting the sound system to improve the quality of the reproduced speech.
[0010] In this disclosure, the term "speech" is used to refer to sound or audio, including speech. The term "input audio data" refers to digital audio data about sound or audio, including speech, originating from a source (e.g., a human speaker) and detected by a microphone (herein "input microphone") of a sound system. The term "output audio data" refers to digital audio data about sound or audio reproduced by one or more audio speakers of a sound system and detected by a microphone (herein "output microphone"). Audio data therefore "represents" or "includes" sound or audio, including speech, received by a microphone of a sound system. Reference to "recording" audio data refers to storing audio data in data storage, which includes long-term data storage of audio data as an audio data file as well as transient storage of audio data for communication.
[0011] FIG. 1 is a schematic diagram illustrating an audio system according to one embodiment of the present invention. Audio system 100 includes a host processing system 110, multiple microphones 120, and multiple audio speakers 130 interconnected by a data communications network 140. In the system shown in FIG. 1, audio system 100 is a distributed system including microphones 120 and audio speakers 130 at different spatial locations (e.g., a meeting or conference room). At least one location (Location 1) includes host processing system 110, which, in the illustrated embodiment, is the location of the speech source input to audio system 100. Other locations (Locations 2 and 3) include audio speakers 130 for reproducing audio in corresponding listening environments.
[0012] The host processing system 110 typically includes a user computing system (e.g., a notebook computer), a dedicated audio system controller, or the like, which can be operated by a user to manage the audio system 100. The multiple microphones 120 include an input microphone 122 for detecting and recording speech from a source (e.g., a human speaker) for playback by the audio system 100. The input microphone 122 can be a dedicated microphone (e.g., a microphone on a lectern) for receiving audio input to the audio system 100, or a microphone on the user computing system that can be "switched on" under the control of the host processing system 110. Multiple audio speakers 130 play the input speech and are distributed at different positions within one or more locations to form a listening environment (Locations 2 and 3). In particular, each audio speaker 130 receives and plays audio data corresponding to the recorded speech from the host processing system 110 over a network 140. The audio speakers 130 may include, at a location, one or more dedicated loudspeakers 134 of a sound system (e.g., speakers at fixed locations in a movie theater), or audio speakers 132 of a user computer, a networked telephone, etc. The communications network 140 may include any suitable wired or wireless network for communication between the host processing system 110, the microphone 120, and the audio speakers 130.
[0013] In one embodiment, the plurality of microphones 120 further includes output microphones 124 positioned at multiple locations within the listening environment (locations 2 and 3) to receive and record speech played from audio speakers 130 for analysis as described herein. The output microphones 124 may include, for example, dedicated microphones of a sound system associated with a location's audio speaker 134. The output microphones 124 may also include microphones or other devices of user computing systems present within the listening environment that are identified by and made available for this purpose to the host processing system 110. In the distributed system of FIG. 1 , the output microphones 124 in one listening environment (location 3) are each associated with a local processing system 150, such as a system device for recording the played speech as output audio data for communication over network 140 to the host processing system 110. An output microphone 124 in another listening environment (Location 2) is configured to communicate audio data over network 140 to host processing system 110, which records the output audio data. As one skilled in the art will appreciate, microphones 120 can include digital microphones that generate digital output signals, or analog microphones that generate analog output signals that are recorded and converted to digital audio data by another component along the audio signal chain, or both.
[0014] The host processing system 110 records speech received by the input microphone 122 from a source as input audio data. Additionally, the host processing system 110 receives output audio data corresponding to the speech recorded by each output microphone 124 and communicated over the network 140 and played back by the audio speaker 130. According to one embodiment of the present invention, the host processing system 110 is configured to perform speech recognition on the input audio data associated with the input microphone 122 and the output audio data associated with each output microphone 124. Speech recognition techniques are known in the art, and the host processing system 110 can implement any suitable speech recognition technique. Speech recognition may provide a transcript of the speech, or a confidence metric value or level, or a combination thereof, indicating the reliability of the speech recognition. Speech recognition may be performed continuously or periodically on the input audio data and the corresponding output audio data. The host processing system 110 further compares the speech recognition results determined for the output audio data associated with each output microphone 124 with the speech recognition results determined corresponding to the input audio data associated with the input microphones 122. If the comparison determines an unacceptable difference between the results determined for the output audio data of one or more output microphones 124 and the results determined for the input audio data of the input microphones 122, the host processing system 110 determines that the speech quality of the recorded speech is unsatisfactory to the listener and takes corrective action. For example, corrective action may include adjusting parameters of components of the sound system (e.g., gain or channel equalization settings of a sound card controlling audio speakers) or sending a message to the user to adopt a certain structure (e.g., instructions to direct a human speaker to move closer or further away from an input microphone).For example, an unacceptable difference can be determined when the compared results obtained from speech recognition, such as a difference in confidence level or a measured difference in speech transcripts, are lower than or equal to a threshold value, as further described below.
[0015] Thus, host processing system 110 can detect when the quality of speech reproduced by sound system 100 is unsatisfactory to the listener (i.e., unclear, distorted, or too quiet) and can take action to improve the quality of the speech. The disclosed techniques can be performed automatically and in real time, improving the listener's experience. In the embodiment illustrated in FIG. 1, the disclosed techniques are implemented in host processing system 110. As one skilled in the art will recognize, the present invention can be implemented in any other processing device or system in communication with sound system 100, such as local processing system 150 or a combination of processing devices.
[0016] 2 is a flowchart of a method 200 for detecting and correcting unsatisfactory speech quality according to one embodiment of the present invention. For example, method 200 may be performed by host processing system 110 of audio system 100 of FIG.
[0017] Method 200 begins at step 205. For example, step 205 can be initiated in response to the initiation of an acoustic check of the sound system, in response to the start of speech, or otherwise.
[0018] In step 210, the sound system receives input audio data for speech input to the sound system from a source. For example, the input audio data may be received from input microphone 122 of sound system 100 of Figure 1 in response to the speech of a human speaker speaking into input microphone 122. The input audio data is typically received and recorded (e.g., stored in a file of input audio data) in substantially real time (i.e., with minimal delay), for example, by host processing system 110 of Figure 1.
[0019] In step 220, the acoustic system performs speech recognition on the input audio data to determine a result of the recognition of the input audio-speech indicative of speech quality. Any suitable speech recognition technique or algorithm can be used to perform the speech recognition. The speech recognition result typically includes a transcript of the speech contained in the associated audio data. The quality of the transcript can be indicative of the quality of the speech. Additionally, the speech recognition result can include a confidence metric value or level indicative of the reliability of the speech recognition. Such confidence metrics are well known in the art of speech recognition. Thus, the confidence level can also be indicative of the quality of the speech. The speech recognition can provide other results indicative of the quality of the speech.
[0020] In step 230, the sound system receives output audio data corresponding to the reproduced speech reproduced by the sound speakers of the sound system at one or more locations within the listening environment. For example, an instance of the output audio data may be received from each of one or more output microphones 124 that detect the reproduced speech reproduced by the audio speakers 130 of the sound system 100 shown in FIG. 1. The output audio data is typically received substantially in real time but is necessarily delayed relative to the input audio data. For example, the delay is due to the communication of the input audio data to the audio speakers and the communication of the corresponding output audio data associated with the output microphones 124 over the network 140 shown in FIG. 1. In some implementations, the output audio data is recorded, for example, by the local processing system 150 or the host processing system 11 shown in FIG. 1.
[0021] In step 240, the acoustic system performs speech recognition on the output audio data to determine an output audio speech recognition result indicative of speech quality. In particular, in step 240, the acoustic system performs speech recognition on each instance of the received output audio data. In step 240, the acoustic system uses the speech recognition technique used in step 220 so that the results of steps 220 and 240 are comparable.
[0022] Thus, in steps 210 and 220, the acoustic system derives speech recognition result(s) for the input audio data, and in steps 230 and 240, the acoustic system performs derivation of speech recognition result(s) for each instance of the output audio. In each case, the speech recognition result may include a transcript of the speech, or a confidence level, or both. As one skilled in the art will recognize, in practice, steps 210 through 240 may be performed continuously, particularly in applications where input and output audio data are received and processed continuously, such as in real time.
[0023] In an implementation of the example of Figure 2, multiple instances of output audio data are received. In particular, each instance of output audio data is associated with a particular output microphone located within the listening environment. In step 250, the sound system selects the results of speech recognition for the first instance of output audio data.
[0024] In step 260, the acoustic system compares the selected speech recognition results for the instance of output audio data with the speech recognition results for the corresponding input audio data and determines a difference in speech quality. The difference is a quantitative value representing the difference in speech quality calculated using the speech recognition result(s). In one implementation, in step 260, the acoustic system can compare the text of the speech recognition transcript determined for the instance of output audio data with the text of the transcript for the corresponding input audio data and determine a difference, such as a raw count difference or percentage difference (e.g., word) in the transcript text. A difference in the text of the transcript speech indicates a degradation in speech quality at the audio speaker compared to the speech quality at the source, and the amount of difference indicates the amount of quality degradation. In another implementation, in step 260, the acoustic system can compare the confidence level of the speech recognition determined for the instance of output audio data with the confidence level determined for the corresponding input audio data and determine a difference. As described above, the confidence level indicates the reliability of the transcribed speech (e.g., the value of the confidence metric expressed as a percentage). Because the reliability of the transcribed speech depends on the quality of the speech reproduced to the listener, differences in the confidence levels indicate a decrease in the quality of the reproduced speech at the audio speaker compared to the quality of the speech at the source, and the amount of difference indicates the amount of quality degradation. In other implementations, the acoustic system can use other criteria derived from the speech recognition result(s) to detect speech quality degradation. As one skilled in the art will recognize, in step 260, the acoustic system can compare the speech recognition results for the output audio data sample with the speech recognition results for the corresponding input audio data sample.In some scenarios, steps 210 through 240 may be performed continuously (e.g., substantially in real time) on the input and playback speech, in which case corresponding samples of the input audio data and the output audio data contain identical input and playback speech and may be identified using any suitable technique, such as one or more of time synchronization and audio matching (identifying identical sections of audio in the audio data) or speech recognition transcript matching (identifying identical sections by matching words and phrases in the transcript text). In other scenarios, steps 210 through 240 may be performed by periodically sampling the input and playback speech (e.g., sampling the input and playback speech over synchronized, temporally separated time windows), such that the speech recognition result(s) relate to corresponding samples of the input and output audio data.
[0025] In step 270, the sound system determines whether the difference is greater than or equal to a threshold value. The threshold value is a difference metric (e.g., a number / percentage, or a word in text, or a reliability metric) that indicates an unacceptable degradation in speech quality of the reproduced speech compared to the quality of the input speech. The threshold value can be selected according to application requirements and can be changed by the user. For example, in some applications, a difference of up to 5% in speech quality is acceptable, so the threshold is set to 5%, while in other applications, a difference of up to 10% in speech quality is acceptable, so the threshold is set to 10%. In some implementations, the threshold value can be adjusted based on the quality of the input speech, as described below.
[0026] If the difference is less than the threshold (NO branch of step 270), the quality of the reproduced speech in the selected instance of output audio data is satisfactory, and the sound system performs step 280. In step 280, the sound system determines whether there are more instances of output audio data to consider. If there are more instances of output audio data to consider (YES branch of step 280), the sound system begins performing step 250, and then continues looping through steps 260 and 270 until the sound system determines in step 280 that there are no more instances of output audio data to consider. After the sound system determines that there are no more instances of output audio data to consider, the sound system stops processing in step 295.
[0027] If the difference is greater than or equal to the threshold value (YES branch of step 270), the reproduced speech of the selected instance of the output audio data is unsatisfactory, and the sound system performs step 290. In step 290, the sound system performs corrective action to improve the quality of the speech reproduced by the sound system. For example, the corrective action may include changing configuration parameters of the sound system, sending a message to the user, or both, as described below with reference to FIG.
[0028] As those skilled in the art will recognize, many variations on the implementation of the embodiment illustrated in FIG. 2 are possible. For example, speech recognition can be performed on the output audio data associated with each output microphone of a corresponding local processing device or associated user device. Thus, in step 240, the sound system can instead receive the output audio speech recognition results for the output audio data associated with each output microphone over network 140, thereby omitting step 230. Thus, the processing burden of speech recognition is distributed across multiple processing devices. Additionally, prior to step 210, the sound system can identify microphones available at locations within the listening environment and select a set of microphones to use as output audio microphones. For example, the microphones of the user device can be identified based on an established connection from the listening environment to network 140 or Global Positioning System coordinates (or equivalent) within the listening environment. In this case, a message can be sent to the user requesting permission to use the identified microphone of the user device to listen to the reproduced speech, and the user can choose to grant or deny permission. If permission is granted, the sound system then switches on the microphone and any other features necessary to enable the user device to transmit output audio data in step 230. In another embodiment, a standalone networked microphone can be identified within the listening environment and used to transmit output audio data (e.g., if configured to grant permission for the microphone(s)). Furthermore, the sound system can take corrective action in response to determining in step 270 that the quality of the reproduced speech in a single selected instance of the output audio data is unsatisfactory. In other implementations, corrective action can be taken based on other criteria.For example, the sound system may take corrective action in response to determining that the quality of the reproduced speech for a number of instances of the output audio data is unsatisfactory. In another embodiment, the corrective action may be taken based on the location of the output microphone(s) within the listening environment associated with the unsatisfactory reproduced speech.
[0029] Figure 3 is a flowchart of a method for adjusting an audio system to improve speech quality, according to one embodiment of the present invention. For example, method 300 may be performed as a corrective action in step 290 of Figure 2. Method 300 may be performed by the host processing system 110 of audio system 100 shown in Figure 1 or by another processing device of the audio system.
[0030] Method 300 begins at step 305. For example, method 300 may begin in response to determining that a difference between the speech recognition result(s) of the output audio data and the input audio data is greater than or equal to a threshold value, where the difference is indicative of unsatisfactory quality of the reproduced speech.
[0031] In step 310, the acoustic system tests the speech quality of the input speech from the source. For example, in step 310, the acoustic system can compare the results of speech recognition to a threshold for speech quality. The threshold can be a predetermined confidence level (e.g., 66%). The threshold can be user-configurable. A value below the threshold indicates unsatisfactory quality of the input speech. In other embodiments, in step 310, the acoustic system can process the input audio data using one or more techniques to identify problems that adversely affect the quality of the input speech, such as a high volume of background noise compared to the input speech (indicated by a low signal-to-noise ratio), the proximity of a human speaker to the input microphone (indicated by a "popping" effect), input microphone settings (e.g., audio sensitivity or gain / volume level), etc. Accordingly, in step 310, the acoustic system performs a series of tests to identify possible problems associated with the audio input.
[0032] In step 320, the audio system determines whether the quality of the input speech from the source is acceptable. For example, in step 320, the audio system may determine whether the audio system identified a problem with the input speech in 310 that indicated the speech quality was unsatisfactory. In response to determining that the quality of the input speech is acceptable (YES branch of step 320), the audio system continues to step 340. However, in response to determining that the quality of the input speech is unsatisfactory (NO branch of step 320), the audio system proceeds to step 330.
[0033] In step 330, the audio system sends a warning message to the user of the source. In particular, the warning message to the user can include instructions for adjustments based on the results of the test(s) in step 310. For example, if the results of speech recognition performed on the input audio data are below a threshold, the warning message can instruct the human speaker to speak more clearly. In another embodiment, if the test identifies a problem with the human speaker's proximity to the input microphone, the message can include instructions for moving closer or farther away from the input microphone. In yet another embodiment, if the test identifies a problem with the input microphone, the warning message can include instructions for adjusting the microphone settings (e.g., audio sensitivity or gain / volume level). In other implementations, in scenarios where the audio system identifies a problem with the input microphone in step 310, automatic adjustments to the input microphone are performed, e.g., using steps 340 through 370, as described below.
[0034] In step 340, the sound system performs a first parameter adjustment. The parameter adjustment can include any independently configurable parameter or setting of an individual component of the sound system, such as a sound card, audio speakers, or microphone. As one skilled in the art will appreciate, equalization settings include adjustable settings for multiple frequency regions (also called frequency bands or channels) of an audio signal. Thus, with respect to equalization, each adjustable frequency band of a component corresponds to an adjustable parameter. Thus, adjustable parameters of the sound system include configurable settings, such as gain and equalization settings, for each configurable component of the sound system. The parameter adjustment can include a positive or negative increment of the parameter value of a particular audio component, such as an audio speaker. The parameter adjustment can be defined for a component parameter by increasing the parameter's existing value or a new (target) value. In step 340, the sound system can optionally select a first parameter adjustment. Alternatively, the sound system can select the first parameter using an intelligent tuning scheme, which can be predefined or learned as described below. Thus, in step 340, the sound system can include sending configuration instructions to a remote component of the sound system (e.g., a sound card of an audio speaker) to adjust the identified parameter of the component of the sound system by a predefined amount or increment. In some implementations, the sound system can include sending configuration instructions to adjust the identified parameter to a target value in step 340.
[0035] In step 350, the acoustic system determines the effect of the first parameter adjustment in step 340 and stores relationship information of the determined relationship. In particular, in step 350, the acoustic system can iteratively perform steps 210-260 of method 200 shown in FIG. 2 and, after the first parameter adjustment, determine the impact of the adjustment. For example, in step 350, the system can determine the impact of the adjustment by comparing the difference in speech quality of the reproduced speech and the input speech determined in step 260 before and after the parameter adjustment in each iteration of steps 210-260 shown in FIG. 2. In step 350, the acoustic system can determine a positive or negative impact on the quality of the reproduced speech resulting from the first parameter adjustment (e.g., a percentage improvement or degradation in speech quality). The acoustic system can store the parameter, the increment corresponding to the parameter adjustment, and the determined impact on speech quality, which together provide information about the relationship between the first parameter and the quality of the reproduced speech for the acoustic system.
[0036] In step 360, the acoustic system determines whether the quality of the reproduced speech is satisfactory after the first parameter adjustment in step 340. For example, the acoustic system corresponds to step 270 of method 200 shown in FIG. 2. In some embodiments, the threshold value used in step 360 to determine whether the quality of the reproduced speech is satisfactory can be adjusted based on the quality of the input speech. For example, the confidence level (or equivalent) of the speech recognition result for the reproduced speech is necessarily lower than that of the input speech. Therefore, the threshold confidence level (or equivalent) can be adjusted or determined based on the quality of the reproduced speech. For example, the threshold can be a function of the value of the confidence level (or equivalent) for the input speech, e.g., fixed or a variable percentage (e.g., 90% to 95%). In response to determining that the quality of the reproduced speech is satisfactory (YES branch of step 360), the method ends in step 375. However, in response to determining that the quality of the reproduced speech is still unsatisfactory (NO branch of step 360), the method continues with step 370.
[0037] In step 370, the audio system determines whether there are any more configurable parameter adjustments that can be made. In particular, in some implementations, the audio system may cycle through a predetermined set of parameter adjustments only once as part of the corrective action of step 290 of the method shown in FIG. 2. In response to determining that there are more parameter adjustments that can be made (YES branch of step 370), the method returns to step 340 to make the next parameter adjustment. The audio system continues in a loop through steps 350-370 until the audio system determines that there are no more parameter adjustments to make. In response to determining that there are no more parameter adjustments to make (NO branch of step 370), the method ends at step 375. In other implementations, the audio system may repeatedly cycle through a set of audio system parameter adjustments until a predefined condition is met. For example, this condition may be that the quality of the reproduced speech is satisfied, that successive parameter adjustments do not result in a significant improvement in the quality of the reproduced speech, or that a timer has expired. In this case, step 370 is omitted and the sound system determines whether one or more conditions have been met in step 360; if not, method 300 returns to step 340 to make the next parameter adjustment. Method 300 then continues in a loop through steps 350-360 until the sound system determines that the quality of the audio output is satisfactory (or other conditions are met), at which point the method ends in step 375.
[0038] Thus, method 300 causes the audio system to improve the quality of the reproduced speech once the input speech is of acceptable quality by automatically adjusting the configuration of the audio system, particularly by automatically adjusting configurable parameters of the audio system to improve the quality of the reproduced speech.
[0039] Additionally, the sound system determines and stores information regarding the relationship between one or more configurable parameters of the sound system's components and speech quality. Over time, this information can be used for more intelligent adjustment of the sound system's configurable parameters. For example, in response to detecting unsatisfactory speech quality at one or more particular locations within the listening environment, this information can be used to predict the particular adjustment(s) required for a parameter or group of parameters that will provide the smallest expected difference between input and reproduced speech quality.
[0040] As one skilled in the art will recognize, there may be interdependencies among the configurable parameters of the sound system in terms of their influence or effect on speech quality. For example, adjusting a first parameter, including a positive increase in the first parameter, may induce improved but unsatisfactory speech quality, while adjusting a second parameter, including a positive increase in the second parameter, may induce degradation of speech quality, followed by adjusting a third parameter, including a negative increase in the first parameter—below its original level—to induce satisfactory speech quality. In this embodiment, the first and second parameters are interdependent—to improve speech quality, a negative adjustment of the first parameter should be combined with a positive adjustment of the second parameter. Such patterns of interdependence between the configurable parameters of the sound system and speech quality may be determined from stored information collected over a time period and used to develop an intelligent scheme for adjusting the sound system parameters in steps 340-370.
[0041] In some implementations, intelligent schemes for adjusting sound systems can be developed using machine learning. In particular, the information stored in step 350, in response to one or a series of incremental parameter adjustments, can be stored in a centralized database for one or more sound systems and used as training data for a machine learning model. This training data can additionally include information about the input audio data (e.g., input microphone type / quality, gain / amplification / volume, background noise, etc.), or information about the input speech (e.g., pitch, language, accent, etc.), or both, in addition to information about the type and configuration of the associated sound system. In this manner, a machine learning model can be developed to accurately predict the best configuration for a particular sound system for a particular type of input speech (e.g., a particular category of human speaker). The model can then be used to intelligently, simultaneously, or both, adjust multiple configurable parameters (e.g., associated with the same, different, or both audio components) to optimize output speech quality. Concurrent parameter adjustment to achieve the best predicted configuration can reduce or eliminate the need for multiple incremental parameter adjustments and iterations of steps 340 through 370. After the model is developed, the information recorded in step 350 can be used as feedback to improve model performance.
[0042] 4 is a block diagram of a system 400 according to one embodiment of the present invention. In particular, the system 400 includes processing components for an audio system as described herein.
[0043] System 400 includes a host processing system 410, a database 470, and processing devices 450 (e.g., local processing devices and user devices) at a listening location that communicate with host processing system 410 over a network 440. Network 440 can include any suitable wired or wireless data communication network, such as a mobile communication network, a local area network (LAN), a wide area network (WAN), or the Internet. Host processing system 410 can include a user interface device 460 connected to I / O unit 416. User interface device 460 can include one or more displays (e.g., screen or touch screen), a printer, a keyboard, a pointing device (e.g., a mouse, joystick, touchpad), an audio device (e.g., a microphone or speaker, or both), and any other type of user interface device.
[0044] The memory unit 414 contains audio data files 420 and one or more processing modules 430 for performing methods according to the present disclosure. The audio data files 420 include input audio data 420A associated with input microphones of a sound system. Additionally, the audio data files 420 include output audio data 420B associated with output microphones at distributed locations within the listening environment and received via the I / O unit 412 over a network 440. Each processing module 430 contains instructions for execution by the processing unit 412 for processing data, such as the audio data files 420, and / or instructions stored in the I / O unit 416 and / or memory unit 414.
[0045] According to an implementation of an embodiment of the present disclosure, the processing module 430 includes a speech evaluation module 432, a composition module 434, and a feedback module 436.
[0046] The speech evaluation module 432 is configured to evaluate the quality of reproduced speech corresponding to audio data reproduced by an audio speaker of a sound system. In particular, the speech evaluation module 432 includes a speech recognition module 432A and a detection module 432B. The speech recognition module 432A is configured to perform speech recognition on the input audio data 420A and the output audio data 420B from the audio data file 420, as exemplified by steps 220 and 240 of the method 200 of FIG. 2. The detection module 432B is configured to detect whether the quality of the reproduced speech in the output audio data 420B is unsatisfactory to a listener, as exemplified by steps 250 to 270 of the method 200 shown in FIG. 2. Accordingly, the speech evaluation module 432 retrieves and processes the audio data file 420 to perform the method 200 shown in FIG. 2. In particular, the processing performed by the speech evaluation module 432 may be performed in real time using the input audio data 420A and the output audio data 420B as described herein as it is received from the sound system.
[0047] The configuration module 434 is configured to adjust configurable parameters of the acoustic system to optimize the quality of the reproduced speech. The configuration module 434 includes a calibration module 434A, a parameter adjustment module 434B, and an adjustment evaluation module 434C. The calibration module 434A is configured to calibrate the acoustic system, for example, during setup and subsequently as needed. In particular, the calibration module 434A, in combination with the speech evaluation module 432, can calibrate the acoustic system using a pre-recorded input audio data file 420A containing speech that is considered to be “perfect” for speech recognition purposes. If the detection module 432B detects that the quality of the reproduced speech is unsatisfactory for the listener, parameter adjustments are made and evaluated using the parameter adjustment module 434B and the adjustment evaluation module 434C, as described below, until the quality of the reproduced speech is maximized. As those skilled in the art will recognize, calibration of the acoustic system using a "perfect" speech sample determines the difference that the acoustic system can achieve in a best-case scenario in determining the quality of reproduced speech compared to input speech. This can be used, for example, to set an initial threshold for satisfactory reproduced speech quality used in step 270 of method 200 shown in FIG. 2. As described above, the threshold can be adjusted during use based on the actual input speech quality. Parameter adjustment module 434B is configured to adjust configurable parameters of the acoustic system. For example, parameter adjustment module 434B may iteratively adjust parameters using an arbitrary or intelligent scheme, as exemplified by step 340 of method 300 shown in FIG. 3. Adjustment evaluation module 434C is configured to evaluate the impact of parameter adjustments made by parameter adjustment module 434B. In particular, adjustment evaluation module 434C is configured to determine whether the quality of reproduced speech is satisfactory after parameter adjustments, as exemplified by steps 350-360 of method 300 shown in FIG. 3.As described above, the parameter adjustment module 434B and the adjustment evaluation module 434C are invoked by the calibration module 434A and the detection module 432B to calibrate and reconfigure the parameters of the sound system, respectively, to optimize the quality of the reproduced speech.
[0048] The feedback module 436 is configured to provide feedback based on information obtained from the speech evaluation module 432, the calibration module 434, or both. For example, the feedback module 436 may provide feedback (e.g., a warning message) to a human speaker indicating that the quality of the input speech is unsatisfactory, as determined by the speech recognition module 432A or other analysis of the input audio data, as illustrated in steps 310 and 320 of the method 300 shown in FIG. 3 . Additionally or alternatively, the feedback module 436 may provide information regarding the impact of parameter adjustments on the quality of reproduced speech to a system or model for use in step 340 of the method 300 shown in FIG. 3 to develop or improve an intelligent parameter adjustment scheme for optimizing performance. For example, the feedback module 436 may send feedback, including information regarding the relationship between acoustic system parameters and speech quality stored in step 350 of the method 300 shown in FIG. 3 , over the network 440 to a centralized database 470 or another data storage. The stored data can be used as training data for machine learning models to optimize the performance of the sound system or as feedback to refine existing machine learning models. Additionally, the feedback module 434 can provide feedback to a user of the host processing system 410 in the event that the sound system parameters cannot be optimized to provide reproduced speech with satisfactory speech quality. For example, if the sound system determines in step 280 of the method 200 shown in FIG. 2 that there are no more instances of audio data to consider, a warning message can be sent before the method 200 ends in step 295. The warning message can provide recommendations to the sound system owner, such as suggested actions to improve the sound system's performance.For example, the warning message may instruct the owner to perform a manual check of a component of the sound system (e.g., a sound card), change the number or location of components, or change the overall output of the sound system, or a combination thereof. Techniques for recommending manual checks and changes to the sound system are known in the art and may use any suitable technique now known or developed in the future.
[0049] 4, a computer program product 480 is provided. The computer program product includes a computer-readable medium 482 having a storage medium 484 and program instructions 486 (i.e., program code) embodied therein. The program instructions 486 are configured to be loaded onto a memory unit 414 of a host processing system 410 via an I / O unit 416, such as one of a user interface device 460 or another device 450 connected to a network 440. In an example implementation, the program instructions 486 are configured to perform one or more method steps disclosed herein, such as the method steps illustrated in FIGS. 2 and 3, as described above.
[0050] While the present disclosure has been described and illustrated above with reference to example implementations, those skilled in the art will recognize that the present disclosure itself lends itself to many different variations and modifications not specifically exemplified herein.
[0051] The present invention may be embodied in a system, a method, a computer program product, or a combination thereof, and may include a computer-readable recording medium or media having computer-readable program instructions thereon for causing a processor to perform features of the present invention.
[0052] A computer-readable storage medium may be any tangible device capable of holding and storing a plurality of instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electro-magnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples of computer-readable storage media include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile diskette (DVD), a memory stick, a floppy disk, a punch card, or a mechanically encoded device having protruding structures within grooves that record instructions, and any suitable combination thereof. As used herein, a computer-readable recording medium is not to be construed as a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave such as a wave guide or other communication medium (e.g., light pulses passing through a fiber optic cable), or an electrical signal communicated through a wire.
[0053] The computer-readable programs described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or can be downloaded to an external computer or storage device via a network, such as the Internet, a local area network (LAN), a wide area network (WAN), or a wireless network, or a combination thereof. The network can include copper cables, fiber optics, wireless networks, routers, firewalls, switches, gateway computers, and edge servers, or a combination thereof. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions to a computer-readable storage medium within the computing / processing device for storage.
[0054] Computer-readable program instructions for carrying out the operations of the present invention can be either source code or object code written in any combination of programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine language instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or one or more conventional procedural programming languages such as Smalltalk®, object-oriented programming languages such as C++, the "C" programming language, or similar programming languages. The computer-readable program instructions can execute entirely on the user computer, partially on the user computer as a stand-alone software package, partially on the user computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer can be connected to the user computer through any type of network, including a local area network (LAN), a wide area network (WAN), or the connection can be to an external computer (e.g., through an Internet service provider). In some embodiments, computer-readable program instructions can be executed by electrical circuitry, including, for example, programmable logic circuitry, field programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), using state information from the computer-readable program instructions to personalize the electrical circuitry to perform features of the present invention.
[0055] Aspects of the invention described herein have been described with reference to flowchart instructions and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that any combination of flowchart illustrations and / or block diagrams and blocks in flowchart illustrations and / or block diagrams can be implemented by computer-readable program instructions.
[0056] Computer-readable program instructions can be provided to a general-purpose, special-purpose computer, or other programmable data processing device to create a computer processor or machine, which, when executed by the computer processor or other programmable data processing device, creates means for implementing the functions / operations specified in a block or blocks of the flowcharts and block diagrams, or a combination thereof. These computer-readable program instructions, which direct a computer, programmable data processing device, or other device, or a combination thereof, to function in a particular manner, can also be stored on a computer-readable recording medium, and the computer-readable recording medium having instructions stored thereon constitutes an article of manufacture containing instructions that implement the functional / operational features specified in a block or blocks of the flowcharts and block diagrams, or a combination thereof.
[0057] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device and cause a computer-implemented process to perform a series of operational steps on the computer, other programmable apparatus, or other device to implement the functions / acts identified in a block or blocks of the flowcharts and block diagrams, or a combination thereof, on the computer, other programmable apparatus, or other device.
[0058] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, the flowcharts or block diagrams may represent modules, segments, or portions of instructions, which contain one or more executable instructions for implementing a specific logical function(s). In some alternative implementations, the functions described in the blocks may be performed other than as illustrated. For example, two blocks shown in succession may actually be performed as a single step, or may be performed simultaneously, substantially simultaneously, or partially or completely overlapping in time, depending on the functionality involved, or the blocks may sometimes be performed in reverse order. It should also be noted that block diagram and / or flowchart illustrations and / or combinations thereof may be implemented by special-purpose hardware-based systems that perform specific functions or operations or execute specific-purpose hardware and computer instructions.
[0059] The description of various embodiments of the present disclosure has been reproduced for illustrative purposes, but is not intended to be exclusive or limited to the disclosed embodiments. Many modifications or variations will be apparent to those skilled in the art without departing from the scope and spirit of the present disclosure. The terms used herein have been selected to best explain the principles, practical applications, or technical improvements of the present embodiments beyond those found in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. 1. A computer-implemented method comprising: performing speech recognition on input audio data including speech input to an acoustic system; performing speech recognition on at least one instance of output audio data comprising speech reproduced by one or more audio speaker devices of the sound system; determining a difference between the speech recognition result for the input audio data and the speech recognition result for the at least one instance of the output audio data; and determining that the quality of the reproduced speech is unsatisfactory if the difference is greater than or equal to a threshold value; 11. A computer-implemented method comprising:
2. the difference comprises a quantitative value calculated from the result of the speech recognition on the input audio data as samples of input speech and the result of the speech recognition on the output audio data as samples of the reproduced speech corresponding to the samples of the input speech. The computer-implemented method of claim 1 .
3. Determining the difference between the result of the speech recognition for the input audio data and the result of the speech recognition for the at least one instance of the output audio data includes: comparing a first confidence level determined by the speech recognition for the input audio data with a second confidence level determined by the speech recognition for the at least one instance of the output audio data, wherein the first confidence level comprises a value of a confidence measure indicative of the confidence of the speech recognition for the input audio data and the second confidence level comprises a value of a confidence measure indicative of the confidence of the speech recognition for the at least one instance of the output audio data; and Determining a difference between the first confidence level and the second confidence level. The computer-implemented method of claim 1 , comprising:
4. moreover, In response to determining that the quality of the reproduced speech is unsatisfactory, performing an adjustment of one or more parameters of the sound system to improve the quality of the reproduced speech. Including, A computer-implemented method according to any one of claims 1 to 3.
5. performing adjustment of the one or more parameters of the sound system; performing a first parameter adjustment, which includes adjusting a parameter of the sound system by a defined increment or to a target value; and and in response to determining that the quality of the reproduced speech is still unsatisfactory, performing further parameter adjustments until the quality of the reproduced speech is satisfied, successive parameter adjustments do not result in an improvement in the quality of the reproduced speech, or a timer expires. The computer-implemented method of claim 4 , comprising:
6. The quality of the reproduced speech is satisfied or maximized by executing an intelligent parameter adjustment scheme of a predefined set of parameter adjustments or according to a machine learning optimization model. The computer-implemented method of claim 5 .
7. moreover, determining an impact of parameter adjustment on the quality of the reproduced speech based on a difference between the result of the speech recognition for the input audio data and the result of the speech recognition for the at least one instance of the output audio data, wherein the input audio data and the output audio data comprise input speech and the reproduced speech after the parameter adjustment; and storing relationship information including the parameters, increments corresponding to the parameter adjustments, and the impact on the quality of the reproduced speech for use as feedback to an intelligent parameter adjustment selection scheme or machine learning model for optimizing the reproduced speech; Including, 7. A computer-implemented method according to claim 4, 5 or 6.
8. moreover, determining whether the quality of speech input to the sound system is acceptable; and Responsive to determining that the quality of the speech input to the sound system is unacceptable, sending a message to a user to make changes related to the speech input to the sound system.
8. A computer-implemented method according to any one of claims 1 to 7, comprising:
9. a parameter in the one or more parameter adjustments selected from the group consisting of audio gain and equalization of an audio channel for each frequency band of a component of the sound system; A computer-implemented method according to any one of claims 4 to 7.
10. 1. An apparatus comprising: a processor and storage, the processor performing speech recognition on input audio data including speech input to the acoustic system; performing speech recognition on at least one instance of output audio data comprising speech reproduced by one or more audio speaker devices of the sound system; determining a difference between the speech recognition result for the input audio data and the speech recognition result for the at least one instance of the output audio data; and determining that the quality of the reproduced speech is unsatisfactory if the difference is greater than or equal to a threshold value; The apparatus is configured to:
11. the difference comprises a quantitative value calculated from the result of the speech recognition on the input audio data as samples of input speech and the result of the speech recognition on the output audio data as samples of the reproduced speech corresponding to the samples of the input speech.
11. The apparatus of claim 10.
12. the processor determining the difference between the result of the speech recognition for the input audio data and the result of the speech recognition for the at least one instance of the output audio data as: comparing a first confidence level determined by the speech recognition for the input audio data with a second confidence level determined by the speech recognition for the at least one instance of the output audio data, the first confidence level comprising a confidence measure value indicative of the confidence of the speech recognition for the input audio data, and the second confidence level comprising a confidence measure value indicative of the confidence of the speech recognition for the at least one instance of the output audio data; configured to determine by determining a difference between the first confidence level and the second confidence level.
11. The apparatus of claim 10.
13. The processor is further configured, in response to determining that the quality of the reproduced speech is unsatisfactory, to perform adjustment of one or more parameters of the sound system to improve the quality of the reproduced speech. An apparatus according to any one of claims 10 to 12.
14. the processor performing the adjustment of the one or more parameters of the sound system; performing a first parameter adjustment, which includes adjusting a parameter of the sound system by a defined increment or to a target value; configured to, in response to determining that the quality of the reproduced speech remains unsatisfactory, perform parameter adjustments until the quality of the reproduced speech is satisfied, successive parameter adjustments do not result in an improvement in the quality of the reproduced speech, or a timer expires.
14. The apparatus of claim 13.
15. The quality of the reproduced speech is satisfied or maximized by executing an intelligent parameter adjustment scheme of a predefined set of parameter adjustments or according to a machine learning optimization model.
15. The apparatus of claim 14.
16. The processor further comprises: determining an impact of parameter adjustment on the quality of the reproduced speech based on a difference between the result of the speech recognition for the input audio data and the result of the speech recognition for the at least one instance of the output audio data, wherein the input audio data and the output audio data comprise input speech and the reproduced speech after the parameter adjustment; storing relationship information including the parameters, increments corresponding to the parameter adjustments, and the impact on the quality of the reproduced speech for use as feedback to an intelligent parameter adjustment selection scheme or machine learning model for optimizing the reproduced speech; 16. The apparatus of claim 13, 14 or 15, configured to:
17. The processor further comprises: determining whether the quality of speech input to the sound system is acceptable; In response to determining that the quality of the speech input to the sound system is unacceptable, sending a message to a user to make changes related to the speech input to the sound system. The device according to any one of claims 10 to 16, configured to:
18. a parameter in the one or more parameter adjustments selected from the group consisting of audio gain and equalization of an audio channel for each frequency band of a component of the sound system; 17. Apparatus according to any one of claims 13 to 16.
19. A computer program for causing a computer to execute the computer-implemented method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Method and apparatus for adjusting audio signal, terminal and computer readable storage medium
CN107484081A
Method and system for speech quality perception evaluation based on speech semantic recognition technology
CN108877839A
Technique for displaying contents of speech synchronously with reproduction of speech
JP2012198552A
Detecting potential significant errors in speech recognition results
US20140012579A1