IMPROVING THE AUDIO QUALITY OF SPEECH IN SOUND SYSTEMS

DE112020003875B4Active Publication Date: 2026-09-03INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
DE112020003875
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-10-05
Filing Date
2020-10-05
Publication Date
2026-09-03
Estimated Expiration
2040-10-05

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method (200) implemented by a computer (410), comprising: performing (220) speech recognition from input audio data (420A) comprising speech input into a sound reinforcement system; performing (240) speech recognition from at least one instance of output audio data (420B) comprising speech reproduced by one or more loudspeakers of the sound reinforcement system; determining (260) a difference between a speech recognition result from the input audio data (420A) and a speech recognition result from the at least one instance of the output audio data; determining (270) that the quality of the reproduced speech is unsatisfactory if the difference is greater than or equal to a threshold value;and performing (290) one or more parameter adjustments of the sound reinforcement system to improve the quality of the speech reproduced, in response to a finding that the quality of the speech reproduced is unsatisfactory, wherein parameters of the one or more parameter adjustments are selected from a group which necessarily includes equalization of the audio channels for each frequency band of a component of the sound reinforcement system and optionally additionally includes audio amplification.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND The present invention relates generally to techniques for improving the quality of the audio output of a sound reinforcement system and in particular to improving the speech quality for listeners of the audio output of the sound reinforcement system. Public address systems are frequently used to reproduce speech via loudspeakers for listeners at geographically dispersed locations, such as participants in conferences, lectures, and performances in a theater or auditorium, or participants in telephone conferences and webinars, over a data transmission network. In such systems, speech input via a microphone is received and optionally recorded by a host system. Audio data is then transmitted from the host system to one or more loudspeakers, and the loudspeaker(s) reproduce the speech to the listeners (i.e., in the form of a "replay"). In many cases, the speech reproduced via the loudspeakers is not a perfect reproduction of the input speech (e.g., the speech may be unclear).For example, if a loudspeaker's settings are not optimal, the reproduced sound, and therefore the speech, may be distorted, making it difficult for listeners to hear and / or understand. In other cases, the input speech itself may be flawed, for example, due to the microphone's position relative to the speech source or its suboptimal settings. This, in turn, makes it difficult for listeners to hear or understand the speech reproduced by the loudspeakers. Typically, such speech intelligibility issues can be resolved by adjusting the sound system. For instance, if listeners inform the presenter that the speech is difficult to hear or understand, the presenter can adjust the sound system's settings or ask the speaker to change their position relative to the microphone.However, this leads to interruptions and delays while the adjustments are made. Since these are manual adjustments, the listeners' difficulties may not be completely resolved. US 2016 / 0293180A1 discloses a system in which a loudspeaker transmits sound with dynamic volume and a receiving device provides feedback to adjust the volume based on the received sound. DE 10 2008 030 086 A1 discloses a system in which a speech recognition system is used on the receiver side for the automatic evaluation of speech quality, whereby a confidence measure is determined and transmitted as feedback to the speaker via a separate transmission channel. SUMMARY According to one aspect of the present invention, a computer-implemented method is provided. The computer-implemented method comprises performing speech recognition based on input audio data comprising speech input into a sound system. The computer-implemented method further comprises performing speech recognition based on at least one instance of output audio data comprising speech reproduced by one or more loudspeakers of the sound system. The computer-implemented method further comprises determining a difference between a result of speech recognition based on the input audio data and a result of speech recognition based on the at least one instance of the output audio data.The computer-implemented procedure further comprises determining that the quality of the reproduced speech is unsatisfactory when the difference is greater than or equal to a threshold value; and performing one or more parameter adjustments of the sound reinforcement system to improve the quality of the reproduced speech in response to a determination that the quality of the reproduced speech is unsatisfactory, wherein parameters of the one or more parameter adjustments are selected from a group that necessarily includes equalization of the audio channels for each frequency band of a component of the sound reinforcement system and optionally additional audio amplification. According to a further aspect of the present invention, a device is provided. The device comprises a processor and a data storage device. The processor is configured to perform speech recognition based on input audio data comprising speech input into a sound system. The processor is further configured to perform speech recognition based on at least one instance of output audio data comprising speech reproduced by one or more loudspeakers of the sound system. The processor is further configured to determine a difference between a result of speech recognition based on the input audio data and a result of speech recognition based on the at least one instance of the output audio data.The processor is further configured to detect that the quality of the reproduced speech is unsatisfactory when the difference is greater than or equal to a threshold value, and to perform one or more parameter adjustments of the sound reinforcement system to improve the quality of the reproduced speech in response to a detection that the quality of the reproduced speech is unsatisfactory, whereby parameters of the one or more parameter adjustments are selected from a group that necessarily includes equalization of the audio channels for each frequency band of a component of the sound reinforcement system and optionally additional audio amplification. According to a further aspect of the present invention, a computer program product is provided. The computer program product comprises a computer-readable storage medium containing program instructions.The program instructions are executable by a processor to cause the processor to: perform speech recognition from input audio data containing speech input into a sound reinforcement system; perform speech recognition from at least one instance of output audio data containing speech reproduced by one or more loudspeakers of the sound reinforcement system; determine a difference between a result of speech recognition from the input audio data and a result of speech recognition from the at least one instance of output audio data; and determine that the quality of the reproduced speech is unsatisfactory if the difference is greater than or equal to a threshold value.The processor is further configured to perform one or more parameter adjustments of the sound reinforcement system to improve the quality of the reproduced speech, in response to a finding that the quality of the reproduced speech is unsatisfactory, whereby parameters of one or more parameter adjustments are selected from a group that necessarily includes equalization of the audio channels for each frequency band of a component of the sound reinforcement system and optionally additional audio amplification. Brief description of the different views of the drawings Exemplary implementations of the present disclosure are described below with reference to the accompanying drawings. Fig. 1 is a schematic representation of a sound reinforcement system according to an embodiment of the present invention. Fig. 2 is a flowchart for a method for detecting and correcting unsatisfactory speech quality according to an embodiment of the present invention. Fig. 3 is a flowchart for a method for adjusting a sound reinforcement system according to an embodiment of the present invention to improve speech quality. Fig. 4 is a block diagram of a sound reinforcement system according to an embodiment of the present invention. DETAILED DESCRIPTION This disclosure provides systems and methods for detecting when the quality of speech reproduced by a sound system is unsatisfactory (e.g., difficult to hear or unclear / incoherent for the listeners), and for reconfiguring the sound system to improve the quality of the reproduced speech. The techniques of this disclosure can be performed automatically and in real time to limit interruptions and improve the listeners' experience. According to the present disclosure, one or more microphones are used, distributed at positions in the listening environment, to capture speech reproduced by one or more loudspeakers of a sound reinforcement system. The output audio data, containing the reproduced speech received by each of the one or more microphones, is recorded. Speech recognition is performed using the output audio data belonging to each of the one or more microphones to determine the quality of the speech reproduced at the corresponding microphone position. Furthermore, speech recognition is performed using the input audio data, containing the speech input into the sound reinforcement system, to determine the quality of the speech originating from the source.For each microphone, a comparison is made between the speech recognition result performed on the input audio data and the speech recognition result performed on the corresponding output audio data. Based on this comparison, it is determined whether the speech quality is unsatisfactory for the listeners, and if so, corrective measures are taken, such as adjustments to the sound system, to improve the quality of the reproduced speech. In this disclosure, the term "speech" is used to refer to sounds or audio data containing speech. The term "input audio data" refers to digital audio data for sounds or audio data containing speech originating from a source (e.g., a human speaker) that is captured by a microphone of the sound reinforcement system (here, "input microphone"). The term "output audio data" refers to digital audio data for sounds or audio data containing speech that is reproduced by one or more loudspeakers of a sound reinforcement system and captured by a microphone of a sound reinforcement system (here, "output microphone"). Audio data thus "represents" or "contains" sounds or audio data containing speech that is received by a microphone of a sound reinforcement system.The term "recording" audio data refers to storing audio data in data storage devices, which includes both caching audio data for data transmission over a network and temporary and long-term storage of audio data as audio files. Fig. 1 is a schematic representation of a sound reinforcement system according to an embodiment of the present invention. The sound reinforcement system 100 comprises a host processing system 110, a plurality of microphones 120, and a plurality of loudspeakers 130, which are interconnected via a data transmission network 140. In the system shown in Fig. 1, the sound reinforcement system 100 is a distributed system that has the microphones 120 and the loudspeakers 130 at different locations (e.g., in meeting or conference rooms). At least one location (location 1) has the host processing system 110, which in the illustrated example is the location of the source of the speech input into the sound reinforcement system 100. Other locations (locations 2 and 3) have the loudspeakers 130, which reproduce the audio data for listeners in the respective listening environments. The host processing system 110 typically includes a user data processing system (e.g., a notebook), a dedicated control unit for the sound system, or similar equipment that can be used by a user to operate the sound system 100. Most microphones 120 have an input microphone 122 to capture and record speech from a source (e.g., a human speaker) for playback by the sound system 100. The input microphone 122 can be a dedicated microphone for receiving sound input into the sound system 100 (e.g., a microphone on a lectern) or similar equipment, or it can be a microphone of a user data processing system that can be "turned on" under the control of the host processing system 110.The majority of the loudspeakers 130 reproduce the input speech and are located at various positions in one or more locations that constitute listening environments (locations 2 and 3). Each loudspeaker 130 receives audio data corresponding to the recorded speech from the host processing system 110 via the network 140 and reproduces it. The loudspeakers 130 may include one or more dedicated loudspeakers 134 of a sound reinforcement system at the location (e.g., a loudspeaker at a fixed location in a theater) or loudspeakers 132 of user data processing systems, networked telephones, or similar devices. The data transmission network 140 may comprise any suitable wired or wireless network for data transmission between the host processing system 110, the microphones 120, and the loudspeakers 130. According to one embodiment, the majority of the microphones 120 also include output microphones 124 located at positions in the listening environments (locations 2 and 3) to receive and record the speech reproduced by the loudspeakers 130 for analysis, as described herein. The output microphones 124 may be dedicated microphones of a sound reinforcement system, for example, belonging to the loudspeakers 134 at the location. The output microphones 124 may also be microphones of user data processing systems or other units in the listening environments, which can be recognized and used by the host processing system 110. In the distributed system of Fig. 1, the output microphones 124 in a listening environment (location 3) each belong to a local processing system 150, for example, a system unit, to record the reproduced speech as output audio data for transmission to the host processing system 110 via the network 140.The output microphones 124 in a different listening environment (location 2) are configured to transmit audio data over the network 140 to the host processing system 110, which records the output audio data. As can be seen by a person skilled in the art, the majority of the microphones 120 may be digital microphones that generate digital output signals and / or analog microphones that generate analog output signals, which are received by another component along the audio signal chain and converted into digital audio data. The host processing system 110 records the speech received from the source via the input microphone 122 as input audio data. In addition, the host processing system 110 receives output audio data corresponding to the speech reproduced by the loudspeakers 130, which is recorded by each output microphone 124 and transmitted via the network 140. According to one embodiment of the present invention, the host processing system 110 is configured to perform speech recognition based on input audio data belonging to the input microphone 122 and on output audio data belonging to each output microphone 124. Speech recognition techniques are known in the art, and the host processing system 110 can implement any suitable speech recognition technique.Speech recognition can produce a transcript of the speech and / or a value or grade of reliability measurement, or similar, indicating the reliability of the speech recognition. Speech recognition can be performed continuously or periodically using the input audio data and the corresponding output audio data. The host processing system 110 further compares a speech recognition result obtained for the output audio data associated with each output microphone 124 with a speech recognition result obtained for the corresponding input audio data associated with the input microphone 122.If the comparison reveals an unacceptable difference between the result determined for the output audio data of one or more output microphones 124 and the result determined for the input audio data of the input microphone 122, the host processing system 110 determines that the speech quality of the reproduced speech is unsatisfactory for the listeners and takes corrective action. Corrective action may include, for example, adjusting parameters of components of the sound reinforcement system (e.g., gain or channel equalization settings of a sound card controlling a loudspeaker) or sending a message to a user to take specific actions (e.g., instructions for a human speaker to move closer to or further away from the input microphone). An unacceptable difference might, for example,a threshold will be determined if a difference in the comparison results derived from speech recognition, e.g., a difference in the reliability level or a measured difference in the speech transcript, as described below, is less than or equal to a threshold value. Accordingly, the host processing system 110 can detect when the quality of the speech reproduced by the sound system 100 is unsatisfactory for the listeners (i.e., unclear, distorted, or too quiet, etc.) and take measures to improve the speech quality. Since the disclosed techniques can be performed automatically and in real time, the listening experience is enhanced. In the embodiment shown in Fig. 1, the disclosed techniques are performed in the host processing system 110. As is apparent to those skilled in the art, the present invention can be implemented in any other processing unit or system that exchanges data with the sound system 100, e.g., in a local processing system 150 or in a combination of processing units. Fig. 2 shows a flowchart for a method 200 for detecting and correcting unsatisfactory speech quality according to an embodiment of the present invention. The method 200 can, for example, be carried out by the host processing system 110 of the sound reinforcement system 100 of Fig. 1. Procedure 200 begins at step 205. Step 205 may be initiated, for example, in response to the start of a sound check of the sound system, in response to the start of a conversation, or in some other way. In step 210, the sound reinforcement system receives input audio data from a source for the speech input into the system. The input audio data can be received, for example, by the input microphone 122 of the sound reinforcement system 100 shown in Fig. 1 in response to the speech of a human speaker speaking into the input microphone 122. The input audio data is typically received and recorded (e.g., stored in a file containing input audio data) essentially in real time (i.e., with minimal time delay), for example, by the host processing system 110 shown in Fig. 1. In step 220, the sound reinforcement system performs speech recognition on the input audio data to determine a speech recognition result that indicates the speech quality. Any suitable speech recognition technique or algorithm can be used to perform the speech recognition. The speech recognition result typically includes a transcript of the speech contained in the corresponding audio data. The quality of the transcript can be an indicator of the speech quality. Furthermore, the speech recognition result can include a value or grade of reliability measurement or similar indicator that specifies the reliability of the speech recognition. Such reliability measurements are known in the field of speech recognition. The reliability grade can therefore also be an indicator of speech quality. The speech recognition process can provide further results that indicate speech quality.In step 230, the sound reinforcement system receives output audio data for corresponding speech, which is reproduced by loudspeakers of the sound reinforcement system at one or more positions in the listening environments. For example, each of the one or more output microphones 124 can receive an instance of output audio data capturing reproduced speech, which is reproduced by the loudspeakers 130 of the sound reinforcement system 100 shown in Fig. 1. The output audio data is generally received essentially in real time, but is necessarily delayed with respect to the input audio data. For example, the transmission of the input audio data to the loudspeakers and the transmission of the corresponding output audio data belonging to the output microphones 124 via the network 140 shown in Fig. 1 introduce a time delay. In some implementations, the output audio data is recorded, e.g.by the local processing system 150 and / or the host processing system 110 shown in Fig. 1. In step 240, the sound system performs speech recognition on the output audio data to determine a speech recognition result that indicates the speech quality. Specifically, in step 240, the sound system performs speech recognition on each received instance of output audio data. In step 240, the sound system uses the same speech recognition technique employed in step 220, ensuring that the results in step 220 and step 240 are comparable. Accordingly, in steps 210 and 220, the sound reinforcement system derives one or more speech recognition results for the input audio data, and in steps 230 and 240, the sound reinforcement system derives one or more speech recognition results for each instance of the output audio data. In each case, the speech recognition result includes a speech transcript and / or a reliability grade or similar. As is evident to a person skilled in the art, steps 210 to 240 can be performed simultaneously in practice, particularly in applications where the input and output audio data are received and processed continuously, e.g., in real time. In the exemplary implementation shown in Fig. 2, multiple instances of output audio data are received. Each instance of output audio data is specifically associated with a particular output microphone positioned in a listening environment. In step 250, the sound reinforcement system selects a speech recognition result for a first instance of output audio data. In step 260, the sound reinforcement system compares the selected speech recognition result for the output audio instance with the speech recognition result for the corresponding input audio and determines a difference in speech quality. This difference is a quantitative value representing the difference in speech quality calculated from the speech recognition result(s). In one implementation, the sound reinforcement system in step 260 might compare the text of a speech recognition transcript obtained for the output audio instance with the text of a transcript of the corresponding input audio and determine a difference, such as a rough numerical difference or a percentage difference in the transcribed text (e.g., in the number of words).Differences in the text of the transcribed speech indicate that the speech quality at the loudspeakers is reduced compared to the speech quality at the source, and the magnitude of the difference indicates the extent of the quality degradation. In another implementation, the sound reinforcement system in step 260 can compare a speech recognition reliability score determined for the output audio data instance with the reliability score determined for the corresponding input audio data and identify any difference. As described above, the reliability score indicates the reliability of the transcribed speech (e.g., as a reliability value expressed as a percentage).Since the reliability of the transcribed speech depends on the quality of the reproduced speech for the listeners, a difference in the reliability level indicates that the quality of the reproduced speech at the loudspeakers is reduced compared to the quality of the speech at the source, and the magnitude of the difference indicates the extent of the quality degradation. In other implementations, the sound reinforcement system can compare other measurements derived from the speech recognition results to detect a reduction in speech quality. As can be seen by those skilled in the art, in step 260, the sound reinforcement system can compare a speech recognition result for a sample of the output audio data with a speech recognition result for a corresponding sample of the input audio data. In some scenarios, steps 210 to 240 can be performed continuously using the input and reproduced speech (e.g.,(essentially in real time). In this case, corresponding samples of input and output audio data containing the same input and playback language can be identified using any suitable technique, such as time synchronization and / or audio matching (identifying the same audio segments in the audio data) or speech recognition transcript matching (identifying the same text segments by matching words and phrases in the transcribed text). In other scenarios, steps 210 to 240 can be performed by periodically sampling the input and playback language (e.g., sampling the input and playback language over synchronized, time-separated time windows) so that the speech recognition result(s) relate to corresponding samples of the input and output audio data. In step 270, the sound reinforcement system determines whether the difference is greater than or equal to a threshold value. The threshold value is a difference measurement (e.g., count / percentage or words in text, or reliability measurement) that represents an unacceptable reduction in the speech quality of the reproduced speech compared to the quality of the input speech. The threshold value can be selected according to the application requirements and can be changed by the user. For example, in some applications, a difference in speech quality of up to 5% may be acceptable, so the threshold is set to 5%, while in other applications, a difference in speech quality of up to 10% may be acceptable, so the threshold is set to 10%. In some example implementations, the threshold can be adjusted based on the quality of the input speech, as described below. If the difference is less than the threshold (NO branch from step 270), the quality of the reproduced speech in the selected instance of the output audio data is satisfactory, and the sound reinforcement system executes step 280. In step 280, the sound reinforcement system determines whether further instances of output audio data need to be considered. If further instances of output audio data need to be considered (YES branch from step 280), the sound reinforcement system begins executing step 250 and then loops through steps 260 and 270 until, in step 280, the system determines that there are no further instances of output audio data to consider. After determining that there are no further instances of output audio data to consider, the sound reinforcement system terminates execution in step 295. If the difference is greater than or equal to the threshold (branching YES from step 270), the quality of the reproduced speech in the selected instance of the output audio data is unsatisfactory, and the sound reinforcement system executes step 290. In step 290, the sound reinforcement system performs a corrective action to improve the quality of the speech reproduced by the system. The corrective action may include, for example, changing configurable parameters of the sound reinforcement system and / or sending messages to the users, as described below with reference to Figure 3. As can be seen by those skilled in the art, many variations of the exemplary implementation shown in Fig. 2 are possible. For example, speech recognition can be performed using output audio data belonging to each output microphone at a corresponding local processing unit or user unit associated with it. Thus, in step 240, the sound reinforcement system can instead receive a speech recognition result for output audio data belonging to each output microphone via network 140, and step 230 can be omitted. In this way, the processing load for speech recognition is distributed across multiple processing units. Furthermore, prior to step 210, the sound reinforcement system can identify available microphones at positions in the listening environments and select a set of microphones to use as output microphones.For example, microphones on user units can be identified based on an existing connection to Network 140 from a listening environment or based on the coordinates of a global positioning system (or a comparable system) in a listening environment. In this case, a message can be sent to the user requesting permission to use an identified microphone on the user unit to listen to the speech being played back, and the user can decide whether to grant or deny permission. If permission is granted, the public address system can then activate the microphone and any other functions required for the user unit to send the output audio data in step 230. In another example, standalone microphones connected to the network in the listening environment can be identified and used to send output audio data (e.g.,(if the configured permissions of the microphone(s) allow it). Furthermore, the sound reinforcement system performs corrective action if, in step 270, it determines that the quality of the reproduced speech in a single selected instance of the output audio data is unsatisfactory. In other implementations, corrective action may be performed based on other criteria. For example, the sound reinforcement system performs corrective action if it determines that the quality of the reproduced speech is unsatisfactory for multiple instances of output audio data. In another example, corrective action may be performed based on the position of the output microphone(s) in the listening environment associated with the unsatisfactorily reproduced speech. Fig. 3 shows a flowchart for a method 300 for adjusting a sound reinforcement system according to an embodiment of the present invention to improve speech quality. The method 300 can be carried out, for example, as a corrective measure in step 290 shown in Fig. 2. The method 300 can be carried out, for example, by the host processing system 110 of the sound reinforcement system 100 shown in Fig. 1 or by another processing unit of the sound reinforcement system. Procedure 300 begins at step 305. Procedure 300 might begin, for example, in response to the detection of a difference between the speech recognition result(s) of the output audio data and the input audio data that is greater than or equal to a threshold. If the difference is greater than or equal to a threshold, this indicates that the quality of the reproduced speech is unsatisfactory. In step 310, the sound system checks the speech quality of the input from the source. For example, in step 310, the sound system might compare a speech recognition result with a speech quality threshold. The threshold could be a predefined reliability level (e.g., 60%). The threshold can also be configured by a user. Falling below the threshold indicates that the input speech quality is unacceptable. In other examples, in step 310, the sound system might process the input audio data using one or more techniques that identify problems negatively impacting the input speech quality, such as...A high volume of background noise relative to the input speech (recognizable by a low signal-to-noise ratio), the speaker's position relative to the input microphone (recognizable by "pops"), the input microphone settings (e.g., sensitivity or gain / volume level), and similar factors. Thus, in step 310, the sound system can perform a series of tests to identify potential problems related to the audio input. In step 320, the sound system determines whether the quality of the speech input from the source is acceptable. For example, in step 320, the sound system may determine whether it detected a problem with the input speech in step 310 indicating that the speech quality is unacceptable. If the system determines that the input speech quality is acceptable (YES branch from step 320), it proceeds to step 340. However, if it determines that the input speech quality is unacceptable (NO branch from step 204), the sound system proceeds to step 330. In step 330, the public address system sends a warning message to a user at the source. This warning message may contain instructions to make adjustments based on the test result(s) from step 310. For example, if the speech recognition result performed on the input audio data is below the threshold, the warning message may instruct the speaker to speak more clearly. In another example, if a test detects a problem with the speaker's position relative to the input microphone, the message may instruct the speaker to move closer to or further away from the microphone. In yet another example, if a test detects a problem with the input microphone, the warning message may provide instructions for adjusting the microphone settings (e.g., tone sensitivity or gain / volume level).In other implementations, in scenarios where the sound reinforcement system detects a problem with the input microphone in step 310, automatic adjustment of the input microphone can be performed, e.g., using steps 340 to 370 described below. In step 340, the sound reinforcement system performs an initial parameter adjustment. The adjusted parameter can be any independently configurable parameter or setting of an individual component of the sound reinforcement system, such as a sound card, loudspeaker, or microphone. Suitable parameters might include gain and equalization settings, and similar adjustments, for an audio component. As is evident to the expert, equalization settings comprise adjustable settings for multiple frequency ranges (also referred to as frequency bands or channels) of the audio signals. Thus, with regard to equalization, each adjustable frequency band of a component corresponds to an adjustable parameter. Accordingly, the adjustable parameters of a sound reinforcement system have configurable settings, such as gain and equalization settings, for each configurable component of the system.Parameter adjustment can involve a positive or negative increment of the parameter value of a specific audio component, such as a loudspeaker. Parameter adjustment can be defined as an increment of an existing parameter value or as a new (target) value for the component's parameter. In step 340, the sound reinforcement system can randomly select the first parameter adjustment. Alternatively, the sound reinforcement system can select the first parameter adjustment using an intelligent selection scheme, which can be predefined or learned as described below. Similarly, in step 340, the sound reinforcement system can send configuration instructions to a remotely located component of the sound reinforcement system (e.g., a loudspeaker's sound card) to adjust an identified parameter of a component of the sound reinforcement system by a specific value or increment.In some implementations, the sound reinforcement system may include sending configuration instructions in step 340 to adjust an identified parameter to a target value. In step 350, the sound reinforcement system determines the effects of the first parameter adjustment in step 340 and stores information about the determined relationship. In particular, in step 350, the sound reinforcement system can perform an iteration from step 210 to step 260 of the procedure 200 shown in Fig. 2 after the first parameter adjustment and determine the effects of the adjustment. For example, in step 350, the system can determine the effects of the adjustment by comparing the difference in speech quality between the reproduced speech and the input speech, determined in step 260 in the iteration from step 210 to step 260 of the procedure 200 shown in Fig. 2, before and after the parameter adjustment. In step 350, the sound reinforcement system can determine positive or negative effects on the quality of the reproduced speech resulting from the first parameter adjustment (e.g.,(a percentage improvement or deterioration in speech quality). The sound system can store the parameter and the increment corresponding to the parameter adjustment and the determined effects on speech quality, which together provide information about the relationship between the first parameter and the quality of the reproduced speech for the sound system. In step 360, the sound reinforcement system determines whether the quality of the reproduced speech is satisfactory after the initial parameter adjustment in step 340. The sound reinforcement system can, for example, correspond to step 270 of method 200 shown in Fig. 2. In some implementations, the threshold used in step 360 to determine whether the quality of the reproduced speech is satisfactory can be adjusted based on the quality of the input speech. For example, the reliability level (or equivalent) of the speech recognition result for the reproduced speech is necessarily lower than for the input speech. The threshold for the reliability level (or equivalent) can therefore be adjusted or determined based on the quality of the reproduced speech. The threshold can, for example,The reliability level (or equivalent) for the input language could be a function of that level, for example, a fixed or variable percentage (e.g., 90% to 95%). If the procedure determines that the quality of the reproduced language is satisfactory (branching YES from step 360), it terminates at step 375. However, if the procedure determines that the quality of the reproduced language is still unsatisfactory (branching NO from step 360), it continues at step 370. In step 370, the sound system determines whether further configurable parameter adjustments are required. In some implementations, particularly as part of the corrective action in step 290 of the procedure shown in Fig. 2, the sound system may only execute a predefined set of parameter adjustments once. Upon determining that further parameter adjustments are possible (branching YES from step 370), the procedure returns to step 340, where the next parameter adjustment is made. The sound system continues in a loop from step 350 to step 370 until it determines that no further parameter adjustments are required. Upon determining that no further parameter adjustments are required (branching NO from step 370), the procedure terminates at step 375.In other implementations, the sound reinforcement system can repeatedly cycle through a set of parameter adjustments until a predefined condition is met. This condition might be, for example, that the speech quality is satisfactory, that successive parameter adjustments do not significantly improve the speech quality (i.e., speech quality is maximized), or that a timer expires. In this case, step 370 can be omitted. The sound reinforcement system determines in step 360 whether one or more of the conditions are met, and if not, the process returns to step 340, where the next parameter adjustment is made.Procedure 300 then continues in a loop from step 350 to step 360 until the sound system determines that the audio output quality is satisfactory (or another condition is met), and the procedure ends at step 375. Accordingly, the sound system using method 300 improves the quality of the reproduced speech, provided the input speech is of acceptable quality, by automatically adjusting the system's configuration. Specifically, the sound system automatically adjusts configurable parameters to improve the quality of the reproduced speech. Furthermore, the sound reinforcement system collects and stores information about the relationship between one or more configurable parameters of the system's components and speech quality. This information can then be used to more intelligently adjust the system's configurable parameters. For example, in response to the detection of unsatisfactory speech quality at one or more specific locations in the listening environment, the system can predict the specific adjustment(s) required for a particular parameter or group of parameters to achieve a minimal expected difference between the input and reproduced speech quality. As is evident to the expert, the configurable parameters of the sound system can be interdependent with regard to their effects on speech quality. For example, an initial parameter adjustment involving a positive increment of the first parameter leads to improved, but not satisfactory, speech quality; a second parameter adjustment involving a positive increment of the second parameter leads to decreased speech quality; however, a subsequent third parameter adjustment involving a negative increment of the first parameter—below its original value—results in satisfactory speech quality. In this example, the first and second parameters are interdependent—a negative adjustment of the first parameter should be combined with a positive adjustment of the second parameter to improve speech quality.Such patterns of dependence between configurable parameters of the sound system and speech quality can be determined from the stored information collected over a certain period of time and used to develop intelligent schemes for adjusting the parameters of the sound system from step 340 to step 370. In some implementations, intelligent schemes for adapting the sound reinforcement system can be developed using machine learning. The information stored in step 350 in response to one or more incremental parameter adjustments can be stored in a central database for one or more sound reinforcement systems and used as training data for a machine learning model. This training data can additionally include information about the input audio data (e.g., type / quality of the input microphone, gain / amplitude / volume, background noise, etc.) and / or information about the input speech (e.g., pitch, speech, accent, etc.), as well as information about the type and layout of the sound reinforcement system in question.In this way, a machine learning model can be developed to accurately predict the best configuration for a given sound reinforcement system for a specific type of input speech (e.g., a particular category of human speaker). The model can then be used to intelligently and / or simultaneously adjust multiple configurable parameters of the sound reinforcement system (e.g., relating to the same and / or different audio components) to optimize the output speech quality. Simultaneous parameter adjustments to achieve a predicted best configuration can increase or eliminate the need for multiple incremental parameter adjustments and iterations from step 340 to step 370. After the model is developed, the information recorded in step 350 can be used as feedback to improve the model's performance. Fig. 4 is a block diagram of a system 400 according to an embodiment of the present invention. The system 400 includes, in particular, processing components for a sound reinforcement system as described herein. System 400 comprises a host processing system 410, a database 470, and processing units 450 (e.g., local processing units and user units) at listener locations that exchange data with the host processing system 410 via a network 440. The network 440 can be any suitable wired or wireless data transmission network, e.g., a cellular network, a local area network (LAN), a wide area network (WAN), or the Internet. The host processing system 410 comprises a processing unit 412, a storage unit 414, and an input / output (I / O) unit 416. The host processing system 410 can include input / output interfaces 460 connected to the I / O unit 416. The user interface units 460 can include one or more displays (e.g., screen or touchscreen), a printer, a keyboard, a pointing device (e.g., mouse, joystick, touchpad), an audio device (e.g., microphone, microphone, etc.).microphone and / or speaker) and any other type of user interface unit. The storage unit 414 comprises the audio data files 420 and one or more processing modules 430 for performing methods according to the present disclosure. The audio data files 420 include the input audio data 420A, which is associated with an input microphone of the sound reinforcement system. In addition, the audio data files 420 include output audio data 420B, which is associated with output microphones at distributed positions in the listening environments and is received via the I / O unit 416 over the network 440. Each processing module 430 contains instructions for execution by the processing unit 412 to process data and / or instructions received by the I / O unit 416 and / or stored in the storage unit 414, e.g., the audio data files 420. According to exemplary implementations of the present disclosure, the processing modules 430 comprise a language evaluation module 432, a configuration module 434 and a feedback module 436. The speech evaluation module 432 is configured to evaluate the quality of the reproduced speech in the output audio data 420A, which corresponds to the audio data reproduced by the loudspeakers of the sound reinforcement system. The speech evaluation module 432 specifically comprises a speech recognition module 432A and a capture module 432B. The speech recognition module 432A is configured to perform speech recognition from input audio data 420A and output audio data 420B from audio data files 420, for example, as in steps 220 and 240 of method 200 shown in Fig. 2. The capture module 432B is configured to detect when the quality of the reproduced speech in the output audio data 420B is unsatisfactory for listeners, for example, as in steps 250 to 270 of method 200 shown in Fig. 2.Accordingly, the speech evaluation module 432 retrieves and processes the audio data files 420 to perform the method 200 shown in Fig. 2. The processing by the speech evaluation module 432 can, as described here, be carried out in real time, in particular, using the input audio data 420A and the output audio data 420B as received by the sound reinforcement system. Configuration Module 434 is configured to adjust the configurable parameters of the sound reinforcement system to optimize the quality of the reproduced speech. Configuration Module 434 comprises a Calibration Module 434A, a Parameter Adjustment Module 434B, and an Adjustment Evaluation Module 434C. Calibration Module 434A is configured to calibrate the sound reinforcement system, for example, during setup and later as needed. In particular, Calibration Module 434A, in conjunction with Speech Evaluation Module 432, can calibrate the sound reinforcement system using a pre-recorded input audio data file 420A containing speech considered "perfect" for speech recognition purposes.If the acquisition module 432B determines that the quality of the reproduced speech is unsatisfactory for listeners, parameter adjustments are made and evaluated using the parameter adjustment module 434B and the adjustment evaluation module 434C as described below, until the quality of the reproduced speech is maximized. As can be seen by those skilled in the art, calibrating the sound reinforcement system using a "perfect" speech sample determines the difference in the determined quality of the reproduced speech compared to the input speech that can be achieved with the sound reinforcement system under optimal conditions. This can be used, for example, to establish the initial threshold for satisfactory quality of the reproduced speech, which is used in step 270 of method 200 shown in Fig. 2.As described above, the threshold can be adjusted during use based on the quality of the actual speech input. The parameter adjustment module 434B is configured to adjust the configurable parameters of the sound reinforcement system. For example, the parameter adjustment module 434B can iteratively perform parameter adjustments according to an arbitrary or intelligent scheme, such as in step 340 of procedure 300 shown in Fig. 3. The adjustment evaluation module 434C is configured to evaluate the effects of a parameter adjustment performed by the parameter adjustment module 434B. In particular, the adjustment evaluation module 434C is configured to determine whether the quality of the reproduced speech is satisfactory after a parameter adjustment, as in steps 350 and 360 of procedure 300 shown in Fig. 3.As described above, the parameter adjustment module 434B and the adjustment evaluation module 434C can be called by the calibration module 434A and the acquisition module 432B to configure or reconfigure the parameters of the sound reinforcement system so that the quality of the reproduced speech is optimized. The feedback module 436 is configured to provide feedback on information obtained from the operation of the speech evaluation module 432 and / or the configuration module 434. For example, the feedback module 436 can provide feedback (e.g., a warning message) to a human speaker indicating that the quality of the input speech, as determined by the speech recognition module 432A or by other analysis of the input audio data, is unsatisfactory, as in steps 310 and 320 of procedure 300 shown in Fig. 3. Additionally or alternatively, the feedback module 436 can provide a system or model with information about the effects of parameter adjustments on the quality of the reproduced speech, in order to develop or improve an intelligent parameter adjustment scheme so that performance is optimized, as in step 340 of procedure 300 shown in Fig. 3.The feedback module 436 can be used in the procedure 300 shown in Fig. 3. For example, the feedback module 436 can send feedback containing information about the relationship between parameters of the sound reinforcement system and the speech quality, which was stored in step 350 of the procedure 300 shown in Fig. 3, via the network 440 to a centralized database 470 or to another data storage device. The stored data can be used as training data for a machine learning model to optimize the performance of a sound reinforcement system or as feedback to refine an existing machine learning model. In addition, the feedback module 436 can provide feedback to a user of the host processing system 410 if the configuration module 434 is unable to optimize the parameters of the sound reinforcement system to provide the reproduced speech with satisfactory speech quality.For example, a warning message can be sent when the sound reinforcement system, in step 280 of the procedure 200 shown in Fig. 2, detects that no further instances of output audio data need to be considered before the procedure 200 terminates in step 295. This warning message can provide the operator of the sound reinforcement system with recommendations, such as suggested measures to improve the system's performance. The warning message can, for example, instruct the operator to perform manual checks of the sound reinforcement system's components (e.g., sound cards), to change the number or location of the components, and / or to adjust the overall performance of the sound reinforcement system. Techniques for providing recommendations for manual checks and changes to sound reinforcement systems are known in the art, and any suitable technique can be used, whether it is already known or will be developed in the future. With reference to Fig. 4, a computer program product 480 is provided. The computer program product 480 comprises computer-readable media 482 with storage media 484 and program instructions 486 (i.e., program code) stored thereon. The program instructions 486 are configured to be loaded into the storage unit 414 of the host processing system 410 via the I / O unit 416, for example, by one of the user interface units 460 or other units 450 connected to the network 440. In exemplary implementations, the program instructions 486 are configured to perform steps of one or more of the methods disclosed herein as described above, for example, the steps of the method shown in Fig. 2 or Fig. 3. Although the present disclosure has been described and illustrated using exemplary implementations, it is evident to the person skilled in the art that the present disclosure is suitable for many different variants and modifications, which are not specifically shown here. The present invention may be a system, a method, and / or a computer program product. The computer program product may comprise a computer-readable storage medium (or media) on which computer-readable program instructions are stored to induce a processor to execute aspects of the present invention. A computer-readable storage medium can be a physical unit capable of retaining and storing instructions for use by a unit to execute instructions. For example, a computer-readable storage medium can be an electronic storage unit, a magnetic storage unit, an optical storage unit, an electromagnetic storage unit, a semiconductor storage unit, or any suitable combination thereof, without limitation. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer disk, a hard disk, random-access memory (RAM), read-only memory (ROM), and erasable programmable read-only memory (EPROM).Flash memory), static random-access memory (SRAM), portable compact storage disk-read-only memory (CD-ROM), a DVD (digital versatile disc), a USB flash drive, a floppy disk, a mechanically coded unit such as punched cards or raised structures in a groove on which instructions are stored, and any suitable combination thereof. For the purposes of this usage, a computer-readable storage medium shall not be understood as volatile signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses traveling through an optical fiber cable), or electrical signals transmitted by a wire. The computer-readable program instructions described here can be downloaded from a computer-readable storage medium to individual data processing units or, via a network such as the internet, a local area network (LAN), a wide area network (WAN), and / or a wireless network, to an external computer or storage device. The network may include copper cables, fiber optic cables, wireless transmission, routers, firewalls, switching units, gateway computers, and / or edge servers. A network adapter card or network interface in each data processing unit receives computer-readable program instructions from the network and forwards them for storage on a computer-readable storage medium within the respective data processing unit.Computer-readable program instructions for executing work steps of the present invention may be assembly instructions, ISA (Instruction Set Architecture) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as the programming language "C" or similar programming languages.The computer-readable program instructions can be executed entirely on the user's computer, partially on the user's computer as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the latter case, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, via the internet using an internet service provider).In some embodiments, electronic circuits, including, for example, programmable logic circuits, field programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), can execute computer-readable program instructions by using state information from the computer-readable program instructions to personalize the electronic circuits to implement aspects of the present invention. Aspects of the present invention are described here with reference to flowcharts and / or block diagrams or diagrams of methods, devices (systems), and computer program products according to embodiments of the invention. It is pointed out that each block of the flowcharts and / or block diagrams or diagrams, as well as combinations of blocks in the flowcharts and / or block diagrams or diagrams, can be executed by means of computer-readable program instructions. These computer-readable program instructions can be provided to a processor of a general-purpose computer, a specialized computer, or another programmable data processing device to create a machine such that the instructions executed by the processor of the computer or other programmable data processing device generate a means of implementing the functions / steps specified in the block(s) of the flowcharts and / or block diagrams or charts.These computer-readable program instructions may also be stored on a computer-readable storage medium capable of controlling a computer, programmable data processing device, and / or other units to function in a particular manner, such that the computer-readable storage medium on which instructions are stored has a manufacturing item, including instructions that implement aspects of the function / step specified in the block(s) of the flowchart and / or block diagrams or charts. The computer-readable program instructions can also be loaded onto a computer, other programmable data processing device, or other unit to cause the execution of a series of process steps on the computer or other programmable device or other unit in order to generate a process executed on a computer, such that the instructions executed on the computer, other programmable device, or other unit implement the functions / steps specified in the block(s) of the flowcharts and / or block diagrams or charts. The flowcharts and block diagrams or charts in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this context, each block in the flowcharts or block diagrams or charts can represent a module, a segment, or a part of instructions that includes one or more executable instructions for performing the specific logical function(s). In some alternative implementations, the functions specified in the block may occur in a different order than shown in the figures. For example, two blocks shown consecutively may in reality be executed essentially simultaneously, or the blocks may sometimes be executed in reverse order depending on the corresponding functionality.It should also be noted that each block of the block diagrams or charts and / or flowcharts, as well as combinations of blocks in the block diagrams or charts and / or flowcharts, can be implemented by special hardware-based systems that perform the specified functions or steps, or execute combinations of special hardware and computer instructions. The descriptions of the various exemplary implementations of the present disclosure are presented for illustrative purposes only and are not intended to be exhaustive or limited to the disclosed implementations. It is obvious to those skilled in the art that many modifications and adaptations are possible without deviating from the scope and inventive concept of the described implementations. The terminology used here has been chosen to best explain the basic ideas of the exemplary implementations, their practical application, or technical improvements over technologies already on the market, or to enable those skilled in the art to understand the implementations described herein.

Claims

A method (200) implemented by a computer (410), comprising: performing (220) speech recognition from input audio data (420A) comprising speech input into a sound reinforcement system; performing (240) speech recognition from at least one instance of output audio data (420B) comprising speech reproduced by one or more loudspeakers of the sound reinforcement system; determining (260) a difference between a speech recognition result from the input audio data (420A) and a speech recognition result from the at least one instance of the output audio data; determining (270) that the quality of the reproduced speech is unsatisfactory if the difference is greater than or equal to a threshold value;and performing (290) one or more parameter adjustments of the sound reinforcement system to improve the quality of the speech reproduced, in response to a finding that the quality of the speech reproduced is unsatisfactory, wherein parameters of the one or more parameter adjustments are selected from a group which necessarily includes equalization of the audio channels for each frequency band of a component of the sound reinforcement system and optionally additionally includes audio amplification. A method (200) implemented by a computer (410) according to claim 1, wherein the difference has a quantitative value calculated from the result of speech recognition based on the input audio data (420A) for a sample of the input language and the result of speech recognition based on the output audio data (420B) for a sample of the reproduced language corresponding to the sample of the input language. A method (200) implemented by a computer (410) according to claim 1, wherein determining the difference between the result of speech recognition based on the input audio data (420A) and the result of speech recognition based on the at least one instance of the output audio data (420B) comprises: comparing text of a transcript of speech recognition based on the input audio data (420A) with text of a transcript of speech recognition based on the at least one instance of the output audio data (420B); and determining a quantitative value for the difference, which is selected from a group consisting of: number of characters that differ, number of words that differ, and percentage of characters or words that differ. A method (200) implemented by a computer (410) according to claim 1, wherein determining the difference between the result of speech recognition based on the input audio data (420A) and the result of speech recognition based on the at least one instance of the output audio data (420B) comprises: comparing a first reliability level determined by the speech recognition based on the input audio data (420A) with a second reliability level determined by the speech recognition based on the at least one instance of the output audio data (420B), wherein the first reliability level comprises a value of a reliability measurement indicating the reliability of the speech recognition based on the input audio data (420A), and the second reliability level comprises a value of a reliability measurement indicating the reliability of the speech recognition based on the at least one instance of the output audio data (420B);and determining a difference between the first reliability grade and the second reliability grade. A method (200) implemented by a computer (410) according to claim 1, wherein performing one or more parameter adjustments of the sound reinforcement system comprises: performing a first parameter adjustment, which includes adjusting a parameter of the sound reinforcement system by a defined increment or to a target value; and, in response to a finding that the quality of the reproduced speech is still not satisfactory, performing a further parameter adjustment until a predefined condition is met. A method (200) implemented by a computer (410) according to claim 5, wherein the predefined condition is selected from a group consisting of: the quality of the reproduced speech is satisfactory, the quality of the reproduced speech is maximized, a predefined set of parameter adjustments has been performed, the predefined set of parameter adjustments has been performed according to an intelligent parameter adjustment scheme, the predefined set of parameter adjustments has been adjusted according to an optimization model, and a timer is running. A method (200) implemented by a computer (410) according to claim 1, further comprising: determining the effect of a parameter adjustment on the quality of the reproduced speech based on the difference between the result of speech recognition using the input audio data (420A) and the result of speech recognition using the at least one instance of the output audio data (420B), wherein the input audio data (420A) and the output audio data (420B) comprise input speech and the reproduced speech after the parameter adjustment, and storing ratio information comprising a parameter and an increment corresponding to the parameter adjustment and the effect on the quality of the reproduced speech, in order to use this as feedback for an intelligent scheme for selecting the parameter adjustment or a machine learning model for optimizing the quality of the reproduced speech. A method (200) implemented by a computer (410) according to claim 1, further comprising: determining whether the quality of speech input into the sound system is acceptable, and in response to a finding that the quality of speech input into the sound system is not acceptable, sending a message to a user to make changes with respect to the speech input into the sound system. Device (400) comprising: a processor (412) and a data storage device (414), wherein the processor (412) is configured to: perform speech recognition based on input audio data (420A) comprising speech input into a sound reinforcement system; perform speech recognition based on at least one instance of output audio data (420B) comprising speech reproduced by one or more loudspeakers of the sound reinforcement system; determine (260) a difference between a result of speech recognition based on the input audio data (420A) and a result of speech recognition based on the at least one instance of the output audio data (420B); determine (270) that the quality of the reproduced speech is unsatisfactory if the difference is greater than or equal to a threshold value;and performs one or more parameter adjustments of the sound reinforcement system (290) to improve the quality of the reproduced speech in response to a finding that the quality of the reproduced speech is unsatisfactory, wherein parameters of the one or more parameter adjustments are selected from a group which necessarily includes equalization of the audio channels for each frequency band of a component of the sound reinforcement system and optionally additional audio amplification. Device (400) according to claim 9, wherein the difference has a quantitative value calculated from the result of speech recognition based on the input audio data (420A) for a sample of the input language and the result of speech recognition based on the output audio data (420B) for a sample of the reproduced language corresponding to the sample of the input language. Device (400) according to claim 9, wherein the processor (414) is configured to determine the difference between the result of speech recognition based on the input audio data (420A) and the result of speech recognition based on the at least one instance of the output audio data (420B) by: comparing text of a transcript of speech recognition based on the input audio data (420A) with text of a transcript of speech recognition based on the at least one instance of the output audio data (420B); and determining a quantitative value for the difference, which is selected from a group consisting of: number of characters that differ, number of words that differ, and percentage of characters or words that differ. Device (400) according to claim 11, wherein the processor (414) is configured to perform one or more parameter adjustments of the sound reinforcement system by: performing a first parameter adjustment, which includes adjusting a parameter of the sound reinforcement system with a defined increment or to a target value; and, in response to a finding that the quality of the reproduced speech is still not satisfactory, performing a further parameter adjustment until a predefined condition is met. Device (400) according to claim 12, wherein the predefined condition is selected from a group comprising: the quality of the reproduced speech is satisfactory, the quality of the reproduced speech is maximized, a predefined set of parameter adjustments has been performed, the predefined set of parameter adjustments has been performed according to an intelligent parameter adjustment scheme, the predefined set of parameter adjustments has been adjusted according to an optimization model, and a timer is running. Device (400) according to claim 9, wherein the processor (414) is configured to determine the difference between the result of speech recognition based on the input audio data (420A) and the result of speech recognition based on the at least one instance of the output audio data (420B) by: comparing a first reliability level determined by the speech recognition based on the input audio data (420A) with a second reliability level determined by the speech recognition based on the at least one instance of the output audio data (420B), wherein the first reliability level comprises a value of a reliability measurement indicating the reliability of the speech recognition based on the input audio data (420A), and the second reliability level comprises a value of a reliability measurement indicating the reliability of the speech recognition based on the at least one instance of the output audio data (420B);and determined a difference between the first reliability grade and the second reliability grade. Device (400) according to claim 9, wherein the processor (412) is further configured to: determine whether the quality of a speech input into the sound system is acceptable, and in response to a finding that the quality of the speech input into the sound system is not acceptable, send a message to a user to make changes with respect to the speech input into the sound system. A computer program product (480) comprising a non-volatile, computer-readable storage medium (484) containing program instructions (486), wherein the program instructions (486) are executable by a processor (412) to cause the processor (412) to: perform speech recognition based on input audio data (420A) comprising speech input into a sound reinforcement system; perform speech recognition based on at least one instance of output audio data (420B) comprising speech reproduced by one or more loudspeakers of the sound reinforcement system (240); determine a difference between a result of speech recognition based on the input audio data (420A) and a result of speech recognition based on the at least one instance of the output audio data (420B) (260);to determine (270) that the quality of the reproduced speech is unsatisfactory when the difference is greater than or equal to a threshold value; and to perform one or more parameter adjustments of the sound reinforcement system (290) to improve the quality of the reproduced speech in response to a finding that the quality of the reproduced speech is unsatisfactory, wherein parameters of one or more parameter adjustments are selected from a group which necessarily includes equalization of the audio channels for each frequency band of a component of the sound reinforcement system and optionally additional audio amplification.

Citation Information

Patent Citations

  • Method for providing automatic feed back about quality of voice signal to speaker, involves providing feed back about quality of voice signal to speaker during exceeding or lowering of preset confidence threshold per confidence level

    DE102008030086A1

  • Sound verification

    US20160293180A1