Speech processing device, speech input / output system, and speech processing method
The audio processing device enhances speech clarity for hard of hearing individuals by generating and phase-aligning audio signals, addressing communication challenges in call centers and similar settings.
Patent Information
- Application Number
- PCT/JP2024/042990
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-28
- Filing Date
- 2024-12-05
- Publication Date
- 2025-12-04
AI Technical Summary
Existing voice processing technologies struggle to provide sufficient clarity for hard of hearing individuals, leading to communication difficulties in call centers and similar settings.
An audio processing device that generates two audio signals with the same spectrum, filters out specific frequency components, and applies a phase correction signal to align the phases of these signals, enhancing clarity through a synthesis process.
The device significantly improves speech clarity by correcting phase delays and enhancing the signal-to-noise ratio, enabling clearer communication for hard of hearing users.
Smart Images

Figure JP2024042990_04122025_PF_FP_ABST
Abstract
Description
Audio processing device, audio input / output system, and audio processing method
[0001] The present invention relates to an audio processing device, an audio input / output system, and an audio processing method.
[0002] Many companies have established contact points known as call centers or support centers to handle questions and requests from users. In the following, a call center will be used as an example of a contact point. A call center is equipped with one or more telephones. Each operator at the call center communicates with users using the telephones to understand their questions and requests and provide answers to them. Some users who call call centers are hard of hearing. Hard of hearing users may not be able to fully understand what the operator is saying. As a result, communication between the operator and the user may not go smoothly. Therefore, various technologies have been proposed to solve these problems. For example, Patent Document 1 discloses a voice processing device that makes it easier to clearly hear voices coming from speakers in audio equipment, broadcasting facilities, etc.
[0003] Japanese Patent Application Publication No. 2016-110050
[0004] However, in recent years, there has been a demand for a voice clarity technology that surpasses the conventional voice processing technology described in Patent Document 1.
[0005] One aspect of the present invention has been made in consideration of the above-mentioned problems, and an object of the present invention is to provide a voice processing technique that can make voices clearer than ever before.
[0006] An audio processing device according to one aspect of the present invention comprises an audio signal generation unit that generates a first audio signal and a second audio signal having equal spectra from an audio signal representing audio; a filter that removes at least a first frequency component below a predetermined first frequency from the first audio signal; a synthesis unit that adds and synthesizes the first audio signal from which the first frequency component has been removed and the second audio signal to generate a synthesized audio signal; and a correction signal generation unit that generates a phase correction signal based on the second audio signal, wherein the filter has an operational amplifier that amplifies the first audio signal, and the operational amplifier is configured to receive the first audio signal at its inverting input terminal, and to output the first audio signal whose phase has been corrected to approach the phase of the second audio signal when the phase correction signal is input at its non-inverting input terminal.
[0007] Another aspect of the present invention provides an audio processing method including: an audio signal generation step in which an audio signal generator generates, from an audio signal representing audio, a first audio signal and a second audio signal having identical spectra; a filtering step in which a filter removes at least a first frequency component below a predetermined first frequency from the first audio signal; an audio synthesis step in which an audio synthesizer adds and synthesizes the first audio signal from which the first frequency component has been removed and the second audio signal to generate a synthesized audio signal; and a correction signal generation step in which a correction signal generator generates a phase correction signal based on the second audio signal, wherein the filter has an operational amplifier that amplifies the first audio signal, and the operational amplifier is configured to receive the first audio signal at its inverting input terminal; and in the filtering step, the phase correction signal is input to the non-inverting input terminal of the operational amplifier, causing the operational amplifier to output the first audio signal whose phase has been corrected so as to approach the phase of the second audio signal.
[0008] According to one aspect of the present invention, it is possible to make speech clearer than ever before.
[0009] 1 is a block diagram illustrating an example of a configuration of an audio processing device according to an embodiment of an aspect of the present invention. FIG. 2 is a diagram illustrating formant components contained in an audio signal. FIG. 3 is a graph illustrating adjustment characteristics of an input buffer unit included in the audio processing device of FIG. 1. FIG. 4 is a circuit diagram illustrating an example of a specific configuration of a filter and a correction signal generation unit included in the audio processing device of FIG. 1. FIG. 5 is a graph illustrating a phase correction characteristic of a filter included in the audio processing device of FIG. 1. FIG. 6 is a graph illustrating a case where a first audio signal and a second audio signal are synthesized without phase correction. FIG. 7 is a graph illustrating a frequency characteristic of a synthesized audio signal when phase correction is not performed. FIG. 8 is a graph illustrating an example of adjustment of an audio signal by an adjustment unit included in the audio processing device of FIG. 1. FIG. 9 is a graph illustrating changes over time in phase and output of an audio signal, a first audio signal before and after correction, and a synthesized audio signal. FIG. 10 is a graph illustrating overall output characteristics of a synthesis unit. FIG. 11 is a schematic diagram illustrating an example of a specific configuration of an auxiliary device in which the audio processing device of FIG. 1 is incorporated. FIG. 12 is a schematic diagram illustrating an example of a specific configuration of an audio input / output system in which the audio processing device of FIG. 1 is incorporated. FIG. 13 is a schematic diagram illustrating another example of a specific configuration of an audio input / output system. FIG. 14 is a schematic diagram illustrating another example of a specific configuration of an audio input / output system. It is a schematic diagram showing another example of a specific configuration of the voice input / output system.It is a schematic diagram showing another example of a specific configuration of the voice input / output system.It is a flowchart showing an example of the flow of the voice processing method according to an embodiment of another aspect of the present invention.
[0010] <Audio Processing Device> Hereinafter, an audio processing device according to an embodiment of the present invention will be described with reference to the drawings.
[0011] 1, the audio processing device 1 includes an audio signal generating unit 11, a correction signal generating unit 12, a filter 13, and a synthesizing unit 14. The audio processing device 1 according to this embodiment further includes a first port 15, an input buffering unit 16, an adjusting unit 17, and a second port 18. The audio processing device 1 may further include a power supply unit, a detecting unit, and a control unit, which are not shown.
[0012] (First Port) The first port 15 outputs an audio signal x representing audio supplied from a device connected to the first port 15 to the input buffer unit 16. As shown in Fig. 2, the audio signal x includes a first-order formant component h1, a second-order formant component h2, ..., and an nth-order formant component hn. Note that n depends on the user on the transmitting side and, although there may be some individual differences, is a positive integer of at least 4 or more. Note that the first-order formant component h1, the second-order formant component h2, ..., and the nth-order formant component hn are described in Patent Document 1, and therefore will not be described in this embodiment.
[0013] (Input Buffer Unit) As shown in FIG. 1 , the input buffer unit 16 is provided between the first port 15 and the audio signal generation unit 11. The input buffer unit 16 can be configured with a microphone amplifier, a preamplifier that reduces output impedance, or the like, depending on the usage situation. The input buffer unit 16 according to this embodiment adjusts the audio signal x by narrowing the upper and lower frequency limits and increasing the sound pressure, as shown in FIG. 3 , for example. Then, as shown in FIG. 1 , the input buffer unit 16 outputs the audio signal x that has been adjusted as necessary to the audio signal generation unit 11. Note that if the audio signal supplied from the first port does not require adjustment (if the device connected to the first port is configured to output an audio signal that does not require adjustment), the audio processing device 1 does not need to include the input buffer unit 16.
[0014] (Audio Signal Generator) The audio signal generator 11 generates a first audio signal x1 and a second audio signal x2 having the same spectrum from the audio signal x supplied from the input buffer 16. "Having the same spectrum" means that the first audio signal x1 and the second audio signal x2 each contain a first-order formant component h1, a second-order formant component h2, ..., an n-order formant component hn. The audio signal generator 11 according to this embodiment also generates a third audio signal x3 having the same spectrum as the second audio signal x2. The audio signal generator 11 then supplies the first audio signal x1 to the high-pass filter 131, the second audio signal x2 to the synthesis unit 14, and the third audio signal to the correction signal generator 12. The audio signal generator 11 according to this embodiment generates the first audio signal x1, the second audio signal x2, and the third audio signal x3 by branching the audio signal. In this embodiment, the audio signal generation unit 11 is configured so that the intensity ratio between the first audio signal x1, the second audio signal x2, and the third audio signal x3 is 1:1:1, i.e., so that the distribution ratio is 1:1:1. However, the distribution ratio of the audio signals is not limited to 1:1:1 and can be set appropriately. The audio signal generation unit 11 may also duplicate the audio signals supplied from the input buffer unit 16, and use the original audio signal as the first audio signal x1, and the duplicated audio signals as the second audio signal x2 and the third audio signal x3.
[0015] (Correction Signal Generator) The correction signal generator 12 generates a phase correction signal p based on the second audio signal x2 supplied from the audio signal generator 11. As described above, the audio signal generator 11 according to this embodiment generates the third audio signal x3 having the same phase as the second audio signal x2. Therefore, the correction signal generator 12 according to this embodiment generates the phase correction signal p based on the third audio signal x3. As a result, the correction signal generator 12 generates the phase correction signal p based on a dedicated third audio signal. Therefore, it is possible to prevent the second audio signal from being out of phase by using the second audio signal to generate the phase correction signal p. The correction signal generator 12 then outputs the generated phase correction signal p to the operational amplifier 132a of the filter 13.
[0016] As shown in FIG. 4 , the correction signal generator 12 according to this embodiment includes a capacitor 121 and a resistor 122. One end of the capacitor 121 is connected to the audio signal generator 11, and the other end is connected to the non-inverting input terminal (+) of the operational amplifier 132a. One end of the resistor 122 is connected to the other end of the capacitor 121, and a bias voltage is applied to the other end. The correction signal generator 12 configured in this manner outputs a phase correction signal p containing only frequency components equal to or greater than ½πCR (C: capacitance, R: resistance) using a high-pass filter function internally formed by the capacitor 121 and the resistor 122. Hereinafter, the lower limit of the frequency components contained in the phase correction signal p will be referred to as the phase correction start frequency fs. The correction signal generator 12 can adjust the phase correction start frequency fs by changing at least one of the capacitor 121 and the resistor 122. As will be described in detail later, the degree to which the operational amplifier 132a of the filter 13 corrects the phase varies within a range of 0 to −180° depending on the phase correction start frequency fs of the phase correction signal p input to the non-inverting input terminal (+), as shown in Fig. 5. For this reason, the phase correction start frequency fs is adjusted to an optimum value that minimizes the phase difference between the first audio signal x1 and the second audio signal x2 after correction.
[0017] Furthermore, the impedance Z (internal resistance) of the capacitor 121 changes depending on the phase correction start frequency fs. Specifically, the impedance Z of the capacitor 121 is inversely proportional to the phase correction start frequency fs, and its value is expressed as 1 / ωC (ω=2πfs). That is, the impedance Z of the capacitor 121 decreases as the phase correction start frequency fs increases. Therefore, the correction signal generator 12 outputs a phase correction signal p with a higher voltage as the phase correction start frequency fs increases, and outputs a phase correction signal p with a lower voltage as the phase correction start frequency fs decreases. In this way, by using the capacitor 121, the correction signal generator 12 can be configured with a simple configuration that is compatible with audio signals of various frequencies.
[0018] (Filter) The filter 13 removes at least a first frequency component equal to or lower than a predetermined first frequency from the first audio signal. The first frequency is a frequency between the upper limit frequency of the first formant component h1 and the lower limit frequency of the second formant component h2. The filter 13 according to this embodiment is a band-pass filter that also removes a second frequency component equal to or higher than a second frequency that is higher than the first frequency. The filter 13 according to this embodiment is composed of a high-pass filter 131 and a low-pass filter 132.
[0019] The high-pass filter 131 removes a first frequency component from the first audio signal x1 supplied from the audio signal generation unit 11. In the high-pass filter 131 according to this embodiment, the first frequency is set to 400 Hz. The high-pass filter 131 then outputs the first audio signal x1' from which the first frequency component has been removed to the low-pass filter 132. The first audio signal x1' is an audio signal that does not include a first-order formant component h1, but does include a second-order formant component h2, a third-order formant component h3, ..., an n-th order formant component hn.
[0020] The high-pass filter 131 according to this embodiment has an operational amplifier (not shown) that amplifies the first audio signal x1. Therefore, the high-pass filter 131 according to this embodiment outputs the first audio signal x1 from which the first frequency component has been removed and which has been amplified. The operational amplifier amplifies the first audio signal x1 from which the first frequency component has been removed, thereby outputting the first audio signal x1" in which the intensities of the second-order formant component h2, the third-order formant component h3, ..., the n-order formant component hn have been increased. Note that the high-pass filter 131 may be configured so that at least one of the gain of the operational amplifier and the boundary frequency at which components are removed or not can be set arbitrarily depending on the intended use.
[0021] The low-pass filter 132 removes a second frequency component from the first audio signal x1, from which the first frequency component has been removed, supplied from the high-pass filter 131. The second frequency component is a frequency component equal to or greater than the second frequency. The second frequency (cutoff frequency) is a frequency between the upper limit frequency of the fifth-order formant component h5 and the lower limit frequency of the sixth-order formant component h6, or a frequency between the upper limit frequency of the sixth-order formant component h6 and the lower limit frequency of the seventh-order formant component h7. In the low-pass filter 132 according to this embodiment, the second frequency is set to 5 kHz or 7 kHz. Low-pass filter 132 then outputs first speech signal x1" from which the fifth, sixth, or higher frequency components have been removed, i.e., a speech signal containing only second- to fifth-order formant components h2 to h5, or only second- to sixth-order formant components h2 to h6. Higher-order formant components become noise that does not contribute to language understanding. For this reason, by low-pass filter 132 removing the second frequency component, it is possible to improve the signal-to-noise ratio in synthesized speech signal x', which will be described later.
[0022] However, when frequency components are removed from the first speech signal, a phase delay inevitably occurs. If a first speech signal x1, whose phase is delayed by time T, is combined with a second speech signal x2 without correction, the voltage V1 of the first speech signal x1 becomes −V1 due to the phase delay, as shown in FIG. 6 . Therefore, the voltage V of the combined speech signal becomes V2 + (−V1) because the first speech signal x1 and the second speech signal x2 cancel each other out. As a result, as shown in FIG. 7 , a dead zone (deep valley) called a dip occurs in a certain frequency range, causing extremely low sound pressures for the first-order formant component h1 and the second-order formant component h2, which are particularly important for speech comprehension, and other problems that significantly hinder improvements in clarity.
[0023] For this reason, the filter 13 has a function of correcting phase delay. Specifically, as shown in FIG. 4, an operational amplifier 132a is provided in the low-pass filter 132 together with other elements 132b and 132c. The operational amplifier 132a amplifies the first audio signal x1. The operational amplifier 132a is configured so that the first audio signal is input to its inverting input terminal (-). The operational amplifier 132a receives a phase correction signal p at its non-inverting input terminal (+), and outputs a first audio signal x1' whose phase has been corrected to approach the phase of the second audio signal x2. Specifically, the operational amplifier 132a compares the voltage of the first audio signal input to its inverting input terminal (-) and the voltage of the phase correction signal p input to its non-inverting input terminal (+). If the voltage of the phase correction signal p is higher, the operational amplifier 132a outputs a first audio signal x1' whose phase has been advanced in accordance with the voltage difference. On the other hand, if the voltage of the phase correction signal p is lower, the operational amplifier 132a outputs a first audio signal x1' whose phase has been delayed in accordance with the voltage difference. As described above, the voltage of the phase correction signal p varies in accordance with the phase correction start frequency fs determined by the capacitor 121 and resistor 122 of the correction signal generation unit 12. Therefore, the degree to which the operational amplifier 132a advances or delays the phase varies within a range of 0 to −180° depending on the phase correction start frequency fs, as shown in FIG. 5, for example. As described above, the phase correction start frequency fs is adjusted in advance. Therefore, a phase correction signal p having a voltage corresponding to the adjusted phase correction start frequency fs is input to the non-inverting input terminal (+) of the operational amplifier 132a. The operational amplifier 132a then outputs the first audio signal x1' corrected so that the difference in phase with the second audio signal x2 is minimized.
[0024] The filter 13 may be configured such that the low-pass filter 132 first removes the second frequency component from the first audio signal x1. Then, the high-pass filter 131 removes the first frequency component from the first audio signal x1 from which the second frequency component has been removed. In this case, the phase correction signal p is input to the non-inverting input terminal (+) of the operational amplifier included in the high-pass filter 131.
[0025] (Adjustment Unit) The adjustment unit 17 adjusts the sound pressure of the phase-corrected first audio signal x1'. Then, the adjustment unit 17 supplies the adjusted first audio signal x1' to the synthesis unit 14. The adjustment unit 17 adjusts the first audio signal based on the user's setting operation according to the usage environment and the degree of hearing loss. Therefore, the adjustment unit 17 can adjust only the first to fifth formant components h1 to h5, which are important for language comprehension, to a level that exceeds the masking line MK', which is an audible level. In addition, the adjustment unit 17 can freely select the output level curve of the synthesized audio signal x' to be "minimum," "maximum," or "intermediate," as shown in FIG. 8.
[0026] (Synthesis Unit) The synthesis unit 14 generates a synthetic audio signal x' by additively synthesizing the first audio signal x1', from which the first frequency component and the second frequency component have been removed and whose phase has been corrected, with the second audio signal x2 supplied from the audio signal generation unit 11. As described above, the audio processing device 1 according to this embodiment includes the adjustment unit 17. Therefore, the synthesis unit 14 according to this embodiment additively synthesizes the first audio signal x1, from which the sound pressure has been adjusted, with the second audio signal x2. The synthesis unit 14 then supplies the generated synthetic audio signal x' to the second port 18. As shown in FIG. 9 , the phase of the corrected first audio signal x1' in the synthesis unit 14 leads the phase of the uncorrected first audio signal x1 by T'. Therefore, the corrected first audio signal x1' is in phase with the second audio signal x2. As a result, the frequency characteristics of the synthetic audio signal x' output by the synthesis unit 14 are as shown in FIG. 10 .
[0027] (Second Port) The second port 18 outputs the synthesized voice signal x′ supplied from the synthesis unit 14 to a device connected to the second port 18 .
[0028] (Power Supply Unit) The power supply unit supplies power to each unit requiring power in the audio processing device 1. The form of the power supply unit is not limited, and may be an AC / DC converter such as an AC adapter, or a battery.
[0029] (Detection Unit) The detection unit is configured to detect whether or not power is being supplied from the power supply unit to the audio processing unit. If the power supply unit is functioning normally, the detection unit supplies detection information indicating "Yes" to the control unit. On the other hand, if the power supply unit is not functioning normally, the detection unit supplies detection information indicating "No" to the control unit. If the power supply unit is an AC / DC converter, an example of a state in which it is not functioning normally is a power outage. Furthermore, if the power supply unit is a battery, an example of a state in which it is not functioning normally is a dead battery.
[0030] (Selection Unit) The selection unit selects whether or not to apply audio processing to the audio signal x based on control by the control unit. The selection unit is provided, for example, between the low-pass filter 132 and the synthesis unit 14, and is configured by a switch that switches the circuit between ON and OFF. When the selection unit is ON, the first audio signal x1 that has passed through the high-pass filter 131 and the low-pass filter 132 is supplied to the synthesis unit 14, where it is synthesized with the second audio signal x2. On the other hand, when the selection unit is OFF, the first audio signal x1 is not supplied to the synthesis unit 14, and the second audio signal x2 is output from the synthesis unit 14. Note that the selection unit may be provided between the audio signal generation unit 11 and the high-pass filter 131, or between the high-pass filter 131 and the low-pass filter 132.
[0031] (Control unit) The control unit controls the selection unit to perform or not perform audio processing on the audio signal x according to the detection result of the detection unit. If the detection result of the detection unit is "Yes", the control unit controls the selection unit to perform audio processing on the audio signal x (turn ON). On the other hand, if the detection result of the detection unit is "No", the control unit controls the selection unit not to perform audio processing on the audio signal x (turn OFF).
[0032] [Modification of Audio Processing Device] In the audio processing device 1 according to the above embodiment, the high-pass filter 131, the low-pass filter 132, the correction signal generation unit 12, and the synthesis unit 14 are configured as analog circuits. However, at least one of these may be configured as a digital circuit.
[0033] Furthermore, the audio processing device 1 according to the above embodiment includes the high-pass filter 131 and the low-pass filter 132. However, instead of including these filters, the audio processing device 1 may include a band-pass filter that simultaneously removes the first frequency component and the second frequency component.
[0034] The audio processing device 1 may further include a polarity switch and an amplifier interposed between the first port 15 and the audio signal generating unit 11 .
[0035] The voice processing device 1 may further include a selector, an amplifier, a circuit separator, a transmission output adjuster, and a polarity switcher interposed between the synthesis unit 14 and the second port 18 .
[0036] The audio processing device 1 may also include an attenuator between the synthesis unit 14 and the second port 18. The attenuator arbitrarily adjusts the output to match the input sound source signal. By incorporating such an attenuator into the signal line of a microphone in an existing device that transmits audio via wired or wireless communication, or into a television, radio, or other entertainment playback device, it becomes possible to provide a system that accommodates hearing loss symptoms due to hearing impairment.
[0037] [Operation and Effect of Audio Processing Device] The audio processing device 1 described above corrects the phase delay that occurs when the low-pass filter 132 removes the first and second frequency components based on the phase correction signal p. Therefore, the audio processing device 1 can make audio clearer than ever before.
[0038] <Application Examples of Audio Processing Device> Next, a description will be given of various devices incorporating the audio processing device 1. The audio processing device 1 can be incorporated into an auxiliary device 2, audio input / output systems 3 to 6, and the like, which will be described below.
[0039] [Configuration of Auxiliary Device] As shown in FIG. 11, the auxiliary device 2 includes the audio processing device 1, a housing 21, an input terminal 22, and an output terminal 23.
[0040] The housing 21 is formed in a box shape capable of housing the voice processing device 1. The housing 21 is easily portable.
[0041] The input terminal 22 is provided on the surface of the housing 21. The input terminal 22 is connected to a device that outputs an audio signal, and receives an audio signal from the device. The input terminal 22 is connected to the first port 15 of the audio processing device 1. In other words, the input terminal 22 supplies the audio signal input from the device that outputs the audio signal to the first port 15.
[0042] The output terminal 23 is provided on the surface of the housing 21. The output terminal 23 is connected to a device to which the synthetic speech signal x' is input, and outputs a speech signal to the device. The output terminal 23 is connected to the second port 18 of the speech processing device 1. In other words, the output terminal 23 outputs the synthetic speech signal x' supplied from the speech processing device 1 to the device to which the synthetic speech signal x' is input.
[0043] [Effects and Functions of the Auxiliary Device] The auxiliary device 2 described above is not only easy to install and remove, but also easy to carry and store, making it easily applicable to broadcasting equipment at outdoor event venues and other audio output devices.
[0044] 12 , the audio input / output system 3 includes, in addition to the audio processing device 1, an audio input unit 31 and an audio output unit 32. The audio input / output system 3 according to this embodiment further includes an amplifier unit 33.
[0045] (Audio Input Unit) The audio input unit 31 acquires the speaker's voice and generates an audio signal. The audio input unit 31 according to this embodiment is configured with a microphone. The audio input unit 31 is connected to the first port of the audio processing device 1. The audio input unit 31 generates an audio signal x based on the input voice and supplies it to the first port 15. The audio input unit 31 may be, for example, a device that plays back video, such as a television, as shown in FIG. 13 .
[0046] (Audio Output Unit) The audio output unit 32 outputs audio based on the synthetic audio signal x' generated by the audio processing device 1. The audio input unit 31 according to this embodiment is configured with a speaker, headphones, earphones, etc. The audio output unit 32 is connected to the second port 18 of the audio processing device 1. The audio output unit 32 then emits audio based on the synthetic audio signal supplied from the second port.
[0047] [Effects of the Audio Input / Output System (1)] The audio input / output system 3 described above can be configured with existing broadcasting equipment, audio transmission devices, etc., except for the audio processing device 1. Therefore, by incorporating the audio input / output system 3 into existing broadcasting equipment, audio transmission devices, etc., it is possible to easily clarify the audio emitted by these devices.
[0048] [Configuration of Audio Input / Output System (2)] As shown in FIG. 14 , the audio input / output system 4 includes an audio processing device 1, an audio input unit 31, an audio output unit 32, and an amplifier unit 33 similar to those of the audio input / output system 3, as well as a transmitter 41 and a receiver 42.
[0049] The transmitter 41 according to this embodiment is configured as a wireless communication module. The transmitter 41 is connected to the second port 18 of the voice processing device 1. The transmitter 41 converts the synthesized voice signal supplied from the second port into radio waves. The transmitter 41 then transmits the radio waves from the antenna 41 a.
[0050] The receiver 42 according to this embodiment is configured as a wireless communication module. The receiver 42 is connected to the amplifier 33. The receiver 42 receives radio waves from the transmitter 41 via an antenna 42a. The receiver 42 converts the received radio waves into a synthesized voice signal and supplies the synthesized voice signal to the amplifier 33.
[0051] 15, the voice input / output system 4 may have the voice input unit 31 and the voice output unit 32 configured as a telephone receiver, and the transmitter 41 and the receiver 42 configured as the main body of the telephone. In this case, the transmitter 41 transmits the synthesized voice signal to the receiver 42 via a fixed telephone line, the Internet line, etc.
[0052] [Effects of Voice Input / Output System (2)] According to the voice input / output system 4 described above, clear voice can be transmitted to the listener even if the speaker and listener are far apart.
[0053] [Configuration of Audio Input / Output System (3)] The audio input / output system 5 constitutes a stereo device. As shown in Fig. 16 , the audio input / output system 5 includes a pair of input buffer units 16, a pair of audio signal generators 11, a correction signal generator 12, a filter 13, an adjuster 17, and a pair of synthesizers 14, similar to those of the audio processing device 1. The audio input / output system 5 also includes a pair of audio output units 32, similar to those of the audio input / output system 3. The audio input / output system 5 also includes a pair of audio input terminals 51, a second synthesizer 52, a second audio signal generator 53, and a third audio signal generator 54.
[0054] The pair of audio input terminals 51 are each connected to the input buffer unit 16. A right audio signal x(R) and a left audio signal x(L) are respectively input to the pair of audio input terminals 51. The input right audio signal x(R) and left audio signal x(L) are each supplied to the input buffer unit 16.
[0055] The audio signal generating unit 11 in the audio input / output system 5 does not generate the third audio signal x3, but generates the first audio signals x1(R), x1(L) and the second audio signals x2(R), x2(L).
[0056] The second synthesis unit 52 additively synthesizes the right first audio signal x1(R) and the left first audio signal x1(L) generated by the pair of audio signal generation units 11. The synthesized first audio signal x1 is supplied to the second audio signal generation unit 53.
[0057] The second audio signal generating unit 53 generates, from the first audio signal x1, a third audio signal having the same spectrum as the first audio signal.
[0058] The third audio signal generation unit 54 generates a right first audio signal x1'(R) and a left first audio signal x1'(L) from the phase-corrected first audio signal x1' supplied from the adjustment unit 17. The first audio signals x1'(R), x1'(L) are supplied to a pair of synthesis units 14, respectively.
[0059] [Effects of Audio Input / Output System (3)] According to the audio input / output system 5 described above, stereo audio (right audio signal and left audio signal) can be made clearer than ever before.
[0060] [Configuration of Audio Input / Output System (4)] The audio input / output system 6 constitutes a so-called remote conference system. As shown in Fig. 17 , the audio input / output system 6 includes a pair of audio processing devices 1, a pair of audio input units 31, and a pair of audio output units 32 similar to those of the audio input / output system 3, as well as a pair of dedicated terminals 61 for the remote conference system. The audio input / output system 6 may also include a camera and a monitor (not shown).
[0061] The audio input unit 31 is connected to a first port of the audio processing device 1. The audio input unit 31 generates an audio signal x based on the input audio and supplies it to the first port 15. The second port 18 of the audio processing device 1 is connected to a dedicated terminal 61. The audio processing device 1 generates a synthesized audio signal from the audio signal x supplied from the audio input unit 31 and supplies it to the dedicated terminal 61.
[0062] The dedicated terminal 61 is connected to another dedicated terminal 61 via an internet line. One dedicated terminal 61 supplies the synthesized voice signal supplied from the voice processing device 1 to the other dedicated terminal. Furthermore, one dedicated terminal 61 supplies the synthesized voice signal supplied from the other dedicated terminal 61 to the voice output unit 32.
[0063] The voice output unit 32 is connected to the dedicated terminal 61. The voice output unit 32 then emits voice based on the synthesized voice signal supplied from the dedicated terminal 61.
[0064] [Effects of the Voice Input / Output System (4)] The voice input / output system 6 allows the voice uttered by a speaker at one (other) dedicated terminal to be clearly transmitted to a listener at the other (one) dedicated terminal. As a result, a smooth conversation (discussion) can be carried out between a person at one dedicated terminal and a person at the other dedicated terminal.
[0065] <Audio Processing Method> Next, an audio processing method according to an embodiment of another aspect of the present invention will be described.
[0066] As shown in FIG. 18, the audio processing method includes an audio signal generating step S1, a correction signal generating step S2, a filtering step S3, and an audio synthesis step S4.
[0067] (Audio Signal Generation Step) In the initial audio signal generation step S1, an audio signal generator generates a first audio signal x1 and a second audio signal x2 having the same spectrum from an audio signal representing audio. Note that in the audio signal generation step S1, the audio signal generator may also generate a third audio signal having the same phase as the second audio signal. Furthermore, the audio signal generator may be a component of the audio signal generation unit 11 of the audio processing device 1, a component of another device, or an independent device.
[0068] (Correction Signal Generation Step) After generating the first audio signal x1 and the second audio signal x2, the process proceeds to the correction signal generation step S2. In the correction signal generation step S2, the correction signal generator generates a phase correction signal p based on the second audio signal x2. Note that if a third audio signal is generated in the audio signal generation step S1, the correction signal generator may generate the phase correction signal based on the third audio signal in the correction signal generation step S2. The correction signal generator used in the correction signal generation step S2 may include a capacitor having one end connected to the audio signal generation unit and the other end connected to the non-inverting input terminal. The correction signal generator may be part of the correction signal generation unit 12 of the audio processing device 1, part of another device, or an independent device.
[0069] (Filtering Step) After generating the first audio signal x1 and the second audio signal x2, a filtering step S3 is also performed. In the filtering step S3, a filter removes a first frequency component below a predetermined first frequency from the first audio signal x1 (S31). Also, in the filtering step S3, a phase correction signal p is input to the non-inverting input terminal of the operational amplifier, causing the operational amplifier to output the first audio signal x1 whose phase has been corrected to approach the phase of the second audio signal x2 (S32). Note that in the filtering step S3, a band-pass filter that also removes a second frequency component above a second frequency higher than the first frequency may be used as the filter. Alternatively, in the filtering step S3, a filter may be used that includes a high-pass filter that removes the first frequency component from the first audio signal and a low-pass filter that removes the second frequency component from the first audio signal from which the first frequency component has been removed. In this case, it is preferable to use an operational amplifier that is provided in the low-pass filter. In addition, in the filtering step S3, a high-pass filter including an operational amplifier that amplifies the first audio signal may be used. The filter may be a component of the audio signal generating unit 11 of the audio processing device 1, a component of another device, or an independent device. In addition, in the filtering step S3, the removal of the second frequency component may be performed before the removal of the first frequency component. In addition, the removal of the first frequency component and the second frequency component may be performed simultaneously.
[0070] (Speech Synthesis Step) After removing the second frequency component from the first speech signal x1, the process proceeds to speech synthesis step S4. In speech synthesis step S4, a speech synthesizer additively synthesizes the phase-corrected first speech signal x1 and the second speech signal x2 to generate a synthesized speech signal x'. The speech synthesizer may be a component of the speech signal generation unit 11 of the speech processing device 1, a component of another device, or an independent device.
[0071] [Variation of the Audio Processing Method] The audio processing method may further include an adjustment step. The adjustment step is preferably performed between the filtering step and the audio synthesis step. The adjustment step adjusts the sound pressure of the phase-corrected first audio signal. When the adjustment step is included, in the audio synthesis step, the audio synthesizer additively synthesizes the first audio signal whose sound pressure has been adjusted and the second audio signal.
[0072] [Effects of the Audio Processing Method] In the audio processing method described above, in filtering step S3, the filter corrects the phase delay that occurs when the first and second frequency components are removed based on the phase correction signal p. Therefore, the audio processing method can make audio clearer than ever before.
[0073] <Additional Notes> The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in the description of the invention are also included in the technical scope of the present invention.
[0074] <Summary> An audio processing device according to aspect 1 of the present invention comprises an audio signal generation unit that generates a first audio signal and a second audio signal, the first audio signal and the second audio signal having the same spectrum, from an audio signal representing audio; a filter that removes at least a first frequency component that is equal to or less than a predetermined first frequency from the first audio signal; a synthesis unit that adds and synthesizes the first audio signal from which the first frequency component has been removed and the second audio signal to generate a synthesized audio signal; and a correction signal generation unit that generates a phase correction signal based on the second audio signal, wherein the filter has an operational amplifier that amplifies the first audio signal, and the operational amplifier is configured to receive the first audio signal at its inverting input terminal, and to output the first audio signal whose phase has been corrected to approach the phase of the second audio signal when the phase correction signal is input at its non-inverting input terminal.
[0075] According to the above-described voice processing device, voice can be made clearer than ever before.
[0076] A sound processing device according to a second aspect of the present invention may be configured in the above-mentioned first aspect such that the filter is a band-pass filter that also removes a second frequency component equal to or greater than a second frequency that is higher than the first frequency.
[0077] According to the above configuration, the signal-to-noise ratio can be improved.
[0078] An audio processing device according to aspect 3 of the present invention may be configured in the above-mentioned aspect 2 such that the filter is composed of a high-pass filter that removes the first frequency component from the first audio signal and a low-pass filter that removes the second frequency component from the first audio signal from which the first frequency component has been removed, and the operational amplifier is provided in the low-pass filter.
[0079] According to the above configuration, the signal is passed through the operational amplifier after all frequency components have been removed (after the phase is no longer delayed), so that the phase delay can be corrected more reliably.
[0080] An audio processing device according to Aspect 4 of the present invention may be configured in the above-described Aspect 1 such that the audio signal generation unit generates a third audio signal having an equal spectrum to the second audio signal, and the correction signal generation unit generates the phase correction signal based on the third audio signal.
[0081] According to the above configuration, the correction signal generator generates the correction signal based on the dedicated third audio signal, which prevents the second audio signal from being out of phase when the second audio signal is used to generate the correction signal.
[0082] An audio processing device according to aspect 5 of the present invention may be configured in the above-mentioned aspect 1 such that the correction signal generation unit has a capacitor having one end connected to the audio signal generation unit and the other end connected to the non-inverting input terminal.
[0083] With the above configuration, the higher the frequency of the signal passing through the filter, the greater the phase delay, while the capacitor increases the output as the frequency of the input signal increases, and the low-pass filter advances the phase of the output signal as the input to the non-inverting input terminal increases. Therefore, with a simple configuration, it is possible to configure a correction signal generator that can handle audio signals of various frequencies.
[0084] A voice processing device according to Aspect 6 of the present invention may be configured in the above-mentioned Aspect 1 such that the first frequency is a frequency between an upper limit frequency of a first-order formant component and a lower limit frequency of a second-order formant component.
[0085] According to the above configuration, the first formant components are omitted from the first speech signal, and only the second formant components, the third formant components, etc. are left, so that only the second formant components and the third formant components can be amplified.
[0086] An audio processing device according to aspect 7 of the present invention may be configured in the above-mentioned aspect 6 such that the second frequency is a frequency between the upper limit frequency of the fifth formant component and the lower limit frequency of the sixth formant component, or a frequency between the upper limit frequency of the sixth formant component and the lower limit frequency of the seventh formant component.
[0087] According to the above configuration, the signal-to-noise ratio can be further improved.
[0088] An audio processing device according to aspect 8 of the present invention may be configured in the above-mentioned aspect 3 such that the high-pass filter includes an operational amplifier that amplifies the first audio signal, and outputs the amplified first audio signal from which the first frequency component has been removed.
[0089] According to the above configuration, the clarity of the voice can be improved.
[0090] An audio processing device according to aspect 9 of the present invention may be configured in accordance with aspect 1 above, further comprising an adjustment unit that adjusts the sound pressure of the first audio signal whose phase has been corrected, and the synthesis unit additively synthesizes the first audio signal whose sound pressure has been adjusted and the second audio signal.
[0091] According to the above configuration, it is possible to output sound according to the usage environment and the degree of hearing loss.
[0092] The voice input / output system according to aspect 10 of the present invention may be configured as in any one of aspects 1 to 6 above, comprising a voice input unit that acquires the voice of a speaker and generates a voice signal, the voice processing device, and a voice output unit that outputs voice based on the synthesized voice signal generated by the voice processing device.
[0093] According to the above configuration, the sound can be made clearer than ever before.
[0094] An audio processing method according to aspect 11 of the present invention includes: an audio signal generation step in which an audio signal generator generates a first audio signal and a second audio signal having identical spectra from an audio signal representing audio; a filtering step in which a filter removes at least a first frequency component below a predetermined first frequency from the first audio signal; an audio synthesis step in which an audio synthesizer adds and synthesizes the first audio signal from which the first frequency component has been removed and the second audio signal to generate a synthesized audio signal; and a correction signal generation step in which a correction signal generator generates a phase correction signal based on the second audio signal, wherein the filter has an operational amplifier that amplifies the first audio signal, and the operational amplifier is configured to receive the first audio signal at its inverting input terminal; and in the filtering step, the phase correction signal is input to the non-inverting input terminal of the operational amplifier, causing the operational amplifier to output the first audio signal whose phase has been corrected so as to approach the phase of the second audio signal.
[0095] According to the above-described voice processing method, voice can be made clearer than ever before.
[0096] 1 Audio processing device 11 Audio signal generation unit 12 Correction signal generation unit 121 Capacitor 122 Resistor 13 Filter 131 High-pass filter 132 Low-pass filter 132a Operational amplifier 132b, 132c Element 14 Synthesis unit 15 First port 16 Input buffer unit 17 Adjustment unit 18 Second port 2 Auxiliary device 21 Housing 22 Input terminal 23 Output terminal 3, 4, 5, 6 Audio input / output system 31 Audio input unit 32 Audio output unit 33 Amplification unit 41 Transmitter 41a, 42a Antenna 42 Receiver 51 Audio input terminal 52 Second synthesis unit 53 Second audio signal generation unit 54 Third audio signal generation unit 61 Dedicated terminal p Phase correction signal t Synthesis audio signal x Audio signal x1, x1' First audio signal x2 Second audio signal x3 Third audio signal
Claims
1. An audio processing device comprising: an audio signal generation unit that generates a first audio signal and a second audio signal having the same spectrum from an audio signal representing speech; a filter that removes at least a first frequency component that is equal to or less than a predetermined first frequency from the first audio signal; a synthesis unit that adds and synthesizes the first audio signal from which the first frequency component has been removed and the second audio signal to generate a synthesized audio signal; and a correction signal generation unit that generates a phase correction signal based on the second audio signal, wherein the filter has an operational amplifier that amplifies the first audio signal, and the operational amplifier is configured to receive the first audio signal at its inverting input terminal, and when the phase correction signal is input at its non-inverting input terminal, outputs the first audio signal that has been corrected so that its phase approaches that of the second audio signal.
2. The audio processing device according to claim 1, wherein the filter is a band-pass filter that also removes a second frequency component equal to or greater than a second frequency that is higher than the first frequency.
3. An audio processing device as described in claim 2, wherein the filter is composed of a high-pass filter that removes the first frequency component from the first audio signal, and a low-pass filter that removes the second frequency component from the first audio signal from which the first frequency component has been removed, and the operational amplifier is provided in the low-pass filter.
4. The audio processing device according to claim 1, wherein the audio signal generation unit generates a third audio signal having an equal spectrum to the second audio signal, and the correction signal generation unit generates the phase correction signal based on the third audio signal.
5. The audio processing device according to claim 1, wherein the correction signal generating section comprises a capacitor having one end connected to the audio signal generating section and the other end connected to the non-inverting input terminal.
6. The audio processing device according to claim 1, wherein the first frequency is a frequency between an upper limit frequency of a first-order formant component and a lower limit frequency of a second-order formant component.
7. The audio processing device according to claim 2, wherein the second frequency is a frequency between the upper limit frequency of the fifth formant component and the lower limit frequency of the sixth formant component, or a frequency between the upper limit frequency of the sixth formant component and the lower limit frequency of the seventh formant component.
8. The audio processing device according to claim 3, wherein the high-pass filter includes an operational amplifier that amplifies the first audio signal, and outputs the amplified first audio signal from which the first frequency component has been removed.
9. The audio processing device according to claim 1, further comprising an adjustment unit that adjusts the sound pressure of the first audio signal whose phase has been corrected, and the synthesis unit additively synthesizes the first audio signal whose sound pressure has been adjusted and the second audio signal.
10. A voice input / output system comprising: a voice input unit that acquires a speaker's voice and generates a voice signal; a voice processing device according to any one of claims 1 to 6; and a voice output unit that outputs voice based on the synthesized voice signal generated by the voice processing device.
11. A sound processing method comprising: a sound signal generation step in which a sound signal generator generates a first sound signal and a second sound signal having identical spectra from a sound signal representing speech; a filtering step in which a filter removes at least a first frequency component below a predetermined first frequency from the first sound signal; a sound synthesis step in which a sound synthesizer generates a synthesized sound signal by adding and synthesizing the first sound signal from which the first frequency component has been removed and the second sound signal; and a correction signal generation step in which a correction signal generator generates a phase correction signal based on the second sound signal, wherein the filter has an operational amplifier that amplifies the first sound signal, and the operational amplifier is configured to receive the first sound signal at its inverting input terminal, and in the filtering step, the phase correction signal is input to the non-inverting input terminal of the operational amplifier, causing the operational amplifier to output the first sound signal whose phase has been corrected so as to approach the phase of the second sound signal.
Citation Information
Patent Citations
Voice processor, voice clearing device, and voice processing method
JP2016110050A