Voice processing device and voice input / output system
The audio processing device enhances voice clarity by filtering and phase-correcting audio signals, addressing communication challenges for hard of hearing individuals in call centers and similar settings.
Patent Information
- Application Number
- JP2024086497
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-28
- Publication Date
- 2025-12-10
- Estimated Expiration
- 2044-05-28
AI Technical Summary
Existing voice processing technologies struggle to provide sufficient clarity for hard of hearing individuals, leading to communication difficulties in call centers and similar settings.
An audio processing device that generates and filters audio signals to remove specific frequency components, applies phase correction, and synthesizes them to enhance clarity, using operational amplifiers and filters to adjust and align phases of audio signals.
The device significantly improves voice clarity by correcting phase delays and enhancing signal-to-noise ratios, making speech clearer and more understandable for individuals with hearing impairments.
Smart Images

Figure 2025179618000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an audio input / output system and an audio processing device connectable to the audio input / output system. [Background technology]
[0002] Many companies have established contact points known as call centers or support centers to handle questions and requests from users. In the following, a call center will be used as an example of a contact point. A call center is equipped with one or more telephones. Each operator at the call center communicates with users using the telephones to understand their questions and requests and provide answers to them. Some users who call call centers are hard of hearing. Hard of hearing users may not be able to fully understand what the operator is saying. As a result, communication between the operator and the user may not go smoothly. Therefore, various technologies have been proposed to solve these problems. For example, Patent Document 1 discloses a voice processing device that makes it easier to clearly hear voices coming from speakers in audio equipment, broadcasting facilities, etc. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2016-110050 Summary of the Invention [Problem to be solved by the invention]
[0004] However, in recent years, there has been a demand for voice clarity technology that surpasses conventional voice processing technology such as that described in Patent Document 1.
[0005] One aspect of the present invention has been made in consideration of the above-mentioned problems, and an object of the present invention is to provide a voice processing technique that can make voices clearer than ever before. [Means for solving the problem]
[0006] An audio processing device according to a first aspect of the present invention comprises an audio signal generation unit that generates a first audio signal and a second audio signal having equal spectra from an audio signal representing audio; a filter that removes at least a first frequency component below a predetermined first frequency from the first audio signal; a synthesis unit that adds and synthesizes the first audio signal from which the first frequency component has been removed and the second audio signal to generate a synthesized audio signal; and a correction signal generation unit that generates a phase correction signal based on the second audio signal, wherein the filter has an operational amplifier that amplifies the first audio signal, and the operational amplifier is configured to receive the first audio signal at its inverting input terminal, and to output the first audio signal whose phase has been corrected to approach the phase of the second audio signal when the phase correction signal is input at its non-inverting input terminal.
[0007] According to the above-described voice processing device, voice can be made clearer than ever before.
[0008] An audio processing device according to a second aspect of the present invention may be configured in the above-mentioned first aspect, wherein the filter is a band-pass filter that also removes a second frequency component equal to or greater than a second frequency that is higher than the first frequency.
[0009] According to the above configuration, the signal-to-noise ratio can be improved.
[0010] An audio processing device according to aspect 3 of the present invention may be configured in the above-mentioned aspect 2 such that the filter is composed of a high-pass filter that removes the first frequency component from the first audio signal and a low-pass filter that removes the second frequency component from the first audio signal from which the first frequency component has been removed, and the operational amplifier is provided in the low-pass filter.
[0011] According to the above configuration, the signal is passed through the operational amplifier after all frequency components have been removed (after the phase is no longer delayed), so that the phase delay can be corrected more reliably.
[0012] An audio processing device according to a fourth aspect of the present invention may be configured in the above-mentioned first aspect such that the audio signal generation unit generates a third audio signal having an equal spectrum to the second audio signal, and the correction signal generation unit generates the phase correction signal based on the third audio signal.
[0013] According to the above configuration, the correction signal generator generates the correction signal based on the dedicated third audio signal, which prevents the second audio signal from being out of phase when the second audio signal is used to generate the correction signal.
[0014] An audio processing device according to aspect 5 of the present invention may be configured in the above-mentioned aspect 1 such that the correction signal generation unit includes a capacitor having one end connected to the audio signal generation unit and the other end connected to the non-inverting input terminal.
[0015] With the above configuration, the higher the frequency of the signal passing through the filter, the greater the phase delay, while the capacitor increases the output as the frequency of the input signal increases, and the low-pass filter advances the phase of the output signal as the input to the non-inverting input terminal increases. Therefore, with a simple configuration, it is possible to configure a correction signal generator that can handle audio signals of various frequencies.
[0016] A sixth aspect of the present invention may be configured such that, in the first aspect, the first frequency is a frequency between an upper limit frequency of a first-order formant component and a lower limit frequency of a second-order formant component.
[0017] According to the above configuration, the first formant components are omitted from the first audio signal, and only the second formant components, the third formant components, etc. are left, so that only the second formant components and the third formant components can be amplified.
[0018] An audio processing device according to aspect 7 of the present invention may be configured in the above-mentioned aspect 6 such that the second frequency is a frequency between the upper limit frequency of the 5th formant component and the lower limit frequency of the 6th formant component, or a frequency between the upper limit frequency of the 6th formant component and the lower limit frequency of the 7th formant component.
[0019] According to the above configuration, the signal-to-noise ratio can be further improved.
[0020] An audio processing device according to aspect 8 of the present invention may be configured such that, in aspect 3 above, the high-pass filter includes an operational amplifier that amplifies the first audio signal, and the first frequency component is removed and the amplified first audio signal is output.
[0021] According to the above configuration, the clarity of the voice can be improved.
[0022] An audio processing device according to aspect 9 of the present invention may be configured in accordance with aspect 1 above, further comprising an adjustment unit that adjusts the sound pressure of the first audio signal whose phase has been corrected, and the synthesis unit additively synthesizes the first audio signal whose sound pressure has been adjusted and the second audio signal.
[0023] According to the above configuration, it is possible to output sound according to the usage environment and the degree of hearing loss.
[0024] The voice input / output system according to aspect 10 of the present invention may be configured in any one of aspects 1 to 6 above, including a voice input unit that acquires the voice of a speaker and generates a voice signal, the voice processing device, and a voice output unit that outputs voice based on the synthesized voice signal generated by the voice processing device.
[0025] According to the above configuration, the sound can be made clearer than ever before.
[0026] An audio processing method according to aspect 11 of the present invention includes an audio signal generation step in which an audio signal generator generates a first audio signal and a second audio signal having equal spectra from an audio signal representing audio; a filtering step in which a filter removes at least a first frequency component below a predetermined first frequency from the first audio signal; an audio synthesis step in which an audio synthesizer adds and synthesizes the first audio signal from which the first frequency component has been removed and the second audio signal to generate a synthesized audio signal; and a correction signal generation step in which a correction signal generator generates a phase correction signal based on the second audio signal, wherein the filter has an operational amplifier that amplifies the first audio signal, and the operational amplifier is configured to receive the first audio signal at its inverting input terminal, and in the filtering step, the phase correction signal is input to the non-inverting input terminal of the operational amplifier, causing the operational amplifier to output the first audio signal whose phase has been corrected so as to approach the phase of the second audio signal.
[0027] According to the above-described voice processing method, voice can be made clearer than ever before. [Effects of the Invention]
[0028] According to one aspect of the present invention, it is possible to make speech clearer than ever before. [Brief explanation of the drawings]
[0029] [Figure 1] 1 is a block diagram illustrating an example of a configuration of a voice processing device according to an embodiment of one aspect of the present invention. [Figure 2] FIG. 2 is a diagram illustrating formant components contained in a speech signal. [Figure 3] 2 is a graph showing adjustment characteristics of an input buffer unit included in the audio processing device of FIG. 1. [Figure 4]2 is a circuit diagram showing an example of a specific configuration of a filter and a correction signal generating unit included in the audio processing device of FIG. 1. FIG. [Figure 5] 2 is a graph showing the phase correction characteristics of a filter included in the audio processing device of FIG. [Figure 6] 10 is a graph illustrating a case where a first audio signal and a second audio signal are combined without performing phase correction. [Figure 7] 10 is a graph showing the frequency characteristics of a synthesized speech signal when phase correction is not performed. [Figure 8] 2 is a graph showing an example of adjustment of an audio signal by an adjustment unit included in the audio processing device of FIG. 1. [Figure 9] 10 is a graph showing changes over time in phase and output of an audio signal, a first audio signal before and after correction, and a synthesized audio signal. [Figure 10] 10 is a graph showing the overall output characteristics of a synthesis unit. [Figure 11] 2 is a schematic diagram showing an example of a specific configuration of an auxiliary device in which the audio processing device of FIG. 1 is incorporated. FIG. [Figure 12] 2 is a schematic diagram showing an example of a specific configuration of a voice input / output system in which the voice processing device of FIG. 1 is incorporated. FIG. [Figure 13] FIG. 10 is a schematic diagram showing another example of a specific configuration of the voice input / output system. [Figure 14] FIG. 10 is a schematic diagram showing another example of a specific configuration of the voice input / output system. [Figure 15] FIG. 10 is a schematic diagram showing another example of a specific configuration of the voice input / output system. [Figure 16] FIG. 10 is a schematic diagram showing another example of a specific configuration of the voice input / output system. [Figure 17] One specific configuration of the voice input / output system incorporating the voice processing device of FIG. [Figure 18] 10 is a flowchart illustrating an example of the flow of an audio processing method according to an embodiment of another aspect of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0030] <Sound processing device> Hereinafter, a voice processing device according to an embodiment of the present invention will be described with reference to the drawings.
[0031] [Configuration of voice processing device] 1, the audio processing device 1 includes an audio signal generation unit 11, a correction signal generation unit 12, a filter 13, and a synthesis unit 14. The audio processing device 1 according to this embodiment further includes a first port 15, an input buffer unit 16, an adjustment unit 17, and a second port 18. Note that the audio processing device 1 may further include a power supply unit, a detection unit, and a control unit, which are not shown.
[0032] (Port 1) The first port 15 outputs an audio signal x representing audio supplied from a device connected to the first port 15 to the input buffer unit 16. As shown in FIG. 2, the audio signal x includes a first-order formant component h1, a second-order formant component h2, ..., and an n-order formant component hn. Note that n depends on the user on the transmitting side and, although there may be some individual differences, is a positive integer of at least 4 or more. Note that the first-order formant component h1, the second-order formant component h2, ..., and the n-order formant component hn are described in Patent Document 1, and therefore will not be described in this embodiment.
[0033] (input buffer) As shown in FIG. 1, the input buffer unit 16 is provided between the first port 15 and the audio signal generation unit 11. The input buffer unit 16 can be configured with a microphone amplifier, a preamplifier that lowers output impedance, or the like, depending on the usage situation. The input buffer unit 16 according to this embodiment adjusts the audio signal x to narrow the upper and lower frequency limits and increase the sound pressure, for example, as shown in FIG. 3. Then, as shown in FIG. 1, the input buffer unit 16 outputs the audio signal x that has been adjusted as required to the audio signal generation unit 11. Note that if the audio signal supplied from the first port does not require adjustment (if the device connected to the first port is configured to output an audio signal that does not require adjustment), the audio processing device 1 does not need to include the input buffer unit 16.
[0034] (Audio signal generation section) The audio signal generation unit 11 generates a first audio signal x1 and a second audio signal x2 having the same spectrum from the audio signal x supplied from the input buffer unit 16. "Having the same spectrum" means that the first audio signal x1 and the second audio signal x2 each contain a first-order formant component h1, a second-order formant component h2, ..., and an n-order formant component hn. The audio signal generation unit 11 according to this embodiment also generates a third audio signal x3 having the same spectrum as the second audio signal x2. The audio signal generation unit 11 then supplies the first audio signal x1 to the high-pass filter 131, the second audio signal x2 to the synthesis unit 14, and the third audio signal to the correction signal generation unit 12. The audio signal generation unit 11 according to this embodiment generates the first audio signal x1, the second audio signal x2, and the third audio signal x3 by branching the audio signal. In this embodiment, the audio signal generation unit 11 is configured so that the intensity ratio between the first audio signal x1, the second audio signal x2, and the third audio signal x3 is 1:1:1, i.e., so that the distribution ratio is 1:1:1. However, the distribution ratio of the audio signals is not limited to 1:1:1 and can be set appropriately. The audio signal generation unit 11 may also duplicate the audio signals supplied from the input buffer unit 16, and use the original audio signal as the first audio signal x1, and the duplicated audio signals as the second audio signal x2 and the third audio signal x3.
[0035] (Correction signal generation unit) The correction signal generation unit 12 generates a phase correction signal p based on the second audio signal x2 supplied from the audio signal generation unit 11. As described above, the audio signal generation unit 11 according to this embodiment generates the third audio signal x3 having the same phase as the second audio signal x2. Therefore, the correction signal generation unit 12 according to this embodiment generates the phase correction signal p based on the third audio signal x3. As a result, the correction signal generation unit 12 generates the phase correction signal p based on a dedicated third audio signal. Therefore, it is possible to prevent the second audio signal from being out of phase by using the second audio signal to generate the phase correction signal p. Then, the correction signal generation unit 12 outputs the generated phase correction signal p to the operational amplifier 132a of the filter 13.
[0036] As shown in FIG. 4, the correction signal generation unit 12 according to this embodiment includes a capacitor 121 and a resistor 122. One end of the capacitor 121 is connected to the audio signal generation unit 11, and the other end is connected to the non-inverting input terminal (+) of the operational amplifier 132a. One end of the resistor 122 is connected to the other end of the capacitor 121, and a bias voltage is applied to the other end. The correction signal generation unit 12 configured in this manner outputs a phase correction signal p containing only frequency components equal to or greater than 1 / 2πCR (C: capacitance, R: resistance) using a high-pass filter function internally formed by the capacitor 121 and the resistor 122. Hereinafter, the lower limit value of the frequency components of the phase correction signal p will be referred to as the phase correction start frequency fs. The correction signal generation unit 12 can adjust the phase correction start frequency fs by changing at least one of the capacitor 121 and the resistor 122. As will be described in detail later, the degree to which the operational amplifier 132a of the filter 13 corrects the phase varies within a range of 0 to −180° depending on the phase correction start frequency fs of the phase correction signal p input to the non-inverting input terminal (+), as shown in Fig. 5. For this reason, the phase correction start frequency fs is adjusted to an optimum value that minimizes the phase difference between the first audio signal x1 and the second audio signal x2 after correction.
[0037] Furthermore, the impedance Z (internal resistance) of capacitor 121 changes depending on the phase correction start frequency fs. Specifically, the impedance Z of capacitor 121 is inversely proportional to the phase correction start frequency fs, and its value is expressed as 1 / ωC (ω=2πfs). That is, the impedance Z of capacitor 121 decreases as the phase correction start frequency fs increases. Therefore, the correction signal generation unit 12 outputs a phase correction signal p with a higher voltage as the phase correction start frequency fs increases, and outputs a phase correction signal p with a lower voltage as the phase correction start frequency fs decreases. In this way, by using capacitor 121, the correction signal generation unit 12 can be configured with a simple configuration that is compatible with audio signals of various frequencies.
[0038] (filter) The filter 13 removes at least a first frequency component equal to or lower than a predetermined first frequency from the first audio signal. The first frequency is a frequency between the upper limit frequency of the first formant component h1 and the lower limit frequency of the second formant component h2. The filter 13 according to this embodiment is a band-pass filter that also removes a second frequency component equal to or higher than a second frequency higher than the first frequency. The filter 13 according to this embodiment is composed of a high-pass filter 131 and a low-pass filter 132.
[0039] The high-pass filter 131 removes a first frequency component from the first audio signal x1 supplied from the audio signal generation unit 11. In the high-pass filter 131 according to the present embodiment, the first frequency is set to 400 Hz. The high-pass filter 131 then outputs the first audio signal x1′ from which the first frequency component has been removed, i.e., an audio signal that does not include the first-order formant component h1 but includes the second-order formant component h2, the third-order formant component h3, . . . , the n-order formant component hn, to the low-pass filter 132.
[0040] The high-pass filter 131 according to this embodiment has an operational amplifier (not shown) that amplifies the first audio signal x1. Therefore, the high-pass filter 131 according to this embodiment outputs the first audio signal x1 from which the first frequency component has been removed and which has been amplified. The operational amplifier amplifies the first audio signal x1 from which the first frequency component has been removed, thereby outputting the first audio signal x1" in which the intensities of the second-order formant component h2, the third-order formant component h3, ..., and the n-order formant component hn have been increased. Note that the high-pass filter 131 may be configured so that at least one of the gain of the operational amplifier and the boundary frequency for whether or not to remove the components can be set arbitrarily depending on the intended use.
[0041] The low-pass filter 132 removes a second frequency component from the first audio signal x1, which is supplied from the high-pass filter 131 and from which the first frequency component has been removed. The second frequency component is a frequency component equal to or higher than the second frequency. The second frequency (cutoff frequency) is a frequency between the upper limit frequency of the fifth-order formant component h5 and the lower limit frequency of the sixth-order formant component h6, or a frequency between the upper limit frequency of the sixth-order formant component h6 and the lower limit frequency of the seventh-order formant component h7. In the low-pass filter 132 according to this embodiment, the second frequency is set to 5 kHz or 7 kHz. Low-pass filter 132 then outputs first speech signal x1" from which the fifth, sixth, or higher frequency components have been removed, i.e., a speech signal including only second- to fifth-order formant components h2 to h5 or only second- to sixth-order formant components h2 to h6. Higher-order formant components become noise that does not contribute to language understanding. For this reason, by low-pass filter 132 removing the second frequency component, it is possible to improve the signal-to-noise ratio in synthesized speech signal x', which will be described later.
[0042] However, when frequency components are removed from the first speech signal, a phase delay inevitably occurs. If a first speech signal x1, whose phase is delayed by time T, is combined with a second speech signal x2 without correction, the voltage V1 of the first speech signal x1 becomes −V1 due to the phase delay, as shown in FIG. 6. Therefore, the voltage V of the combined speech signal becomes V2 + (−V1) because the first speech signal x1 and the second speech signal x2 cancel each other out. As a result, as shown in FIG. 7, a dead zone (deep valley) called a dip occurs in a certain frequency range. This significantly reduces the sound pressure of the first-order formant component h1 and the second-order formant component h2, which are particularly important for speech comprehension, resulting in a major obstacle to improving clarity.
[0043] For this reason, the filter 13 has a function of correcting the phase delay. Specifically, as shown in FIG. 4, an operational amplifier 132a is provided in the low-pass filter 132 together with other elements 132b and 132c. The operational amplifier 132a amplifies the first audio signal x1. The operational amplifier 132a is configured so that the first audio signal is input to its inverting input terminal (-). The operational amplifier 132a receives a phase correction signal p at its non-inverting input terminal (+), and outputs a first audio signal x1' whose phase has been corrected to approach the phase of the second audio signal x2. Specifically, the operational amplifier 132a compares the voltage of the first audio signal input to its inverting input terminal (-) and the voltage of the phase correction signal p input to its non-inverting input terminal (+). If the voltage of the phase correction signal p is higher, the operational amplifier 132a outputs the first audio signal x1' whose phase has been advanced in accordance with the voltage difference. On the other hand, if the voltage of the phase correction signal p is lower, the operational amplifier 132a outputs a first audio signal x1' whose phase has been delayed in accordance with the voltage difference. As described above, the voltage of the phase correction signal p varies in accordance with the phase correction start frequency fs determined by the capacitor 121 and resistor 122 of the correction signal generation unit 12. Therefore, the degree to which the operational amplifier 132a advances or delays the phase varies within a range of 0 to -180° depending on the phase correction start frequency fs, as shown in FIG. 5, for example. As described above, the phase correction start frequency fs is adjusted in advance. Therefore, a phase correction signal p having a voltage corresponding to the adjusted phase correction start frequency fs is input to the non-inverting input terminal (+) of the operational amplifier 132a. The operational amplifier 132a then outputs the first audio signal x1' corrected so that the difference in phase with the second audio signal x2 is minimized.
[0044] The filter 13 may be configured such that the low-pass filter 132 first removes the second frequency component from the first audio signal, and the high-pass filter 131 removes the first frequency component from the first audio signal x1 from which the second frequency component has been removed. In this case, the phase correction signal p is input to the non-inverting input terminal (+) of the operational amplifier included in the high-pass filter 131.
[0045] (adjustment section) The adjustment unit 17 adjusts the sound pressure of the phase-corrected first audio signal x1'. Then, the adjustment unit 17 supplies the adjusted first audio signal x1' to the synthesis unit 14. The adjustment unit 17 adjusts the first audio signal based on the user's setting operation according to the usage environment and the degree of hearing loss. Therefore, the adjustment unit 17 can adjust only the first to fifth formant components h1 to h5, which are important for language comprehension, to a level that exceeds the masking line MK', which is an audible level. In addition, the adjustment unit 17 can freely select the output level curve of the synthesized audio signal x' to be "minimum," "maximum," or "middle," as shown in FIG. 8.
[0046] (Synthesis section) The synthesis unit 14 generates a synthetic audio signal x' by additively synthesizing the first audio signal x1', from which the first frequency component and the second frequency component have been removed and whose phase has been corrected, with the second audio signal x2 supplied from the audio signal generation unit 11. As described above, the audio processing device 1 according to this embodiment includes an adjustment unit 17. Therefore, the synthesis unit 14 according to this embodiment additively synthesizes the first audio signal x1, from which the sound pressure has been adjusted, with the second audio signal x2. The synthesis unit 14 then supplies the generated synthetic audio signal x' to the second port 18. As shown in FIG. 9, the phase of the corrected first audio signal x1' in the synthesis unit 14 leads the phase of the uncorrected first audio signal x1 by T'. Therefore, the corrected first audio signal x1' is in phase with the second audio signal x2. As a result, the frequency characteristics of the synthetic audio signal x' output by the synthesis unit 14 are as shown in FIG. 10.
[0047] (2nd port) The second port 18 outputs the synthesized voice signal x′ supplied from the synthesis unit 14 to a device connected to the second port 18.
[0048] (Power supply part) The power supply unit supplies power to each unit requiring power in the audio processing device 1. The type of the power supply unit is not limited, and may be an AC / DC converter such as an AC adapter, or a battery.
[0049] (Detection unit) The detection unit is configured to detect whether power is being supplied from the power supply unit to the audio processing unit. If the power supply unit is functioning normally, the detection unit supplies detection information indicating "Yes" to the control unit. On the other hand, if the power supply unit is not functioning normally, the detection unit supplies detection information indicating "No" to the control unit. If the power supply unit is an AC / DC converter, an example of a state in which the power supply unit is not functioning normally is a power outage. If the power supply unit is a battery, an example of a state in which the power supply unit is not functioning normally is a dead battery.
[0050] (Selection section) The selection unit selects whether or not to perform audio processing on the audio signal x under the control of the control unit. The selection unit is provided, for example, between the low-pass filter 132 and the synthesis unit 14, and is configured by a switch that switches the circuit between ON and OFF. When the selection unit is ON, the first audio signal x1 that has passed through the high-pass filter 131 and the low-pass filter 132 is supplied to the synthesis unit 14, where it is synthesized with the second audio signal x2. On the other hand, when the selection unit is OFF, the first audio signal x1 is not supplied to the synthesis unit 14, and the second audio signal x2 is output from the synthesis unit 14. The selection unit may be provided between the audio signal generation unit 11 and the high-pass filter 131, or between the high-pass filter 131 and the low-pass filter 132.
[0051] (Control unit) The control unit controls the selection unit to apply or not apply audio processing to the audio signal x according to the detection result of the detection unit. If the detection result of the detection unit is "Yes", the control unit controls the selection unit to apply audio processing to the audio signal x (turn ON). On the other hand, if the detection result of the detection unit is "No", the control unit controls the selection unit to not apply audio processing to the audio signal x (turn OFF).
[0052] [Variations of the voice processing device] In the audio processing device 1 according to the above embodiment, the high-pass filter 131, the low-pass filter 132, the correction signal generating unit 12, and the synthesizing unit 14 are configured as analog circuits. However, at least one of these may be configured as a digital circuit.
[0053] Furthermore, the audio processing device 1 according to the above embodiment includes the high-pass filter 131 and the low-pass filter 132. However, instead of including these filters, the audio processing device 1 may include a band-pass filter that simultaneously removes the first frequency component and the second frequency component.
[0054] The audio processing device 1 may further include a polarity switch and an amplifier interposed between the first port 15 and the audio signal generating unit 11.
[0055] The voice processing device 1 also includes a selector, an adjuster, and a selector located between the synthesizer 14 and the second port 18. The amplifier may further include a circuit isolator, a transmitter power regulator, and a polarity switch.
[0056] The audio processing device 1 may also include an attenuator between the synthesis unit 14 and the second port 18. The attenuator adjusts the output to suit the input sound source signal. By including such an attenuator, it becomes possible to provide a system that can accommodate hearing loss symptoms due to hearing impairment by incorporating it into the microphone lines of existing wired or wireless audio transmission devices, or into the signal lines of televisions, radios, and other entertainment playback devices.
[0057] [Effects of the voice processing device] In the voice processing device 1 described above, the low-pass filter 132 corrects the phase delay that occurs when the first and second frequency components are removed based on the phase correction signal p. Therefore, the voice processing device 1 can make voices clearer than ever before.
[0058] <Application examples of voice processing devices> Next, there will be explained various devices incorporating the above-mentioned voice processing device 1. The voice processing device 1 can be incorporated into an auxiliary device 2, voice input / output systems 3 to 6, etc., which will be explained below.
[0059] [Configuration of auxiliary equipment] As shown in FIG. 11, the auxiliary device 2 includes the above-described voice processing device 1, as well as a housing 21, an input terminal 22, and an output terminal 23.
[0060] The housing 21 is formed in a box shape capable of housing the voice processing device 1. The housing 21 is easily portable.
[0061] The input terminal 22 is provided on the surface of the housing 21. The input terminal 22 is connected to a device that outputs an audio signal, and receives an audio signal from the device. The input terminal 22 is connected to the first port 15 of the audio processing device 1. In other words, the input terminal 22 supplies the audio signal input from the device that outputs the audio signal to the first port 15.
[0062] The output terminal 23 is provided on the surface of the housing 21. The output terminal 23 is connected to a device to which the synthetic speech signal x' is input, and outputs a speech signal to that device. The output terminal 23 is connected to the second port 18 of the speech processing device 1. In other words, the output terminal 23 outputs the synthetic speech signal x' supplied from the speech processing device 1 to the device to which the synthetic speech signal x' is input.
[0063] [Action and effect of auxiliary devices] The auxiliary device 2 described above is not only easy to install and remove, but also easy to carry and store, making it easily applicable to broadcasting equipment at outdoor event venues and other audio output devices.
[0064] [Configuration of voice input / output system (1)] 12, the audio input / output system 3 includes, in addition to the audio processing device 1, an audio input unit 31 and an audio output unit 32. The audio input / output system 3 according to this embodiment further includes an amplifier unit 33.
[0065] (Audio input section) The audio input unit 31 acquires the speaker's voice and generates an audio signal. The audio input unit 31 according to this embodiment is configured with a microphone. The audio input unit 31 is connected to a first port of the audio processing device 1. The audio input unit 31 generates an audio signal x based on the input voice and supplies it to the first port 15. The audio input unit 31 may be, for example, a device that plays back video, such as a television, as shown in FIG. 13 .
[0066] (Audio output section) The audio output unit 32 outputs audio based on the synthetic audio signal x' generated by the audio processing device 1. The audio input unit 31 according to this embodiment is configured with a speaker, headphones, earphones, etc. The audio output unit 32 is connected to the second port 18 of the audio processing device 1. The audio output unit 32 then emits audio based on the synthetic audio signal supplied from the second port.
[0067] [Effects of voice input / output systems (1)] The audio input / output system 3 described above can be configured with existing broadcasting equipment, audio transmission devices, etc., except for the audio processing device 1. Therefore, by incorporating the audio input / output system 3 into existing broadcasting equipment, audio transmission devices, etc., it is possible to easily clarify the audio emitted by these devices.
[0068] [Configuration of voice input / output system (2)] As shown in Figure 14, the audio input / output system 4 includes an audio processing device 1, an audio input unit 31, an audio output unit 32, and an amplifier unit 33 similar to those of the above-mentioned audio input / output system 3, as well as a transmitter 41 and a receiver 42.
[0069] The transmitter 41 according to this embodiment is configured as a wireless communication module. The transmitter 41 is connected to the second port 18 of the voice processing device 1. The transmitter 41 converts the synthesized voice signal supplied from the second port into radio waves. The transmitter 41 then transmits the radio waves from the antenna 41a.
[0070] The receiver 42 according to this embodiment is configured as a wireless communication module. The receiver 42 is connected to the amplifier 33. The receiver 42 receives radio waves from the transmitter 41 via an antenna 42a. The receiver 42 converts the received radio waves into a synthesized voice signal and supplies the synthesized voice signal to the amplifier 33.
[0071] 15, the voice input / output system 4 may have voice input unit 31 and voice output unit 32 configured as a telephone receiver, and transmitter 41 and receiver 42 configured as the main body of the telephone. In this case, transmitter 41 transmits the synthesized voice signal to receiver 42 via a fixed telephone line, an internet line, or the like.
[0072] [Effects of voice input / output systems (2)] According to the voice input / output system 4 described above, clear voice can be transmitted to the listener even if the speaker and listener are far apart.
[0073] [Configuration of voice input / output system (3)] 16 , the audio input / output system 5 constitutes a stereo device. The audio input / output system 5 includes a pair of input buffer units 16, a pair of audio signal generators 11, a correction signal generator 12, a filter 13, an adjuster 17, and a pair of synthesizers 14, which are similar to those in the audio processing device 1, and a pair of audio output units 32, which are similar to those in the audio input / output system 3, as well as a pair of audio input terminals 51, a second synthesizer 52, a second audio signal generator 53, and a third audio signal generator 54.
[0074] The pair of audio input terminals 51 are each connected to the input buffer unit 16. A right audio signal x(R) and a left audio signal x(L) are respectively input to the pair of audio input terminals 51. The input right audio signal x(R) and left audio signal x(L) are each supplied to the input buffer unit 16.
[0075] The audio signal generation unit 11 in the audio input / output system 5 does not generate the third audio signal x3, but generates the first audio signals x1(R), x1(L) and the second audio signals x2(R), x2(L).
[0076] The second synthesis unit 52 adds and synthesizes the right first audio signal x1(R) and the left first audio signal x1(L) generated by the pair of audio signal generation units 11. The synthesized first audio signal x1 is supplied to the second audio signal generation unit 53.
[0077] The second audio signal generating unit 53 generates, from the first audio signal x1, a third audio signal having the same spectrum as the first audio signal.
[0078] The third audio signal generation unit 54 generates a right first audio signal x1'(R) and a left first audio signal x1'(L) from the phase-corrected first audio signal x1' supplied from the adjustment unit 17. The first audio signals x1'(R), x1'(L) are supplied to a pair of synthesis units 14, respectively.
[0079] [Effects of voice input / output systems (3)] According to the audio input / output system 5 described above, stereo audio (right audio signal and left audio signal) can be made clearer than ever before.
[0080] [Configuration of voice input / output system (4)] The audio input / output system 6 constitutes a so-called remote conference system. As shown in Fig. 17, the audio input / output system 6 includes a pair of audio processing devices 1, a pair of audio input units 31, and a pair of audio output units 32 similar to those of the audio input / output system 3, as well as a pair of dedicated terminals 61 for the remote conference system. The audio input / output system 6 may also include a camera and a monitor (not shown).
[0081] The audio input unit 31 is connected to a first port of the audio processing device 1. The audio input unit 31 generates an audio signal x based on the input audio and supplies it to the first port 15. The second port 18 of the audio processing device 1 is connected to a dedicated terminal 61. The audio processing device 1 generates a synthesized audio signal from the audio signal x supplied from the audio input unit 31 and supplies it to the dedicated terminal 61.
[0082] The dedicated terminal 61 is connected to another dedicated terminal 61 via an internet line. One dedicated terminal 61 supplies the synthesized voice signal supplied from the voice processing device 1 to the other dedicated terminal. In addition, one dedicated terminal 61 supplies the synthesized voice signal supplied from the other dedicated terminal 61 to the voice output unit 32.
[0083] The audio output unit 32 is connected to the dedicated terminal 61. The audio output unit 32 then emits audio based on the synthesized audio signal supplied from the dedicated terminal 61.
[0084] [Effects of voice input / output systems (4)] According to the voice input / output system 6, the voice uttered by a speaker at one (the other) dedicated terminal can be clearly transmitted to a listener at the other (the other) dedicated terminal. As a result, a person at one dedicated terminal and a person at the other dedicated terminal can smoothly converse (discussion).
[0085] <Audio processing method> Next, a speech processing method according to another embodiment of the present invention will be described.
[0086] As shown in FIG. 18, the voice processing method includes a voice signal generating step S1, a correction signal generating step S2, a filtering step S3, and a voice synthesis step S4.
[0087] (Audio signal generation step) In the initial audio signal generation step S1, an audio signal generator generates a first audio signal x1 and a second audio signal x2 having the same spectrum from an audio signal representing speech. Note that in the audio signal generation step S1, the audio signal generator may also generate a third audio signal having the same phase as the second audio signal. The audio signal generator may be a component of the audio signal generation unit 11 of the audio processing device 1, a component of another device, or an independent device.
[0088] (Correction signal generation step) After generating the first audio signal x1 and the second audio signal x2, the process proceeds to correction signal generation step S2. In correction signal generation step S2, a correction signal generator generates a phase correction signal p based on the second audio signal x2. If a third audio signal is generated in the audio signal generation step S1, the correction signal generator may generate the phase correction signal based on the third audio signal in correction signal generation step S2. The correction signal generator used in correction signal generation step S2 may include a capacitor having one end connected to the audio signal generation unit and the other end connected to the non-inverting input terminal. The correction signal generator may be part of the correction signal generation unit 12 of the audio processing device 1, part of another device, or an independent device.
[0089] (Filtering step) After generating the first audio signal x1 and the second audio signal x2, a filtering step S3 is also performed. In the filtering step S3, a filter removes a first frequency component below a predetermined first frequency from the first audio signal x1 (S31). In addition, in the filtering step S3, a phase correction signal p is input to the non-inverting input terminal of the operational amplifier, causing the operational amplifier to output the first audio signal x1 whose phase has been corrected to approach the phase of the second audio signal x2 (S32). Note that in the filtering step S3, a band-pass filter that also removes a second frequency component above a second frequency higher than the first frequency may be used as the filter. In addition, in the filtering step S3, a filter composed of a high-pass filter that removes the first frequency component from the first audio signal and a low-pass filter that removes the second frequency component from the first audio signal from which the first frequency component has been removed may be used. In this case, it is preferable to use an operational amplifier that is provided in the low-pass filter. Furthermore, in the filtering step S3, a high-pass filter including an operational amplifier for amplifying the first audio signal may be used. The filter may be a component of the audio signal generating unit 11 of the audio processing device 1, a component of another device, or an independent device. Furthermore, in the filtering step S3, the removal of the second frequency component may be performed before the removal of the first frequency component. Furthermore, the removal of the first frequency component and the second frequency component may be performed simultaneously.
[0090] (Speech synthesis step) After the second frequency component is removed from the first audio signal x1, the process proceeds to audio synthesis step S4. In audio synthesis step S4, a audio synthesizer adds and synthesizes the phase-corrected first audio signal x1 and the second audio signal x2 to generate a synthesized audio signal x'. The audio synthesizer may be part of the audio signal generation unit 11 of the audio processing device 1, part of another device, or may be an independent device.
[0091] [Variations of the audio processing method] The voice processing method may further include an adjustment step. The adjustment step is preferably performed between the filtering step and the voice synthesis step. The adjustment step adjusts the sound pressure of the phase-corrected first voice signal. When the adjustment step is included, in the voice synthesis step, the voice synthesizer additively synthesizes the first voice signal whose sound pressure has been adjusted and the second voice signal.
[0092] [Action and effect of audio processing method] In the audio processing method described above, in filtering step S3, the filter corrects the phase delay that occurs when the first and second frequency components are removed based on the phase correction signal p. Therefore, the audio processing method can make audio clearer than ever before.
[0093] <Additional Notes> The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in the description for implementing the invention are also included in the technical scope of the present invention.
[0094] <Summary> An audio processing device according to a first aspect of the present invention comprises an audio signal generation unit that generates a first audio signal and a second audio signal having equal spectra from an audio signal representing audio; a filter that removes at least a first frequency component below a predetermined first frequency from the first audio signal; a synthesis unit that adds and synthesizes the first audio signal from which the first frequency component has been removed and the second audio signal to generate a synthesized audio signal; and a correction signal generation unit that generates a phase correction signal based on the second audio signal, wherein the filter has an operational amplifier that amplifies the first audio signal, and the operational amplifier is configured to receive the first audio signal at its inverting input terminal, and to output the first audio signal whose phase has been corrected to approach the phase of the second audio signal when the phase correction signal is input at its non-inverting input terminal.
[0095] According to the above-described voice processing device, voice can be made clearer than ever before.
[0096] An audio processing device according to a second aspect of the present invention may be configured in the above-mentioned first aspect, wherein the filter is a band-pass filter that also removes a second frequency component equal to or greater than a second frequency that is higher than the first frequency.
[0097] According to the above configuration, the signal-to-noise ratio can be improved.
[0098] An audio processing device according to aspect 3 of the present invention may be configured in the above-mentioned aspect 2 such that the filter is composed of a high-pass filter that removes the first frequency component from the first audio signal and a low-pass filter that removes the second frequency component from the first audio signal from which the first frequency component has been removed, and the operational amplifier is provided in the low-pass filter.
[0099] According to the above configuration, the signal is passed through the operational amplifier after all frequency components have been removed (after the phase is no longer delayed), so that the phase delay can be corrected more reliably.
[0100] An audio processing device according to a fourth aspect of the present invention may be configured in the above-mentioned first aspect such that the audio signal generation unit generates a third audio signal having an equal spectrum to the second audio signal, and the correction signal generation unit generates the phase correction signal based on the third audio signal.
[0101] According to the above configuration, the correction signal generator generates the correction signal based on the dedicated third audio signal, which prevents the second audio signal from being out of phase when the second audio signal is used to generate the correction signal.
[0102] An audio processing device according to aspect 5 of the present invention may be configured in the above-mentioned aspect 1 such that the correction signal generation unit includes a capacitor having one end connected to the audio signal generation unit and the other end connected to the non-inverting input terminal.
[0103] With the above configuration, the higher the frequency of the signal passing through the filter, the greater the phase delay, while the capacitor increases the output as the frequency of the input signal increases, and the low-pass filter advances the phase of the output signal as the input to the non-inverting input terminal increases. Therefore, with a simple configuration, it is possible to configure a correction signal generator that can handle audio signals of various frequencies.
[0104] A sixth aspect of the present invention may be configured such that, in the first aspect, the first frequency is a frequency between an upper limit frequency of a first-order formant component and a lower limit frequency of a second-order formant component.
[0105] According to the above configuration, the first formant components are omitted from the first audio signal, and only the second formant components, the third formant components, etc. are left, so that only the second formant components and the third formant components can be amplified.
[0106] An audio processing device according to aspect 7 of the present invention may be configured in the above-mentioned aspect 6 such that the second frequency is a frequency between the upper limit frequency of the 5th formant component and the lower limit frequency of the 6th formant component, or a frequency between the upper limit frequency of the 6th formant component and the lower limit frequency of the 7th formant component.
[0107] According to the above configuration, the signal-to-noise ratio can be further improved.
[0108] An audio processing device according to aspect 8 of the present invention may be configured such that, in aspect 3 above, the high-pass filter includes an operational amplifier that amplifies the first audio signal, and the first frequency component is removed and the amplified first audio signal is output.
[0109] According to the above configuration, the clarity of the voice can be improved.
[0110] An audio processing device according to aspect 9 of the present invention may be configured in accordance with aspect 1 above, further comprising an adjustment unit that adjusts the sound pressure of the first audio signal whose phase has been corrected, and the synthesis unit additively synthesizes the first audio signal whose sound pressure has been adjusted and the second audio signal.
[0111] According to the above configuration, it is possible to output sound according to the usage environment and the degree of hearing loss.
[0112] The voice input / output system according to aspect 10 of the present invention may be configured in any one of aspects 1 to 6 above, including a voice input unit that acquires the voice of a speaker and generates a voice signal, the voice processing device, and a voice output unit that outputs voice based on the synthesized voice signal generated by the voice processing device.
[0113] According to the above configuration, the sound can be made clearer than ever before.
[0114] An audio processing method according to aspect 11 of the present invention includes an audio signal generation step in which an audio signal generator generates a first audio signal and a second audio signal having equal spectra from an audio signal representing audio; a filtering step in which a filter removes at least a first frequency component below a predetermined first frequency from the first audio signal; an audio synthesis step in which an audio synthesizer adds and synthesizes the first audio signal from which the first frequency component has been removed and the second audio signal to generate a synthesized audio signal; and a correction signal generation step in which a correction signal generator generates a phase correction signal based on the second audio signal, wherein the filter has an operational amplifier that amplifies the first audio signal, and the operational amplifier is configured to receive the first audio signal at its inverting input terminal, and in the filtering step, the phase correction signal is input to the non-inverting input terminal of the operational amplifier, causing the operational amplifier to output the first audio signal whose phase has been corrected so as to approach the phase of the second audio signal.
[0115] According to the above-described voice processing method, voice can be made clearer than ever before. [Explanation of symbols]
[0116] 1. Audio processing device 11 Audio signal generation unit 12 Correction signal generator 121 Capacitor 122 resistor 13 Filters 131 High-pass filter 132 Low-pass filter 132a operational amplifier 132b, 132c elements 14 Synthesis section 15 Port 1 16 Input buffer 17 Adjustment part 18 Second Port 2 Auxiliary equipment 21. Cabinet 22 Input terminal 23 Output terminal 3, 4, 5, 6 Audio Input / Output System 31 Audio input section 32 Audio output section 33 Amplification section 41 Transmitter 41a, 42a antennas 42 Receiver 51 Audio input terminal 52 2nd synthesis section 53 Second audio signal generation unit 54 Third audio signal generation unit 61 Dedicated terminal p phase correction signal t Synthesized speech signal x Audio signal x1, x1' 1st audio signal x2 Second audio signal x3 Third audio signal
Claims
1. an audio signal generation unit that generates a first audio signal and a second audio signal having the same spectrum from an audio signal representing a speech; a filter that removes at least a first frequency component that is equal to or lower than a predetermined first frequency from the first audio signal; a synthesis unit that generates a synthesized audio signal by additively synthesizing the first audio signal from which the first frequency component has been removed and the second audio signal; a correction signal generation unit that generates a phase correction signal based on the second audio signal; Equipped with the filter includes an operational amplifier that amplifies the first audio signal; The operational amplifier comprises: The first audio signal is input to an inverting input terminal, When the phase correction signal is input to a non-inverting input terminal, the first audio signal is output, the phase of which is corrected so as to approach the phase of the second audio signal. Audio processing device.
2. the filter is a band-pass filter that also removes a second frequency component equal to or greater than a second frequency that is higher than the first frequency; The audio processing device according to claim 1 .
3. The filter is a high-pass filter that removes the first frequency component from the first audio signal; a low-pass filter that removes the second frequency component from the first audio signal from which the first frequency component has been removed; It is composed of The operational amplifier is provided in the low-pass filter. The audio processing device according to claim 2 .
4. the audio signal generation unit generates a third audio signal having an equal spectrum to the second audio signal; the correction signal generation unit generates the phase correction signal based on the third audio signal. The audio processing device according to claim 1 .
5. the correction signal generation unit includes a capacitor having one end connected to the audio signal generation unit and the other end connected to the non-inverting input terminal. The audio processing device according to claim 1 .
6. the first frequency is a frequency between an upper limit frequency of a first formant component and a lower limit frequency of a second formant component; The audio processing device according to claim 1 .
7. the second frequency is a frequency between an upper limit frequency of a fifth formant component and a lower limit frequency of a sixth formant component, or a frequency between an upper limit frequency of a sixth formant component and a lower limit frequency of a seventh formant component. The audio processing device according to claim 2 .
8. The high-pass filter is an operational amplifier that amplifies the first audio signal; outputting the first audio signal from which the first frequency component has been removed and which has been amplified; The audio processing device according to claim 3 .
9. an adjustment unit that adjusts the sound pressure of the phase-corrected first audio signal, the synthesis unit additively synthesizes the first audio signal, the sound pressure of which has been adjusted, and the second audio signal. The audio processing device according to claim 1 .
10. a voice input unit that acquires a speaker's voice and generates a voice signal; A voice processing device according to any one of claims 1 to 6; a voice output unit that outputs a voice based on the synthesized voice signal generated by the voice processing device; Equipped with Audio input / output system.
11. an audio signal generating step in which an audio signal generator generates a first audio signal and a second audio signal having the same spectrum from an audio signal representing speech; a filtering step in which a filter removes at least a first frequency component below a predetermined first frequency from the first audio signal; a voice synthesis step in which a voice synthesizer additively synthesizes the first voice signal from which the first frequency component has been removed and the second voice signal to generate a synthesized voice signal; a correction signal generating step in which a correction signal generator generates a phase correction signal based on the second audio signal; Including, the filter includes an operational amplifier that amplifies the first audio signal; the operational amplifier is configured to receive the first audio signal at an inverting input terminal; In the filtering step, the phase correction signal is input to a non-inverting input terminal of the operational amplifier, thereby causing the operational amplifier to output the first audio signal whose phase has been corrected so as to approach the phase of the second audio signal. Audio processing methods.
Citation Information
Patent Citations
Voice processor, voice clearing device, and voice processing method
JP2016110050A