Audio processing device and audio input / output system

The audio processing device addresses the challenge of voice clarity by generating and processing audio signals with equal spectra, removing specific frequency components, and applying phase correction, thereby enhancing voice clarity and communication effectiveness.

JP7682346B1Active Publication Date: 2025-05-23RADIUS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024086497
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-05-28
Publication Date
2025-05-23
Estimated Expiration
2044-05-28

Smart Images

  • Figure 0007682346000001_ABST
    Figure 0007682346000001_ABST
Patent Text Reader

Abstract

To provide a voice processing technology that can make voice clearer than ever before. [Solution] The system includes an audio signal generation unit (11) that generates a first audio signal (x1) and a second audio signal (x2) having equal spectra from an audio signal representing speech, a filter (13) that removes at least a first frequency component below a predetermined first frequency from the first audio signal, a synthesis unit (14) that adds and synthesizes the first audio signal from which the first frequency component has been removed and the second audio signal to generate a synthesized audio signal, and a correction signal generation unit (12) that generates a phase correction signal (p) based on the second audio signal, and the filter has an operational amplifier (132a) that amplifies the first audio signal, and the operational amplifier is configured to receive the first audio signal at its inverting input terminal, and outputs the first audio signal whose phase has been corrected to approach the phase of the second audio signal when the phase correction signal is input to its non-inverting input terminal.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an audio input / output system and an audio processing device connectable to an audio input / output system. [Background technology]

[0002] In order to receive questions and requests from users, many companies have established a contact point called a call center or a support center. In the following, a call center is used as an example of a contact point. One or more telephones are installed in the call center. Each operator of the call center uses the telephone to talk to the user, understands the questions and requests from the user, and conveys answers to those questions to the user. Meanwhile, some users who call the call center are hard of hearing. A hard of hearing user may not be able to fully understand what the operator says. As a result, communication between the operator and the user may not go well. Therefore, various technologies have been proposed to solve such problems. For example, Patent Document 1 discloses a voice processing device for making it easier to clearly hear the voice coming from a speaker of an audio device or a broadcasting facility. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2016-110050 A Summary of the Invention [Problem to be solved by the invention]

[0004] However, in recent years, there has been a demand for voice clarity technology that surpasses conventional voice processing technology such as that described in Patent Document 1.

[0005] One aspect of the present invention has been made in consideration of the above-mentioned problems, and an object of the present invention is to provide a voice processing technique that can make voice clearer than ever before. [Means for solving the problem]

[0006] A first aspect of the present invention provides an audio processing device comprising: an audio signal generation unit that generates a first audio signal and a second audio signal having equal spectra from an audio signal representing audio; a filter that removes at least a first frequency component below a predetermined first frequency from the first audio signal; a synthesis unit that additively synthesizes the first audio signal from which the first frequency component has been removed and the second audio signal to generate a synthesized audio signal; and a correction signal generation unit that generates a phase correction signal based on the second audio signal, wherein the filter has an operational amplifier that amplifies the first audio signal, and the operational amplifier is configured to receive the first audio signal at an inverting input terminal, and outputs the first audio signal whose phase has been corrected to approach the phase of the second audio signal when the phase correction signal is input to a non-inverting input terminal.

[0007] According to the above voice processing device, voice can be made clearer than ever before.

[0008] The audio processing device according to a second aspect of the present invention may be configured in the above-mentioned first aspect such that the filter is a band-pass filter that also removes a second frequency component equal to or greater than a second frequency that is higher than the first frequency.

[0009] According to the above configuration, the signal-to-noise ratio can be improved.

[0010] An audio processing device according to aspect 3 of the present invention may be configured in the above-mentioned aspect 2 such that the filter is composed of a high-pass filter that removes the first frequency component from the first audio signal, and a low-pass filter that removes the second frequency component from the first audio signal from which the first frequency component has been removed, and the operational amplifier is provided in the low-pass filter.

[0011] According to the above configuration, the signal is passed through the operational amplifier after all frequency components have been removed (after the phase is no longer delayed any further), so that the phase delay can be corrected more reliably.

[0012] An audio processing device according to aspect 4 of the present invention may be configured in the above-mentioned aspect 1 such that the audio signal generation unit generates a third audio signal having an equal spectrum to the second audio signal, and the correction signal generation unit generates the phase correction signal based on the third audio signal.

[0013] According to the above configuration, the correction signal generator generates the correction signal based on the dedicated third audio signal, which makes it possible to prevent the second audio signal from being out of phase with the second audio signal when the second audio signal is used to generate the correction signal.

[0014] An audio processing device according to aspect 5 of the present invention may be configured in such a way that, in aspect 1 above, the correction signal generating unit has a capacitor having one end connected to the audio signal generating unit and the other end connected to the non-inverting input terminal.

[0015] With the above configuration, the higher the frequency of the signal passing through the filter, the greater the phase delay, while the capacitor increases the output as the frequency of the input signal increases, and the low-pass filter advances the phase of the output signal as the input to the non-inverting input terminal increases. Therefore, with a simple configuration, it is possible to configure a correction signal generator that is compatible with audio signals of various frequencies.

[0016] The audio processing device according to a sixth aspect of the present invention may be configured in the above-mentioned first aspect, wherein the first frequency is a frequency between an upper limit frequency of a first-order formant component and a lower limit frequency of a second-order formant component.

[0017] According to the above configuration, the first formant components are omitted from the first audio signal, and only the second formant components, the third formant components, etc. are left, so that only the second formant components and the third formant components can be amplified.

[0018] An audio processing device according to aspect 7 of the present invention may be configured in such a way that, in aspect 6 above, the second frequency is a frequency between the upper limit frequency of the 5th formant component and the lower limit frequency of the 6th formant component, or a frequency between the upper limit frequency of the 6th formant component and the lower limit frequency of the 7th formant component.

[0019] According to the above configuration, the signal-to-noise ratio can be further improved.

[0020] An audio processing device according to aspect 8 of the present invention may be configured in such a way that, in aspect 3 above, the high-pass filter includes an operational amplifier that amplifies the first audio signal, and the first frequency component is removed and the amplified first audio signal is output.

[0021] According to the above configuration, the clarity of the voice can be improved.

[0022] An audio processing device according to aspect 9 of the present invention may be configured in accordance with aspect 1 above, further comprising an adjustment unit that adjusts the sound pressure of the first audio signal whose phase has been corrected, and the synthesis unit additively synthesizes the first audio signal whose sound pressure has been adjusted and the second audio signal.

[0023] According to the above configuration, it is possible to output a sound according to the usage environment and the degree of hearing loss.

[0024] The voice input / output system of aspect 10 of the present invention may be configured in any one of aspects 1 to 6 above, comprising a voice input unit that acquires a speaker's voice and generates a voice signal, the voice processing device, and a voice output unit that outputs voice based on the synthetic voice signal generated by the voice processing device.

[0025] According to the above configuration, the sound can be made clearer than ever before.

[0026] An audio processing method according to aspect 11 of the present invention includes an audio signal generation step in which an audio signal generator generates a first audio signal and a second audio signal having equal spectra from an audio signal representing audio; a filtering step in which a filter removes at least a first frequency component below a predetermined first frequency from the first audio signal; a audio synthesis step in which a audio synthesizer additively synthesizes the first audio signal from which the first frequency component has been removed and the second audio signal to generate a synthetic audio signal; and a correction signal generation step in which a correction signal generator generates a phase correction signal based on the second audio signal, wherein the filter has an operational amplifier that amplifies the first audio signal, and the operational amplifier is configured to receive the first audio signal at its inverting input terminal, and in the filtering step, the phase correction signal is input to the non-inverting input terminal of the operational amplifier, causing the operational amplifier to output the first audio signal whose phase has been corrected so as to approach the phase of the second audio signal.

[0027] According to the above voice processing method, voice can be made clearer than ever before. Effect of the Invention

[0028] According to one aspect of the present invention, voice can be made clearer than ever before. [Brief description of the drawings]

[0029] [Figure 1] 1 is a block diagram showing an example of a configuration of a voice processing device according to an embodiment of one aspect of the present invention. [Diagram 2] FIG. 2 is a diagram illustrating formant components contained in a speech signal. [Diagram 3] 2 is a graph showing adjustment characteristics of an input buffer section included in the audio processing device of FIG. 1. [Figure 4]2 is a circuit diagram showing an example of a specific configuration of a filter and a correction signal generating unit included in the audio processing device of FIG. 1. [Diagram 5] 2 is a graph showing a phase correction characteristic of a filter included in the audio processing device of FIG. [Figure 6] 11 is a graph illustrating a case where a first audio signal and a second audio signal are combined without performing phase correction. [Figure 7] 11 is a graph showing frequency characteristics of a synthetic speech signal when phase correction is not performed. [Figure 8] 4 is a graph showing an example of adjustment of an audio signal by an adjustment unit included in the audio processing device of FIG. [Figure 9] 6 is a graph showing changes over time in phase and output of an audio signal, a first audio signal before and after correction, and a synthetic audio signal. [Figure 10] 13 is a graph showing the overall output characteristic of a synthesis unit. [Figure 11] 2 is a schematic diagram showing an example of a specific configuration of an auxiliary device in which the voice processing device of FIG. 1 is incorporated. [Figure 12] 2 is a schematic diagram showing an example of a specific configuration of a voice input / output system in which the voice processing device of FIG. 1 is incorporated. [Figure 13] FIG. 13 is a schematic diagram showing another example of a specific configuration of the voice input / output system. [Figure 14] FIG. 13 is a schematic diagram showing another example of a specific configuration of the voice input / output system. [Figure 15] FIG. 13 is a schematic diagram showing another example of a specific configuration of the voice input / output system. [Figure 16] FIG. 13 is a schematic diagram showing another example of a specific configuration of the voice input / output system. [Figure 17] One specific configuration of a voice input / output system incorporating the voice processing device of FIG. [Figure 18] 10 is a flowchart showing an example of the flow of a sound processing method according to an embodiment of another aspect of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0030] <Audio processing device> DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, a voice processing device according to an embodiment of the present invention will be described with reference to the drawings.

[0031] [Configuration of audio processing device] 1, the audio processing device 1 includes an audio signal generating unit 11, a correction signal generating unit 12, a filter 13, and a synthesis unit 14. The audio processing device 1 according to this embodiment further includes a first port 15, an input buffer unit 16, an adjustment unit 17, and a second port 18. Note that the audio processing device 1 may further include a power supply unit, a detection unit, and a control unit, which are not shown.

[0032] (1st port) The first port 15 outputs an audio signal x representing a voice supplied from a device connected to the first port 15 to the input buffer unit 16. As shown in Fig. 2, the audio signal x includes a first formant component h1, a second formant component h2, ..., and an nth formant component hn. Note that n depends on the user on the transmitting side and is a positive integer of at least 4, although there may be some individual differences. Note that the first formant component h1, the second formant component h2, ..., and the nth formant component hn are described in Patent Document 1, and therefore will not be described in this embodiment.

[0033] (Input buffer) As shown in FIG. 1, the input buffer unit 16 is provided between the first port 15 and the audio signal generating unit 11. The input buffer unit 16 can be configured with a microphone amplifier, a preamplifier for lowering output impedance, etc., depending on the usage situation. The input buffer unit 16 according to this embodiment adjusts the audio signal x to narrow the upper limit and the adjustment of the frequency and to increase the sound pressure, for example, as shown in FIG. 3. Then, the input buffer unit 16 outputs the audio signal x that has been subjected to the necessary adjustment to the audio signal generating unit 11, as shown in FIG. 1. Note that, when adjustment of the audio signal supplied from the first port is not required (when the device connected to the first port is configured to output an audio signal that does not require adjustment), the audio processing device 1 does not need to include the input buffer unit 16.

[0034] (Audio signal generation section) The audio signal generating unit 11 generates a first audio signal x1 and a second audio signal x2 having the same spectrum from the audio signal x supplied from the input buffer unit 16. The "having the same spectrum" means that the first audio signal x1 and the second audio signal x2 each contain a first-order formant component h1, a second-order formant component h2, ..., and an n-order formant component hn. The audio signal generating unit 11 according to this embodiment further generates a third audio signal x3 having the same spectrum as the second audio signal x2. The audio signal generating unit 11 then supplies the first audio signal x1 to the high-pass filter 131, the second audio signal x2 to the synthesis unit 14, and the third audio signal to the correction signal generating unit 12. The audio signal generating unit 11 according to this embodiment generates the first audio signal x1, the second audio signal x2, and the third audio signal x3 by branching the audio signal. In this embodiment, the audio signal generating unit 11 is configured so that the intensity ratio between the first audio signal x1, the second audio signal x2, and the third audio signal x3 is 1:1:1, that is, the distribution ratio is 1:1:1. However, the distribution ratio of the audio signals is not limited to 1:1:1 and can be set appropriately. The audio signal generating unit 11 may also copy the audio signal supplied from the input buffer unit 16, and set the original audio signal as the first audio signal x1, and the copied audio signal as the second audio signal x2 and the third audio signal x3.

[0035] (Correction signal generation unit) The correction signal generating unit 12 generates a phase correction signal p based on the second audio signal x2 supplied from the audio signal generating unit 11. As described above, the audio signal generating unit 11 according to this embodiment generates a third audio signal x3 having the same phase as the second audio signal x2. Therefore, the correction signal generating unit 12 according to this embodiment generates a phase correction signal p based on the third audio signal x3. As a result, the correction signal generating unit 12 generates the phase correction signal p based on a dedicated third audio signal. Therefore, it is possible to prevent the second audio signal from being out of phase by using the second audio signal to generate the phase correction signal p. Then, the correction signal generating unit 12 outputs the generated phase correction signal p to the operational amplifier 132a of the filter 13.

[0036] As shown in FIG. 4, the correction signal generating unit 12 according to the present embodiment includes a capacitor 121 and a resistor 122. One end of the capacitor 121 is connected to the audio signal generating unit 11, and the other end is connected to the non-inverting input terminal (+) of the operational amplifier 132a. One end of the resistor 122 is connected to the other end of the capacitor 121, and a bias voltage is applied to the other end. The correction signal generating unit 12 configured in this manner outputs a phase correction signal p having only frequency components equal to or higher than 1 / 2πCR (C: electrostatic capacitance, R: resistance value) by a high-pass filter function internally configured by the capacitor 121 and the resistor 122. Hereinafter, the lower limit value of the frequency components of the phase correction signal p is referred to as a phase correction start frequency fs. The correction signal generating unit 12 can adjust the phase correction start frequency fs by changing at least one of the capacitor 121 and the resistor 122. As will be described later in detail, the degree to which the operational amplifier 132a of the filter 13 corrects the phase varies within a range of 0 to -180° depending on the phase correction start frequency fs of the phase correction signal p input to the non-inverting input terminal (+) as shown in Fig. 5. For this reason, the phase correction start frequency fs is adjusted to an optimum value that minimizes the phase difference between the corrected first audio signal x1 and second audio signal x2.

[0037] Also, the impedance Z (internal resistance) of the capacitor 121 changes according to the phase correction start frequency fs. Specifically, the impedance Z of the capacitor 121 is inversely proportional to the phase correction start frequency fs, and its value is represented by 1 / ωC (ω = 2πfs). That is, the higher the phase correction start frequency fs, the lower the impedance Z of the capacitor 121. For this reason, the correction signal generation unit 12 outputs a phase correction signal p with a higher voltage when the phase correction start frequency fs is high, and outputs a phase correction signal p with a lower voltage when the phase correction start frequency fs is low. In this way, by using the capacitor 121, the correction signal generation unit 12 corresponding to audio signals of various frequencies can be configured with a simple structure.

[0038] (Filter) The filter 13 removes at least the first frequency components below a predetermined first frequency from the first audio signal. The first frequency is a frequency between the upper limit frequency of the first formant component h1 and the lower limit frequency of the second formant component h2. The filter 13 according to the present embodiment is a band-pass filter that also removes the second frequency components of the second frequency or higher that are higher than the first frequency. The filter 13 according to the present embodiment is composed of a high-pass filter 131 and a low-pass filter 132.

[0039] The high-pass filter 131 removes the first frequency components from the first audio signal x1 supplied from the audio signal generation unit 11. In the high-pass filter 131 according to the present embodiment, the first frequency is set to 400 Hz. Then, the high-pass filter 131 outputs the first audio signal x1' from which the first frequency components have been removed, that is, an audio signal that does not include the first formant component h1 and includes the second formant component h2, the third formant component h3, ···, the nth formant component hn, to the low-pass filter 132.

[0040] The high-pass filter 131 according to this embodiment has an operational amplifier (not shown) that amplifies the first audio signal x1. Therefore, the high-pass filter 131 according to this embodiment outputs the first audio signal x1 from which the first frequency component has been removed and which has been amplified. The operational amplifier amplifies the first audio signal x1 from which the first frequency component has been removed, thereby outputting the first audio signal x1" in which the intensities of the second-order formant component h2, the third-order formant component h3, ..., and the n-th order formant component hn have been increased. Note that the high-pass filter 131 may be configured so that at least one of the gain of the operational amplifier and the frequency that is the boundary between removal and non-removal can be set arbitrarily depending on the purpose of use.

[0041] The low-pass filter 132 removes a second frequency component from the first audio signal x1, from which the first frequency component has been removed and which is supplied from the high-pass filter 131. The second frequency component is a frequency component equal to or higher than the second frequency. The second frequency (cutoff frequency) is a frequency between the upper limit frequency of the fifth formant component h5 and the lower limit frequency of the sixth formant component h6, or a frequency between the upper limit frequency of the sixth formant component h6 and the lower limit frequency of the seventh formant component h7. In the low-pass filter 132 according to this embodiment, the second frequency is set to 5 kHz or 7 kHz. Low-pass filter 132 then outputs first speech signal x1" from which the fifth, sixth and higher frequency components have been removed, i.e., a speech signal including only second- to fifth-order formant components h2-h5 or only second- to sixth-order formant components h2-h6. Higher-order formant components become noise that does not contribute to language understanding. For this reason, by low-pass filter 132 removing the second frequency component, it is possible to improve the signal-to-noise ratio in synthetic speech signal x', which will be described later.

[0042] However, when removing frequency components, a phase delay inevitably occurs in the first voice signal. If the first voice signal x1, whose phase is delayed by time T, is synthesized with the second voice signal x2 without correction, the voltage V1 of the first voice signal x1 becomes -V1 due to the phase delay, as shown in FIG. 6. Therefore, the voltage V of the synthesized voice signal becomes V2+(-V1) because the first voice signal x1 and the second voice signal x2 cancel each other out. As a result, as shown in FIG. 7, a dead zone (deep valley) called a dip occurs in a certain frequency range, and the sound pressure of the first formant component h1 and the second formant component h2, which are particularly important for language understanding, becomes extremely small, which is a major obstacle to improving clarity.

[0043] For this reason, the filter 13 has a function of correcting the phase delay. Specifically, as shown in FIG. 4, an operational amplifier 132a is provided in the low-pass filter 132 together with other elements 132b and 132c. The operational amplifier 132a amplifies the first audio signal x1. The operational amplifier 132a is configured so that the first audio signal is input to an inverting input terminal (-). The operational amplifier 132a outputs a first audio signal x1' whose phase is corrected to approach the phase of the second audio signal x2 by inputting a phase correction signal p to a non-inverting input terminal (+). Specifically, the operational amplifier 132a compares and calculates both the voltage of the first audio signal input to the inverting input terminal (-) and the voltage of the phase correction signal p input to the non-inverting input terminal (+). If the voltage of the phase correction signal p is higher, the operational amplifier outputs the first audio signal x1' whose phase is advanced according to the voltage difference. On the other hand, if the voltage of the phase correction signal p is lower, the first audio signal x1' whose phase is delayed according to the voltage difference is output. As described above, the voltage of the phase correction signal p changes according to the phase correction start frequency fs determined by the capacitor 121 and the resistor 122 of the correction signal generating unit 12. Therefore, the degree to which the operational amplifier 132a advances or delays the phase varies within a range of 0 to -180° depending on the phase correction start frequency fs, as shown in FIG. 5, for example. As described above, the phase correction start frequency fs is adjusted in advance. Therefore, the phase correction signal p whose voltage corresponds to the adjusted phase correction start frequency fs is input to the non-inverting input terminal (+) of the operational amplifier 132a. Then, the operational amplifier 132a outputs the first audio signal x1' corrected so that the difference in phase with the second audio signal x2 is minimized.

[0044] The filter 13 may be configured such that the low-pass filter 132 first removes the second frequency component from the first audio signal, and the high-pass filter 131 removes the first frequency component from the first audio signal x1 from which the second frequency component has been removed. In this case, the phase correction signal p is input to the non-inverting input terminal (+) of the operational amplifier included in the high-pass filter 131.

[0045] (adjustment section) The adjustment unit 17 adjusts the sound pressure of the phase-corrected first audio signal x1'. Then, the adjustment unit 17 supplies the adjusted first audio signal x1' to the synthesis unit 14. The adjustment unit 17 adjusts the first audio signal based on a user's setting operation according to the usage environment and the degree of hearing loss. Therefore, the adjustment unit 17 can adjust only the first to fifth formant components h1 to h5, which are important for language understanding, to a level exceeding the masking line MK', which is an audible level. In addition, the adjustment unit 17 can freely select the output level curve of the synthetic audio signal x' to be "minimum", "maximum", or "middle" as shown in FIG. 8.

[0046] (Synthetic section) The synthesis unit 14 generates a synthetic voice signal x' by additively synthesizing the first voice signal x1' from which the first frequency component and the second frequency component have been removed and the phase has been corrected, and the second voice signal x2 supplied from the voice signal generation unit 11. As described above, the voice processing device 1 according to this embodiment includes an adjustment unit 17. For this reason, the synthesis unit 14 according to this embodiment additively synthesizes the first voice signal x1 from which the sound pressure has been adjusted, and the second voice signal x2. Then, the synthesis unit 14 supplies the generated synthetic voice signal x' to the second port 18. As shown in FIG. 9, the phase of the corrected first voice signal x1' in the synthesis unit 14 leads the phase of the first voice signal x1 before correction by T', so that the corrected first voice signal x1' is in phase with the second voice signal x2. As a result, the frequency characteristic of the synthetic voice signal x' output by the synthesis unit 14 becomes as shown in FIG. 10.

[0047] (2nd port) The second port 18 outputs the synthesized voice signal x′ supplied from the synthesis unit 14 to a device connected to the second port 18 .

[0048] (Power supply part) The power supply unit supplies power to each unit requiring power in the audio processing device 1. The type of the power supply unit is not limited, and may be an AC / DC converter such as an AC adapter, or a battery.

[0049] (Detection unit) The detection unit is configured to detect whether or not power is being supplied from the power supply unit to the audio processing unit. If the power supply unit is functioning normally, the detection unit supplies detection information indicating "Yes" to the control unit. On the other hand, if the power supply unit is not functioning normally, the detection unit supplies detection information indicating "No" to the control unit. If the power supply unit is an AC / DC converter, an example of a state in which the power supply unit is not functioning normally is a power outage. If the power supply unit is a battery, an example of a state in which the power supply unit is not functioning normally is a dead battery.

[0050] (Selection section) The selection unit selects whether or not to apply audio processing to the audio signal x based on the control of the control unit. The selection unit is provided, for example, between the low-pass filter 132 and the synthesis unit 14, and is configured by a switch that switches the circuit between ON / OFF. When the selection unit is ON, the first audio signal x1 that has passed through the high-pass filter 131 and the low-pass filter 132 is supplied to the synthesis unit 14, and is synthesized with the second audio signal x2 in the synthesis unit 14. On the other hand, when the selection unit is OFF, the first audio signal x1 is not supplied to the synthesis unit 14, and the second audio signal x2 is output from the synthesis unit 14. Note that the selection unit may be provided between the audio signal generation unit 11 and the high-pass filter 131, or between the high-pass filter 131 and the low-pass filter 132.

[0051] (Control unit) The control unit controls the selection unit to apply or not apply audio processing to the audio signal x according to the detection result of the detection unit. When the detection result of the detection unit is "Yes", the control unit controls the selection unit to apply audio processing to the audio signal x (turn ON). On the other hand, when the detection result of the detection unit is "No", the control unit controls the selection unit to not apply audio processing to the audio signal x (turn OFF).

[0052] [Variations of the voice processing device] In the audio processing device 1 according to the above embodiment, the high-pass filter 131, the low-pass filter 132, the correction signal generating unit 12, and the synthesis unit 14 are configured with analog circuits. However, at least any of these may be configured with digital circuits.

[0053] Moreover, the voice processing device 1 according to the above embodiment includes the high-pass filter 131 and the low-pass filter 132. However, instead of including these filters, the voice processing device 1 may include a band-pass filter that simultaneously removes the first frequency component and the second frequency component.

[0054] Moreover, the audio processing device 1 may further include a polarity switch and an amplifier interposed between the first port 15 and the audio signal generating unit 11.

[0055] The voice processing device 1 further includes a selection unit, an audible selector, and an audio input unit, which are disposed between the synthesis unit 14 and the second port 18. The amplifier may further include a circuit isolator, a transmitter power regulator, and a polarity switch.

[0056] The audio processing device 1 may also include an attenuator between the synthesis unit 14 and the second port 18. The attenuator arbitrarily adjusts the output in accordance with the input audio source signal. By including such an attenuator, it becomes possible to provide a system that can accommodate hearing loss symptoms due to hearing impairment by incorporating the microphone line of an existing wired or wireless audio output device, or the signal line of a television, radio, or other entertainment playback device.

[0057] [Effects of the audio processing device] In the voice processing device 1 described above, the low-pass filter 132 corrects the phase delay that occurs when the first and second frequency components are removed based on the phase correction signal p. Therefore, the voice processing device 1 can make the voice clearer than ever before.

[0058] <Applications of voice processing devices> Next, there will be described various devices incorporating the above-mentioned voice processing device 1. The voice processing device 1 can be incorporated into an auxiliary device 2, voice input / output systems 3 to 6, etc., which will be described below.

[0059] [Configuration of auxiliary equipment] As shown in FIG. 11, the auxiliary device 2 includes a housing 21, an input terminal 22, and an output terminal 23 in addition to the above-mentioned voice processing device 1.

[0060] The housing 21 is formed in a box shape capable of housing the voice processing device 1. The housing 21 is freely portable.

[0061] The input terminal 22 is provided on the surface of the housing 21. The input terminal 22 is connected to a device that outputs an audio signal, and receives an audio signal from the device. The input terminal 22 is connected to the first port 15 of the audio processing device 1. In other words, the input terminal 22 supplies the audio signal input from the device that outputs an audio signal to the first port 15.

[0062] The output terminal 23 is provided on the surface of the housing 21. The output terminal 23 is connected to a device to which the synthetic voice signal x' is input, and outputs a voice signal to the device. The output terminal 23 is connected to the second port 18 of the voice processing device 1. In other words, the output terminal 23 outputs the synthetic voice signal x' supplied from the voice processing device 1 to the device to which the synthetic voice signal x' is input.

[0063] [Effects of auxiliary devices] The auxiliary device 2 described above is not only easy to install and remove, but also easy to carry and store, making it easily applicable to broadcasting equipment and other audio output devices at outdoor event venues.

[0064] [Configuration of voice input / output system (1)] 12, in addition to the above-mentioned audio processing device 1, the audio input / output system 3 includes an audio input unit 31 and an audio output unit 32. The audio input / output system 3 according to this embodiment further includes an amplifier unit 33.

[0065] (Voice input unit) The voice input unit 31 acquires the voice of the speaker and generates a voice signal. The voice input unit 31 according to the present embodiment is composed of a microphone. Further, the voice input unit 31 is connected to the first port of the voice processing device 1. Then, the voice input unit 31 generates a voice signal x based on the input voice and supplies it to the first port 15. Note that the voice input unit 31 may be a device that plays back video such as a TV as shown in FIG. 13, for example.

[0066] (Voice output unit) The voice output unit 32 outputs a voice based on the synthesized voice signal x' generated by the voice processing device 1. The voice input unit 31 according to the present embodiment is composed of a speaker, headphones, earphones, etc. Further, the voice output unit 32 is connected to the second port 18 of the voice processing device 1. Then, the voice output unit 32 emits a voice based on the synthesized voice signal supplied from the second port.

[0067] [Operating effects of the voice input / output system (1)] In the voice input / output system 3 described above, components other than the voice processing device 1 can be configured with existing broadcast equipment, voice transmission devices, etc. Therefore, according to the voice input / output system 3, by incorporating it into existing broadcast equipment, voice transmission devices, etc., the voices emitted by these can be easily clarified.

[0068] [Configuration of the voice input / output system (2)] As shown in FIG. 14, the voice input / output system 4 includes, in addition to the same voice processing device 1, voice input unit 31, voice output unit 32, and amplification unit 33 as the voice input / output system 3 described above, a transmitter 41 and a receiver 42.

[0069] The transmitter 41 according to this embodiment is configured as a wireless communication module. The transmitter 41 is connected to the second port 18 of the voice processing device 1. The transmitter 41 converts the synthesized voice signal supplied from the second port into a radio wave. The transmitter 41 then transmits the radio wave from the antenna 41a.

[0070] The receiver 42 according to this embodiment is configured with a wireless communication module. The receiver 42 is connected to the amplifier 33. The receiver 42 receives radio waves from the transmitter 41 via an antenna 42a. The receiver 42 converts the received radio waves into a synthetic voice signal and supplies the synthetic voice signal to the amplifier 33.

[0071] 15, the voice input / output system 4 may be configured such that the voice input unit 31 and the voice output unit 32 are configured as a handset of a telephone, and the transmitter 41 and the receiver 42 are configured as the main body of the telephone. In this case, the transmitter 41 transmits a synthesized voice signal to the receiver 42 via a fixed telephone line, an Internet line, or the like.

[0072] [Effects of voice input / output systems (2)] According to the voice input / output system 4 described above, clear voice can be transmitted to the listener even if the speaker and the listener are far apart.

[0073] [Configuration of voice input / output system (3)] The audio input / output system 5 constitutes a stereo device. As shown in Fig. 16, the audio input / output system 5 includes a pair of input buffer units 16, a pair of audio signal generators 11, a correction signal generator 12, a filter 13, an adjustment unit 17, and a pair of synthesizers 14 similar to those of the audio processing device 1, and a pair of audio output units 32 similar to those of the audio input / output system 3, as well as a pair of audio input terminals 51, a second synthesizer 52, a second audio signal generator 53, and a third audio signal generator 54.

[0074] The pair of audio input terminals 51 are each connected to the input buffer unit 16. A right audio signal x(R) and a left audio signal x(L) are respectively input to the pair of audio input terminals 51. The input right audio signal x(R) and left audio signal x(L) are each supplied to the input buffer unit 16.

[0075] The audio signal generating unit 11 in the audio input / output system 5 does not generate the third audio signal x3, but generates the first audio signals x1(R), x1(L) and the second audio signals x2(R), x2(L).

[0076] The second synthesis unit 52 additively synthesizes the right first audio signal x1(R) and the left first audio signal x1(L) generated by the pair of audio signal generation units 11. The synthesized first audio signal x1 is supplied to the second audio signal generation unit 53.

[0077] The second audio signal generating unit 53 generates, from the first audio signal x1, a third audio signal having the same spectrum as the first audio signal.

[0078] The third audio signal generating unit 54 generates a right first audio signal x1'(R) and a left first audio signal x1'(L) from the phase-corrected first audio signal x1' supplied from the adjustment unit 17. The first audio signals x1'(R), x1'(L) are supplied to a pair of synthesis units 14, respectively.

[0079] [Effects of voice input / output systems (3)] According to the audio input / output system 5 described above, stereo audio (right audio signal and left audio signal) can be made clearer than ever before.

[0080] [Configuration of voice input / output system (4)] The audio input / output system 6 constitutes a so-called remote conference system. As shown in Fig. 17, the audio input / output system 6 includes a pair of audio processing devices 1, a pair of audio input units 31, and a pair of audio output units 32 similar to those of the audio input / output system 3, as well as a pair of dedicated terminals 61 for the remote conference system. The audio input / output system 6 may also include a camera and a monitor (not shown).

[0081] The voice input unit 31 is connected to a first port of the voice processing device 1. The voice input unit 31 generates a voice signal x based on the input voice and supplies it to the first port 15. The second port 18 of the voice processing device 1 is connected to a dedicated terminal 61. The voice processing device 1 generates a synthetic voice signal from the voice signal x supplied from the voice input unit 31 and supplies it to the dedicated terminal 61.

[0082] The dedicated terminal 61 is connected to another dedicated terminal 61 via an Internet line. Then, one dedicated terminal 61 supplies the synthetic voice signal supplied from the voice processing device 1 to the other dedicated terminal. Also, one dedicated terminal 61 supplies the synthetic voice signal supplied from the other dedicated terminal 61 to the voice output unit 32.

[0083] The voice output unit 32 is connected to the dedicated terminal 61. Then, the voice output unit 32 outputs a voice based on the synthesized voice signal supplied from the dedicated terminal 61.

[0084] [Effects of voice input / output systems (4)] According to the voice input / output system 6, the voice uttered by the speaker at one (the other) dedicated terminal can be clearly transmitted to the listener at the other (the other) dedicated terminal. As a result, the person at one dedicated terminal and the person at the other dedicated terminal can smoothly converse (discuss).

[0085] <Audio processing method> Next, a voice processing method according to an embodiment of another aspect of the present invention will be described.

[0086] As shown in FIG. 18, the voice processing method includes a voice signal generating step S1, a correction signal generating step S2, a filtering step S3, and a voice synthesis step S4.

[0087] (Audio signal generation step) In an initial voice signal generation step S1, a voice signal generator generates a first voice signal x1 and a second voice signal x2 having equal spectra from a voice signal representing a voice. Note that in the voice signal generation step S1, the voice signal generator may generate a third voice signal having the same phase as the second voice signal. Also, the voice signal generator may be one that constitutes the voice signal generation unit 11 of the voice processing device 1, one that constitutes another device, or one that is an independent device.

[0088] (Correction signal generation step) After generating the first audio signal x1 and the second audio signal x2, the process proceeds to a correction signal generation step S2. In the correction signal generation step S2, the correction signal generator generates a phase correction signal p based on the second audio signal x2. If a third audio signal is generated in the audio signal generation step S1, the correction signal generator may generate the phase correction signal based on the third audio signal in the correction signal generation step S2. The correction signal generator used in the correction signal generation step S2 may include a capacitor having one end connected to the audio signal generation unit and the other end connected to the non-inverting input terminal. The correction signal generator may be a component of the correction signal generation unit 12 of the audio processing device 1, a component of another device, or an independent device.

[0089] (Filtering step) After generating the first audio signal x1 and the second audio signal x2, a filtering step S3 is also performed. In the filtering step S3, a filter removes a first frequency component below a predetermined first frequency from the first audio signal x1 (S31). In addition, in the filtering step S3, a phase correction signal p is input to the non-inverting input terminal of the operational amplifier, causing the operational amplifier to output the first audio signal x1 whose phase has been corrected to approach the phase of the second audio signal x2 (S32). In addition, in the filtering step S3, a band-pass filter that also removes a second frequency component above a second frequency higher than the first frequency may be used as the filter. In addition, in the filtering step S3, a filter may be used that is composed of a high-pass filter that removes the first frequency component from the first audio signal and a low-pass filter that removes the second frequency component from the first audio signal from which the first frequency component has been removed. In this case, it is preferable to use an operational amplifier that is provided in the low-pass filter. In addition, in the filtering step S3, a high-pass filter including an operational amplifier for amplifying the first audio signal may be used. The filter may be a component of the audio signal generating unit 11 of the audio processing device 1, a component of another device, or an independent device. In addition, in the filtering step S3, the removal of the second frequency component may be performed prior to the removal of the first frequency component. Furthermore, the removal of the first frequency component and the removal of the second frequency component may be performed simultaneously.

[0090] (Speech synthesis step) After removing the second frequency component from the first voice signal x1, the process proceeds to voice synthesis step S4. In voice synthesis step S4, a voice synthesizer generates a synthetic voice signal x' by additively synthesizing the phase-corrected first voice signal x1 and the second voice signal x2. The voice synthesizer may be one that constitutes the voice signal generating unit 11 of the voice processing device 1, may be one that constitutes another device, or may be an independent device.

[0091] [Variations of the audio processing method] The voice processing method may further include an adjustment step. The adjustment step is preferably performed between the filtering step and the voice synthesis step. The adjustment step adjusts the sound pressure of the first voice signal whose phase has been corrected. When the adjustment step is included, in the voice synthesis step, the voice synthesizer adds and synthesizes the first voice signal whose sound pressure has been adjusted and the second voice signal.

[0092] [Effects of the audio processing method] In the above-described voice processing method, in the filtering step S3, the filter corrects the phase delay that occurs when the first and second frequency components are removed based on the phase correction signal p. Therefore, according to the voice processing method, the voice can be made clearer than ever before.

[0093] <Additional Notes> The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in the description of the invention are also included in the technical scope of the present invention.

[0094] <Summary> An audio processing device according to a first aspect of the present invention includes an audio signal generation unit that generates a first audio signal and a second audio signal having equal spectra from an audio signal representing audio, a filter that removes at least a first frequency component below a predetermined first frequency from the first audio signal, a synthesis unit that additively synthesizes the first audio signal from which the first frequency component has been removed and the second audio signal to generate a synthesized audio signal, and a correction signal generation unit that generates a phase correction signal based on the second audio signal, wherein the filter has an operational amplifier that amplifies the first audio signal, and the operational amplifier is configured to receive the first audio signal at an inverting input terminal, and to output the first audio signal whose phase has been corrected to approach the phase of the second audio signal when the phase correction signal is input to a non-inverting input terminal.

[0095] According to the above voice processing device, voice can be made clearer than ever before.

[0096] The audio processing device according to a second aspect of the present invention may be configured in the above-mentioned first aspect such that the filter is a band-pass filter that also removes a second frequency component equal to or greater than a second frequency that is higher than the first frequency.

[0097] According to the above configuration, the signal-to-noise ratio can be improved.

[0098] An audio processing device according to aspect 3 of the present invention may be configured in the above-mentioned aspect 2 such that the filter is composed of a high-pass filter that removes the first frequency component from the first audio signal, and a low-pass filter that removes the second frequency component from the first audio signal from which the first frequency component has been removed, and the operational amplifier is provided in the low-pass filter.

[0099] According to the above configuration, the signal is passed through the operational amplifier after all frequency components have been removed (after the phase is no longer delayed any further), so that the phase delay can be corrected more reliably.

[0100] An audio processing device according to aspect 4 of the present invention may be configured in the above-mentioned aspect 1 such that the audio signal generation unit generates a third audio signal having an equal spectrum to the second audio signal, and the correction signal generation unit generates the phase correction signal based on the third audio signal.

[0101] According to the above configuration, the correction signal generator generates the correction signal based on the dedicated third audio signal, which prevents the second audio signal from being out of phase when the second audio signal is used to generate the correction signal.

[0102] An audio processing device according to aspect 5 of the present invention may be configured in such a way that, in aspect 1 above, the correction signal generating unit has a capacitor having one end connected to the audio signal generating unit and the other end connected to the non-inverting input terminal.

[0103] With the above configuration, the higher the frequency of the signal passing through the filter, the greater the phase delay, while the capacitor increases the output as the frequency of the input signal increases, and the low-pass filter advances the phase of the output signal as the input to the non-inverting input terminal increases. Therefore, with a simple configuration, it is possible to configure a correction signal generator that is compatible with audio signals of various frequencies.

[0104] The audio processing device according to a sixth aspect of the present invention may be configured in the above-mentioned first aspect, wherein the first frequency is a frequency between an upper limit frequency of a first-order formant component and a lower limit frequency of a second-order formant component.

[0105] According to the above configuration, the first formant components are omitted from the first audio signal, and only the second formant components, the third formant components, etc. are left, so that only the second formant components and the third formant components can be amplified.

[0106] An audio processing device according to aspect 7 of the present invention may be configured in such a way that, in aspect 6 above, the second frequency is a frequency between the upper limit frequency of the 5th formant component and the lower limit frequency of the 6th formant component, or a frequency between the upper limit frequency of the 6th formant component and the lower limit frequency of the 7th formant component.

[0107] According to the above configuration, the signal-to-noise ratio can be further improved.

[0108] An audio processing device according to aspect 8 of the present invention may be configured in such a way that, in aspect 3 above, the high-pass filter includes an operational amplifier that amplifies the first audio signal, and the first frequency component is removed and the amplified first audio signal is output.

[0109] According to the above configuration, the clarity of the voice can be improved.

[0110] An audio processing device according to aspect 9 of the present invention may be configured in accordance with aspect 1 above, further comprising an adjustment unit that adjusts the sound pressure of the first audio signal whose phase has been corrected, and the synthesis unit additively synthesizes the first audio signal whose sound pressure has been adjusted and the second audio signal.

[0111] According to the above configuration, it is possible to output a sound according to the usage environment and the degree of hearing loss.

[0112] The voice input / output system of aspect 10 of the present invention may be configured in any one of aspects 1 to 6 above, comprising a voice input unit that acquires a speaker's voice and generates a voice signal, the voice processing device, and a voice output unit that outputs voice based on the synthetic voice signal generated by the voice processing device.

[0113] According to the above configuration, the sound can be made clearer than ever before.

[0114] An audio processing method according to aspect 11 of the present invention includes an audio signal generation step in which an audio signal generator generates a first audio signal and a second audio signal having equal spectra from an audio signal representing audio; a filtering step in which a filter removes at least a first frequency component below a predetermined first frequency from the first audio signal; a audio synthesis step in which a audio synthesizer additively synthesizes the first audio signal from which the first frequency component has been removed and the second audio signal to generate a synthetic audio signal; and a correction signal generation step in which a correction signal generator generates a phase correction signal based on the second audio signal, wherein the filter has an operational amplifier that amplifies the first audio signal, and the operational amplifier is configured to receive the first audio signal at its inverting input terminal, and in the filtering step, the phase correction signal is input to the non-inverting input terminal of the operational amplifier, causing the operational amplifier to output the first audio signal whose phase has been corrected so as to approach the phase of the second audio signal.

[0115] According to the above voice processing method, voice can be made clearer than ever before. [Explanation of symbols]

[0116] 1 Audio processing device 11 Audio signal generator 12 Correction signal generator 121 Capacitor 122 resistor 13 Filters 131 High Pass Filter 132 Low-pass filter 132a Operational Amplifier 132b, 132c elements 14 Synthesis section 15 First Port 16 Input Buffer 17 Adjustment section 18 Second Port 2 Auxiliary equipment 21 Case 22 Input terminal 23 Output terminal 3, 4, 5, 6 Audio Input / Output System 31 Audio input section 32 Audio output section 33 Amplification section 41 Transmitter 41a, 42a Antenna 42 Receiver 51 Audio input terminal 52 2nd synthesis section 53 Second audio signal generator 54 Third audio signal generator 61 Dedicated terminal p phase correction signal t Synthesized speech signal x Audio Signal x1, x1' 1st audio signal x2 Second audio signal x3 3rd audio signal

Claims

1. a speech signal generating unit that generates a first speech signal and a second speech signal having the same spectrum from a speech signal representing speech; a filter that removes at least a first frequency component that is equal to or lower than a predetermined first frequency from the first audio signal; a synthesis unit that generates a synthetic audio signal by additively synthesizing the first audio signal from which the first frequency component has been removed and the second audio signal; a correction signal generation unit that generates a phase correction signal based on the second audio signal; Equipped with the filter includes an operational amplifier for amplifying the first audio signal; The operational amplifier comprises: The first audio signal having a phase delay caused by removing the first frequency component is input to an inverting input terminal, When the phase correction signal is input to a non-inverting input terminal, the first audio signal having a phase corrected to approach a phase of the second audio signal is output. Audio processing device.

2. The filter is a band-pass filter that also removes a second frequency component equal to or greater than a second frequency that is higher than the first frequency. The audio processing device according to claim 1 .

3. The filter comprises: a high-pass filter that removes the first frequency component from the first audio signal; a low-pass filter that removes the second frequency component from the first audio signal from which the first frequency component has been removed; It is composed of The operational amplifier is provided in the low-pass filter. The audio processing device according to claim 2 .

4. the audio signal generation unit generates a third audio signal having an equal spectrum to the second audio signal; The correction signal generation unit generates the phase correction signal based on the third audio signal. The audio processing device according to claim 1 .

5. the correction signal generating unit includes a capacitor having one end connected to the audio signal generating unit and the other end connected to the non-inverting input terminal; The audio processing device according to claim 1 .

6. The first frequency is a frequency between an upper limit frequency of a first formant component and a lower limit frequency of a second formant component. The audio processing device according to claim 1 .

7. the second frequency is a frequency between an upper limit frequency of a fifth formant component and a lower limit frequency of a sixth formant component, or a frequency between an upper limit frequency of a sixth formant component and a lower limit frequency of a seventh formant component; The audio processing device according to claim 2 .

8. The high pass filter is an operational amplifier for amplifying the first audio signal; outputting the first audio signal from which the first frequency component has been removed and from which the first audio signal has been amplified; The audio processing device according to claim 3 .

9. an adjustment unit that adjusts a sound pressure of the first audio signal whose phase has been corrected, The synthesis unit additively synthesizes the first audio signal, the sound pressure of which has been adjusted, and the second audio signal. The audio processing device according to claim 1 .

10. a voice input unit that acquires a speaker's voice and generates a voice signal; A voice processing device according to any one of claims 1 to 6, a voice output unit that outputs a voice based on the synthetic voice signal generated by the voice processing device; Equipped with Audio input / output system.

11. a speech signal generating step of generating a first speech signal and a second speech signal having equal spectra from a speech signal representing speech by a speech signal generator; a filtering step of removing at least a first frequency component below a predetermined first frequency from the first audio signal using a filter; a voice synthesis step in which a voice synthesizer generates a synthetic voice signal by additively synthesizing the first voice signal from which the first frequency component has been removed and the second voice signal; a correction signal generating step of generating a phase correction signal based on the second audio signal by a correction signal generator; Including, the filter includes an operational amplifier for amplifying the first audio signal; The operational amplifier is configured to receive, at an inverting input terminal, the first audio signal having a phase delay caused by removing the first frequency component; In the filtering step, the phase correction signal is input to a non-inverting input terminal of the operational amplifier, thereby causing the operational amplifier to output the first audio signal whose phase has been corrected to approach the phase of the second audio signal. Audio processing methods.

Citation Information

Patent Citations

  • Voice processor, voice clearing device, and voice processing method

    JP2016110050A