Audio Processing Device

A portable audio processing device enhances communication clarity for hearing-impaired individuals by processing audio signals and being wearable, addressing the limitations of non-portable devices in personal communication scenarios.

JP7719239B2Active Publication Date: 2025-08-05RADIUS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024086498
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-06-02
Filing Date
2024-05-28
Publication Date
2025-08-05
Estimated Expiration
2044-05-28

AI Technical Summary

Technical Problem

Existing audio processing devices are not portable and cannot be used in a mobile state with speakers, limiting their effectiveness in personal communication scenarios, especially for individuals with hearing impairments.

Method used

A portable audio processing device that includes an audio processor, a speaker, and a housing, which can be worn or carried by a user, processing audio signals to enhance clarity and facilitate communication with hearing-impaired individuals.

Benefits of technology

Enables effective communication with hearing-impaired individuals in various situations by clarifying speech through portable audio processing, reducing misunderstandings and improving conversation clarity without the need for additional assistive devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007719239000001
    Figure 0007719239000001
  • Figure 0007719239000002
    Figure 0007719239000002
  • Figure 0007719239000003
    Figure 0007719239000003
Patent Text Reader

Abstract

To provide a voice processing device for enabling a hearing-impaired person to easily hear voice without using an auxiliary device by allowing an utterer to use the auxiliary device.SOLUTION: A portable voice processing device (1) includes: a voice processing unit (10) for executing voice processing to clarify voice with respect to a voice signal supplied from a microphone (21) connected to a support member (24) by which the microphone is positioned close to the mouth of a user; a speaker (22) for converting the voice signal after the execution of the voice processing into voice; and a housing for storing the voice processing unit (10).SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an audio processing device. [Background technology]

[0002] BACKGROUND ART The proportion of the population suffering from hearing loss has been increasing due to the aging population, and assistive devices such as hearing aids and sound amplifiers are now commonly used. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2016-110050 Summary of the Invention [Problem to be solved by the invention]

[0004] Figure 5 of Patent Document 1 discloses an audio processing device for making audio coming from speakers of audio equipment, broadcasting facilities, etc. clear and easy to hear. This audio processing device can be read as the audio processor in this specification. Therefore, hereinafter, the audio processing device disclosed in Patent Document 1 will be referred to as the audio processor.

[0005] However, the voice processing device described in Patent Document 1 is primarily targeted at audio equipment and broadcasting facilities installed in places where large numbers of people gather, such as buildings, public buildings, shopping centers, and sports facilities. Therefore, this voice processing device emits processed voice through speakers pre-installed in places where large numbers of people gather. Therefore, this voice processing device is not portable, and therefore cannot be used in a state where the speaker and the voice processing device can be moved together.

[0006] One aspect of the present invention has been made in consideration of the above-mentioned problems, and an object of the present invention is to provide a voice processing device that can be used in a mobile state together with a speaker. [Means for solving the problem]

[0007] In order to solve the above problems, an information processing device according to a first aspect of the present invention is a portable audio processing device, which includes an audio processor that performs audio processing on an audio signal supplied from a microphone to clarify the audio, a speaker that converts the processed audio signal into audio, and a housing that houses at least the audio processor.

[0008] According to the above configuration, the voice processing device is configured to be portable, and therefore it is possible to provide a voice processing device that can be used in a mobile state together with the speaker.

[0009] Because the speaker and the voice processing device can be moved together, the voice processing device can be used in a variety of situations, such as when a store clerk interacts with customers, when an event venue attendant guides spectators, when a local government official provides information to citizens at a disaster site, etc. Therefore, even if a hearing-impaired person is included among the above-mentioned customers, spectators, citizens, etc., it becomes easy for the speaker to convey the message he or she wants to convey to the hearing-impaired person.

[0010] In these cases, the person to whom the speaker speaks using the processed voice may be a specific person or people facing the speaker, or may be an unspecified person or people such as passersby.

[0011] In addition, in the audio processing device according to the second aspect of the present invention, in addition to the configuration of the audio processing device according to the first aspect described above, a configuration is adopted in which the housing accommodates the speaker in addition to the audio processor.

[0012] According to the above configuration, the voice processor and the speaker are housed integrally in a single housing, making it easy for the speaker to carry the voice processor.

[0013] In addition, in the audio processing device according to the third aspect of the present invention, in addition to the configuration of the audio processing device according to the second aspect described above, a microphone that converts audio into an audio signal is further provided, and the housing is configured to accommodate the microphone in addition to the audio processor and the speaker.

[0014] According to the above configuration, the voice processor, the speaker, and the microphone are housed integrally in a single housing, making it easy for the speaker to carry the voice processor. An example of such a voice processor is a device called an electronic megaphone or loudspeaker.

[0015] In addition, in the audio processing device according to the fourth aspect of the present invention, in addition to the configuration of the audio processing device according to the first or second aspect described above, a configuration is adopted in which the device further includes a microphone that converts audio into an audio signal.

[0016] The information processing device according to one aspect of the present invention can also be sold as a package including a microphone. [Effects of the Invention]

[0017] According to one aspect of the present invention, it is possible to provide a voice processing device that can be used in a mobile state together with a speaker. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a block diagram of a voice processing device according to a first embodiment of the present invention. [Figure 2] The left figure is an example of an external view of the voice processing device shown in FIG. 1, and the right figure is a schematic diagram of a user wearing the voice processing device shown in FIG. [Figure 3] 2 is a graph showing frequency characteristics in the audio processing device shown in FIG. [Figure 4] FIG. 10 is a block diagram showing an example of the configuration of a voice processing device according to a second embodiment of the present invention. [Figure 5] FIG. 2 is a diagram illustrating formant components contained in a speech signal. [Figure 6] 5 is a graph showing adjustment characteristics of an input buffer unit included in the audio processing device of FIG. 4. [Figure 7] 5 is a circuit diagram showing an example of a specific configuration of a filter and a correction signal generating unit included in the audio processing device of FIG. 4. [Figure 8] 5 is a graph showing the phase correction characteristics of a filter included in the audio processing device of FIG. 4. [Figure 9] 10 is a graph illustrating a case where a first audio signal and a second audio signal are combined without performing phase correction. [Figure 10] 10 is a graph showing the frequency characteristics of a synthesized speech signal when phase correction is not performed. [Figure 11] 5 is a graph showing an example of adjustment of an audio signal by an adjustment unit included in the audio processing device of FIG. 4. [Figure 12] 10 is a graph showing changes over time in phase and output of an audio signal, a first audio signal before and after correction, and a synthesized audio signal. [Figure 13] 10 is a graph showing the overall output characteristics of a synthesis unit. DETAILED DESCRIPTION OF THE INVENTION

[0019] [Embodiment 1] Hereinafter, one embodiment of the present invention will be described in detail.

[0020] <Outline of voice processing device> In recent years, with the aging of the population, it has been reported that there are frequent cases in which important information spoken by a speaker cannot be conveyed to a listener with hearing loss.The general idea is to simply increase the speaker's volume, but there are two types of hearing loss: conductive hearing loss and sensorineural hearing loss caused by damage to the inner ear.Sensorineural hearing loss in particular makes it impossible to understand language even at high volumes, and it has been elucidated that this is caused by a decline in the function of the cochlea, the organ of hearing, which converts high frequencies into auditory signals.

[0021] In situations where hearing function has declined due to aging or other factors, communication through conversation in daily life becomes less common, and so assistive devices such as hearing aids and sound amplifiers are commonly used by people with hearing loss. These assistive devices increase the volume and / or perform some frequency processing to make the sound picked up by the microphone easier to hear, and then output it from earphones.

[0022] However, in places such as nursing homes and care facilities where there are a large number of hearing-impaired people, it is inefficient and expensive to have all of the hearing-impaired people wear an assistive device.

[0023] Therefore, the inventors have sought a method for enabling a speaker to use an assistive device to make it easier for a hearing-impaired person to hear speech without using an assistive device.

[0024] A voice processing device 1 according to one embodiment of the present invention will be described below with reference to Figs. 1 to 3. Fig. 1 is a block diagram of the voice processing device 1. The left diagram of Fig. 2 is an example of an external view of the voice processing device 1, and the right diagram of Fig. 2 is a schematic diagram of a user wearing the voice processing device 1. Fig. 3 is a graph showing frequency characteristics of the voice processing device 1.

[0025] <Concept of voice processing device> The voice processing device 1, which is a type of assistive device, is equipped with technology that enables the accurate understanding of language that is difficult to hear by auxiliary correcting frequency components that play an important role in understanding language, and achieves this using a method that is completely opposite to the conventional idea of correcting and processing voice signals on the speaker's side.

[0026] The speech processing device 1 has a function that enables the speaker to arbitrarily set the strength of signal processing components based on the listener's language comprehension ability, thereby enabling a comfortable conversation to be realized.

[0027] <Configuration of voice processing device> 1, the audio processing device 1 includes a microphone 21, an audio processor 10, a speaker 22, and a housing. In this embodiment, the housing houses at least the audio processor 10.

[0028] The voice processing device 1 is configured to be portable. The voice processing device 1 may be configured to be wearable on a part of the speaker's body or clothing, or may be configured to have a holder that can be held by the speaker.

[0029] In this embodiment, the voice processing device 1 has a clip provided on the back surface of the housing (the surface facing the speaker's body), although this is not shown in the left diagram of FIG. 2. The clip is an example of a configuration for attaching the voice processing device 1 to pants, a skirt, or a belt worn by the speaker. Therefore, the voice processing device 1 is configured so that the speaker can carry it while wearing it around their waist (see the right diagram of FIG. 2). Note that the configuration that can be attached to a part of the speaker's body or clothing is not limited to the clip described above, and may be, for example, a belt or a strap. The voice processing device 1 may also be a wristwatch type. A wristwatch type voice processing device 1 is also an example of a wearable device.

[0030] An example of the voice processing device 1 provided with a holding part that can be held by a speaker is an electronic megaphone equipped with the voice processor 10. The grip part of the electronic megaphone that is held by a speaker is an example of a holding part that can be held by a speaker.

[0031] The voice processing device 1 is configured to weigh 10 kg or less so that it can be carried by an average adult. The lighter the weight of the voice processing device 1, the better. Furthermore, if the voice processing device 1 is configured to be wearable on the speaker's body or clothing, the weight of the voice processing device 1 is preferably 1 kg or less.

[0032] (microphone) The microphone 21 converts the voice uttered by the speaker into a voice signal SV and supplies the voice signal SV to the port P11 of the voice processor 10 to which the microphone 21 is connected. In this embodiment, the voice signal SV is an analog signal and an electrical signal. However, the microphone 21 may be configured to convert the voice uttered by the user on the transmitting side into the voice signal SV, which is a digital signal.

[0033] For example, an electric condenser microphone can be used as the microphone 21. A DC voltage is applied to the electric condenser microphone for driving it. Furthermore, electric condenser microphones are often connected using connectors of different types depending on the manufacturer or model.

[0034] (audio processor) An audio processor according to an aspect of the present invention is configured to perform audio processing on an audio signal supplied from a microphone 21 to clarify the audio, and to output the audio signal that has undergone the audio processing. The following describes audio processor 10, which is an example of an audio processor according to an aspect of the present invention.

[0035] As shown in Fig. 1, the audio processor 10 includes a port P11, a port P12, a branching unit 11, a phase corrector 12, a filter (band-pass filter) 13, an amplifier 14, an adjuster 15, and an additive synthesizer 16. The audio processor 10 is disposed between a microphone 21 and a speaker. The branching unit 11, the filter 13, the amplifier 14, and the additive synthesizer 16 are examples of an audio processing unit. The audio processor 10 also includes a power supply unit, a detection unit, and a control unit, which are not shown in Fig. 1.

[0036] Port P11 is a port connected to microphone 21 and is an example of a first port. Port P11 and microphone 21 are connected using a cable or wiring. A voice signal SV converted from the voice of a speaker by microphone 21 is supplied from microphone 21 to a terminal of port P11.

[0037] As shown in FIG. 1 of Patent Document 1, the speech signal SV includes first formant components, second formant components, ..., and nth formant components. Note that n depends on the user on the transmitting side and, although there may be some individual differences, is a positive integer of at least 4. Note that the first formant components, second formant components, ..., and nth formant components can be read as the first formant, second formant, ..., and nth formant shown in FIG. 1 of Patent Document 1, respectively.

[0038] It should be noted that the first formant component, second formant component, . . . , nth formant component are described in Patent Document 1, and therefore, description thereof will be omitted in this embodiment.

[0039] The branching unit 11 branches the audio signal SV into a first audio signal SV1 and a second audio signal SV2. In this embodiment, the branching unit 11 is configured so that the intensity ratio between the first audio signal SV1 and the second audio signal SV2 is 1:1, i.e., so that the distribution ratio is 1:1. However, the distribution ratio of the branching unit 11 is not limited to 1:1 and can be set appropriately.

[0040] The spectrum of each of the first audio signal SV1 and the second audio signal SV2 is similar to the spectrum of the audio signal SV, i.e., each of the first audio signal SV1 and the second audio signal SV2 includes a first formant component, a second formant component, ..., an nth formant component.

[0041] The filter 13 is a high-pass filter that removes low-frequency components, which are components with frequencies less than 400 Hz, from the first audio signal SV1 that has passed through the branching unit 11. The filter 13 outputs a first audio signal SV1' that does not include the low-frequency components. Therefore, the first audio signal SV1' is an audio signal that does not include a first-order formant component, but does include second-order formant components, third-order formant components, ..., n-th order formant components.

[0042] The amplifier 14 amplifies the first audio signal SV1' that has passed through the filter 13, and outputs a first audio signal SV1'' in which the intensities of the second-order formant components, third-order formant components, ..., n-th order formant components have been increased.

[0043] The adjuster 15 adjusts the gain of the amplifier 14. In this embodiment, the adjuster 15 is configured using a switch that can adjust the gain in three stages from 0 to 2. The adjuster 15 is configured so that when "0" is selected by the switch, the gain becomes 1x, when "1" is selected by the switch, the gain becomes 3x, and when "2" is selected by the switch, the gain becomes 5x.

[0044] However, the number of switch stages constituting the adjuster 15 and the gain selected by the switch are not limited to the above example and can be selected as appropriate. For example, the adjuster 15 may be configured so that the gain is 0 when “0” is selected by the switch, 1 when “1” is selected by the switch, and 5 when “2” is selected by the switch. When the gain is set to 0, the voice processing in the voice processor 10 is essentially canceled. Therefore, the voice processing device 1 converts the voice signal supplied from the microphone 21 into voice and outputs it from the speaker without performing voice processing for voice clarity. If it is known that voice clarity processing is unnecessary, the voice processing may be canceled as described above. Alternatively, the voice processing may be normally left canceled, and if a gesture that seems difficult for the listener to hear is recognized, the speaker can adjust the degree of voice processing in the voice processor 10 by appropriately adjusting the adjuster 15. Alternatively, the adjuster 15 may be a volume that continuously changes the gain instead of a switch that discretely changes the gain. Furthermore, if the users who will be listeners can be identified to some extent, the gain is set in advance and adjuster 15 can be omitted.

[0045] The phase corrector 12 corrects the phase of the second audio signal SV2 so that it matches the phase of the first audio signal SV1" that has passed through the amplifier 14. In other words, the phase corrector 12 outputs the second audio signal SV2' that is in phase with the first audio signal SV1".

[0046] The additive synthesizer 16 generates a synthesized audio signal SSV by additively synthesizing the first audio signal SV1″ that has passed through the amplifier 14 and the second audio signal SV2′ that has passed through the phase corrector 12, and supplies the synthesized audio signal SSV to a port P12.

[0047] Port P12 is a port connected to a speaker and is an example of a second port. Port P12 and the speaker are connected using a cable or wiring. A synthesized voice signal SSV is supplied from additive synthesizer 16 to a terminal of port P12.

[0048] In this way, the audio processor 10 performs audio processing on the audio signal SV supplied from the microphone 21 and supplies the synthesized audio signal SSV to the speaker 22.

[0049] In this embodiment, the microphone 21 supplies an analog audio signal SV to the port P11. The port P12 supplies an analog synthesized audio signal SSV to the speaker. However, the microphone 21 may be configured to supply a digital audio signal SV to the port P11, and the port P12 may be configured to supply a digital synthesized audio signal SSV to the speaker.

[0050] In this embodiment, the phase corrector 12, the filter 13, the amplifier 14, and the additive combiner 16 are each configured with an analog circuit. However, the phase corrector 12, the filter 13, the amplifier 14, and the additive combiner 16 may each be configured with a digital circuit.

[0051] Furthermore, in this embodiment, the audio processor 10 includes a filter 13, which is a high-pass filter, that performs filtering on the first audio signal SV1. However, the audio processor 10 may also be configured such that the filter 13 removes (1) low-frequency components, which are components with a frequency less than 400 Hz, and (2) high-frequency components, which are components with a frequency greater than 7 kHz, from the first audio signal SV1 that has passed through the branching unit 11. That is, the filter 13 may be a band-pass filter that passes components with a frequency greater than or equal to 400 Hz and less than or equal to 7 kHz, or may be a band-pass filter that passes components with a frequency greater than or equal to 400 Hz and less than or equal to 5 kHz. In this case, the filter 13 may be realized as a single band-pass filter, or may be realized by connecting a high-pass filter and a low-pass filter in series.

[0052] In addition, in this embodiment, the housing of the audio processor 10 or the audio processing device 1 is made of aluminum. However, the material constituting the housing is not limited to aluminum, and may be metal such as copper or stainless steel. Furthermore, the material constituting the housing may be primarily made of resin with a metal layer provided on the surface. By covering at least the surface of the housing with metal, it is possible to block electromagnetic waves that penetrate from the outside to the inside.

[0053] In one aspect of the present invention, a single housing that houses the audio processor 10 may be configured to house at least one of a microphone 21 and a speaker 22 in addition to the audio processor 10. In this embodiment, as shown in FIG. 2 , the housing is configured to house the audio processor 10 and the speaker 22. Note that an audio processing device 1 that houses the audio processor 10, the microphone 21, and the speaker 22 in a single housing and is portable can be considered a type of device called an electronic megaphone or loudspeaker. Furthermore, when a high-output speaker is used as the speaker 22 to output loud audio, the audio processor 10 and the speaker 22 may be housed in a single housing, and the microphone 21 may be an external device. However, even in such a case, the audio processing device 1 including the audio processor 10, the speaker 22, and the housing is limited in weight and size to be portable.

[0054] The power supply unit, detection unit, and control unit, which are not shown in FIG. 1, are each configured as follows.

[0055] The power supply unit supplies power to the audio processing unit (particularly the amplifier 14) and the control unit. In order to make the audio processing device 1 portable, the power supply unit is preferably a battery. By using a battery as the power supply unit, the speaker can speak while moving freely without worrying about the power cable. However, assuming that the speaker's range of movement is limited, an AC / DC converter, such as an AC adapter, may be used as the power supply unit. By using an AC / DC converter as the power supply unit, commercial power can be used, so the speaker can speak without worrying about the remaining power. Furthermore, the power supply unit may employ both a battery and an AC / DC converter.

[0056] The detection unit is configured to detect whether or not power is being supplied from the power supply unit to the audio processing unit. If the power supply unit is functioning normally, the detection unit supplies detection information indicating "Yes" to the control unit. On the other hand, if the power supply unit is not functioning normally, the detection unit supplies detection information indicating "No" to the control unit. If the power supply unit is an AC / DC converter, an example of a state in which the power supply unit is not functioning normally is a power outage. If the power supply unit is a battery, an example of a state in which the power supply unit is not functioning normally is a dead battery.

[0057] The control unit controls the selection unit to apply or not apply audio processing to the audio signal SV according to the detection result of the detection unit. If the detection result of the detection unit is "Yes", the control unit controls the selection unit to apply audio processing to the audio signal SV. On the other hand, if the detection result of the detection unit is "No", the control unit controls the selection unit not to apply audio processing to the audio signal SV.

[0058] The audio processor 10 may further include a polarity switch and an amplifier interposed between the port P11 and the branching unit 11.

[0059] (speaker) The speaker 22 converts the synthesized voice signal SSV supplied from the port P12 into voice and outputs the voice. In this embodiment, the voice signal supplied from the port P12 is an analog signal and an electrical signal. However, the voice signal supplied from the port P12 may be a digital signal, and the speaker 22 may be configured to convert the digital voice signal into voice. Alternatively, the speaker 22 may have a function to amplify voice, and the voice processor 10 may be realized as a loudspeaker.

[0060] (External view of the audio processing device) The left diagram in FIG. 2 is an example of an external view of the voice processing device 1, and the right diagram is a schematic diagram of a user wearing the voice processing device 1. As shown in FIG. 2, for example, the microphone 21 and the voice processor 10 may be connected by wire or wirelessly. In the example of FIG. 2, the microphone 21 is connected to a support member 24 for positioning the microphone 21 near the user's mouth. The support member 24 is worn along the top of the user's ear and the back of the head. In addition, the voice processor 10 may be fixed to a belt or the like and configured to be integrated with the microphone 21. According to the configuration of FIG. 2, the voice processing device 1 can be used hands-free, thereby improving user convenience.

[0061] <Frequency characteristics of audio processing devices> FIG. 3 is a graph showing the frequency characteristics of the audio processing device 1 equipped with the audio processor 10. Characteristic line a shown in FIG. 3 shows the frequency characteristics of the audio signal SV supplied from the microphone 21. Characteristic line b shown in FIG. 3 shows the frequency characteristics of the audio signal SV after its output value has been adjusted by the amplifier interposed between the port P11 and the branching section 11. Characteristic line c shown in FIG. 3 shows the frequency characteristics of the first audio signal SV1" after passing through the amplifier 14. Characteristic line d shown in FIG. 3 shows the frequency characteristics of the synthesized audio signal SSV. The difference between characteristic line b and characteristic line d, indicated by the arrows in FIG. 3, is the effect of the audio processing performed by the audio processor 10.

[0062] The output level of the frequency characteristic of the first audio signal SV1'' represented by the characteristic line c can be adjusted by adjusting the gain of the amplifier 14 using the adjuster 15.

[0063] <Effects of audio processing devices> The voice processing device 1 is a conversation assistance device equipped with a function that provides a method for reliably communicating the contents of a conversation with a hearing-impaired person in face-to-face conversations. This method was born from a completely opposite idea to conventional hearing aids and sound collectors, and the speaker processes the voice to make it easier for the hearing-impaired person to hear, and then the speaker emits the processed voice via the voice processing device 1, making it possible for the hearing-impaired person to communicate the contents of the conversation without using a hearing aid or the like.

[0064] The voice processing-based voice clarity technology installed in the voice processing device 1 focuses on the fact that a decline in the efficiency of the cochlea, the auditory organ in the inner ear that causes hearing impairment due to aging, in converting sound into electrical signals, particularly in the high-frequency band (called formants), impedes language comprehension, and incorporates the feature of restoring hearing by applying processing that emphasizes the high-frequency components contained in words. In voice spectrum analysis, the highest peak frequency that appears at the lowest frequency is called the first formant, the second highest peak frequency is called the second formant, etc., and each peak frequency appears at an integer multiple point, and although these peak frequencies vary depending on the speaker's skeletal structure, our many findings over many years have clarified that the formant components necessary for language comprehension generally need to be reliably heard from the first to fourth formants, and it has been confirmed that the voice processing device 1 can achieve this effect by focusing on compensation of the 400Hz to 5KHz range.

[0065] As described above, the audio processor 10 of the audio processing device 1 includes the port P11, the branching unit 11, the filter 13, the amplifier 14, the phase corrector 12, the additive combiner 16, and the port P12.

[0066] According to the above configuration, the amplifier 14 amplifies the first audio signal SV1' but does not amplify the second audio signal SV2. Therefore, the audio processor 10 can output a synthetic audio signal in which the first formant components are not amplified and predetermined formant components (e.g., second, third, and fourth formant components) are amplified. Therefore, the audio processing device 1 can output a synthetic audio signal SSV that the listener perceives as clear without using an auxiliary device.

[0067] As a result, the possibility of misunderstandings occurring between the speaker and the listener or of the listener being offended can be reduced, making it easier for communication to proceed smoothly between the speaker and the listener.

[0068] The inventors have found that, when clarifying a voice, it is preferable to remove (1) low-frequency components, which are components with a frequency of less than 400 Hz, and (2) high-frequency components, which are components with a frequency of more than 7 kHz, from the first voice signal SV1, then amplify the first voice signal SV1', and additively combine the first voice signal SV1" and the second voice signal SV2, which are in phase with each other. However, the wavelength of a human voice is usually between 100 Hz and 1 kHz. Therefore, in many cases, it is sufficient for the filter 13 in the voice processor 10 that constitutes the voice processing device 1 to be able to remove only the low-frequency components from the first voice signal SV1. In other words, the filter 13 does not need to be configured to remove the high-frequency components from the first voice signal SV1. This allows the filter 13 to be configured simply, thereby reducing costs.

[0069] In the audio processing device 1, each of the phase corrector 12, the filter 13, the amplifier 14, and the additive combiner 16 is preferably configured by an analog circuit.

[0070] According to the above configuration, the voice processor 10 can be easily configured and can be made smaller. Furthermore, compared to when the filter 13 is configured using a digital circuit, the voice processing device 1 configured in this manner can output a synthesized voice signal SSV that the listener perceives as clearer. This is thought to be because the filter 13 configured using an analog circuit is more likely to have gentler filtering characteristics in the lower limit region of the passband (i.e., the region around 400 Hz) than the filter 13 configured using a digital circuit.

[0071] In the audio processing device 1, the audio processor 10 preferably further includes an adjuster 15.

[0072] According to the above configuration, the speaker user can adjust the gain of the amplifier 14 depending on the state of the listener user, regardless of the listener user's intention. That is, the degree of clarity in the synthetic speech signal can be adjusted. Therefore, the speech processing device 1 configured in this manner can output a synthetic speech signal SSV with the degree of clarity adjusted depending on the listener user.

[0073] In the voice processing device 1, the filter 13 is preferably configured to remove the low-frequency components below a lower limit frequency and the high-frequency components above an upper limit frequency. The band between the lower limit frequency and the upper limit frequency may include second-, third-, and fourth-order formant components.

[0074] The inventors of the present application have found that the frequency band that contributes to the clarity of an audio signal is approximately between 400 Hz and 7 kHz. In the audio processing device 1 configured in this manner, the amplifier 14 amplifies the first audio signal SV1' that falls within the frequency band that contributes to the clarity of an audio signal.

[0075] The audio processor 10 included in the audio processing device 1 is also an aspect of the present invention.

[0076] The speech processor 10 can output a synthesized speech signal in which the first formant components are not amplified and predetermined formant components (for example, second, third, and fourth formant components) are amplified.

[0077] In the audio processor 10, the audio signal SV and the synthesized audio signal SSV are preferably analog signals.

[0078] In addition, in the audio processor 10, each of the phase corrector 12, the filter 13, the amplifier 14, and the additive combiner 16 is preferably configured by an analog circuit.

[0079] This allows for an easy configuration of the voice processor 10 and a reduction in the size of the voice processor. Furthermore, compared to a case where the filter 13 is configured using a digital circuit, the voice processor 10 configured in this manner can output a synthesized voice signal that the listener perceives as clearer.

[0080] [Embodiment 2] Other embodiments of the present invention will be described below. For ease of explanation, the same reference numerals will be used to designate components having the same functions as those described in the above embodiment, and the description thereof will not be repeated.

[0081] <Configuration of voice processing device> As shown in FIG. 4, the voice processing device 1A includes a microphone 21, a speaker 22, and a housing similar to those of the first embodiment, as well as a voice processor 10A.

[0082] [Configuration of voice processing device] The audio processor 10A includes an audio signal generator 11A, a correction signal generator 12A, a filter 13A, and a synthesizer 14A. The audio processor 10A according to this embodiment further includes a first port 15A, an input buffer 16A, an adjuster 17, and a second port 18. The audio processor 10A may further include a power supply, a detector, and a controller, all of which are not shown.

[0083] (Port 1) The first port 15A outputs an audio signal x representing audio supplied from a device connected to the first port 15A to the input buffer unit 16A. As shown in FIG. 5, the audio signal x includes a first-order formant component h1, a second-order formant component h2, ..., and an n-th order formant component hn. Note that n depends on the user on the transmitting side and is a positive integer of at least 4, although there may be some individual differences. Note that the first-order formant component h1, the second-order formant component h2, ..., and the n-th order formant component hn are described in Patent Document 1, and therefore will not be described in this embodiment.

[0084] (input buffer) As shown in FIG. 4, the input buffer unit 16A is provided between the first port 15A and the audio signal generation unit 11A. The input buffer unit 16A can be configured with a microphone amplifier, a preamplifier that lowers output impedance, or the like, depending on the usage situation. The input buffer unit 16A according to this embodiment adjusts the audio signal x by narrowing the upper and lower frequency limits and increasing the sound pressure, for example, as shown in FIG. 6. Then, as shown in FIG. 4, the input buffer unit 16A outputs the audio signal x that has been adjusted as required to the audio signal generation unit 11A. Note that if the audio signal supplied from the first port does not require adjustment (if the device connected to the first port is configured to output an audio signal that does not require adjustment), the audio processor 10A does not need to include the input buffer unit 16A.

[0085] (Audio signal generation section) The audio signal generation unit 11A generates a first audio signal x1 and a second audio signal x2 having the same spectrum from the audio signal x supplied from the input buffer unit 16A. "Having the same spectrum" means that the first audio signal x1 and the second audio signal x2 each contain a first-order formant component h1, a second-order formant component h2, ..., and an n-order formant component hn. The audio signal generation unit 11A according to this embodiment also generates a third audio signal x3 having the same spectrum as the second audio signal x2. The audio signal generation unit 11A then supplies the first audio signal x1 to the high-pass filter 131, the second audio signal x2 to the synthesis unit 14A, and the third audio signal to the correction signal generation unit 12A. The audio signal generation unit 11A according to this embodiment generates the first audio signal x1, the second audio signal x2, and the third audio signal x3 by branching the audio signal. In this embodiment, the audio signal generation unit 11A is configured so that the intensity ratio between the first audio signal x1, the second audio signal x2, and the third audio signal x3 is 1:1:1, i.e., so that the distribution ratio is 1:1:1. However, the distribution ratio of the audio signals is not limited to 1:1:1 and can be set appropriately. The audio signal generation unit 11A may also duplicate the audio signals supplied from the input buffer unit 16A, and use the original audio signal as the first audio signal x1, and the duplicated audio signals as the second audio signal x2 and the third audio signal x3.

[0086] (Correction signal generation unit) The correction signal generation unit 12A generates a phase correction signal p based on the second audio signal x2 supplied from the audio signal generation unit 11A. As described above, the audio signal generation unit 11A according to this embodiment generates the third audio signal x3, which has the same phase as the second audio signal x2. Therefore, the correction signal generation unit 12A according to this embodiment generates the phase correction signal p based on the third audio signal x3. As a result, the correction signal generation unit 12A generates the phase correction signal p based on a dedicated third audio signal. Therefore, it is possible to prevent the second audio signal from being out of phase by using the second audio signal to generate the phase correction signal p. The correction signal generation unit 12A then outputs the generated phase correction signal p to the operational amplifier 132a of the filter 13A.

[0087] As shown in FIG. 7, the correction signal generator 12A according to this embodiment includes a capacitor 121 and a resistor 122. One end of the capacitor 121 is connected to the audio signal generator 11A, and the other end is connected to the non-inverting input terminal (+) of the operational amplifier 132a. One end of the resistor 122 is connected to the other end of the capacitor 121, and a bias voltage is applied to the other end. The correction signal generator 12A configured in this manner outputs a phase correction signal p containing only frequency components equal to or greater than 1 / 2πCR (C: capacitance, R: resistance) using a high-pass filter function internally formed by the capacitor 121 and the resistor 122. Hereinafter, the lower limit value of the frequency components of the phase correction signal p is referred to as the phase correction start frequency fs. The correction signal generator 12A can adjust the phase correction start frequency fs by changing at least one of the capacitor 121 and the resistor 122. As will be described in detail later, the degree to which the operational amplifier 132a of the filter 13 corrects the phase varies within a range of 0 to −180° depending on the phase correction start frequency fs of the phase correction signal p input to the non-inverting input terminal (+), as shown in Fig. 5. For this reason, the phase correction start frequency fs is adjusted to an optimum value that minimizes the phase difference between the first audio signal x1 and the second audio signal x2 after correction.

[0088] Furthermore, the impedance Z (internal resistance) of capacitor 121 changes depending on the phase correction start frequency fs. Specifically, the impedance Z of capacitor 121 is inversely proportional to the phase correction start frequency fs, and its value is expressed as 1 / ωC (ω=2πfs). That is, the impedance Z of capacitor 121 decreases as the phase correction start frequency fs increases. Therefore, the correction signal generating unit 12A outputs a phase correction signal p with a higher voltage as the phase correction start frequency fs increases, and outputs a phase correction signal p with a lower voltage as the phase correction start frequency fs decreases. In this way, by using capacitor 121, the correction signal generating unit 12A can be configured with a simple configuration that is compatible with audio signals of various frequencies.

[0089] (filter) The filter 13A removes at least a first frequency component equal to or lower than a predetermined first frequency from the first audio signal. The first frequency is a frequency between the upper limit frequency of the first formant component h1 and the lower limit frequency of the second formant component h2. The filter 13A according to this embodiment is a band-pass filter that also removes a second frequency component equal to or higher than a second frequency higher than the first frequency. The filter 13A according to this embodiment is composed of a high-pass filter 131 and a low-pass filter 132.

[0090] The high-pass filter 131 removes a first frequency component from the first audio signal x1 supplied from the audio signal generation unit 11A. In the high-pass filter 131 according to the present embodiment, the first frequency is set to 400 Hz. The high-pass filter 131 then outputs the first audio signal x1' from which the first frequency component has been removed, i.e., an audio signal that does not include the first-order formant component h1 but includes the second-order formant component h2, the third-order formant component h3, ..., and the n-order formant component hn, to the low-pass filter 132.

[0091] The high-pass filter 131 according to this embodiment has an operational amplifier (not shown) that amplifies the first audio signal x1. Therefore, the high-pass filter 131 according to this embodiment outputs the first audio signal x1 from which the first frequency component has been removed and which has been amplified. The operational amplifier amplifies the first audio signal x1 from which the first frequency component has been removed, thereby outputting the first audio signal x1" in which the intensities of the second-order formant component h2, the third-order formant component h3, ..., and the n-order formant component hn have been increased. Note that the high-pass filter 131 may be configured so that at least one of the gain of the operational amplifier and the boundary frequency for whether or not to remove the components can be set arbitrarily depending on the intended use.

[0092] The low-pass filter 132 removes a second frequency component from the first audio signal x1, which is supplied from the high-pass filter 131 and from which the first frequency component has been removed. The second frequency component is a frequency component equal to or higher than the second frequency. The second frequency (cutoff frequency) is a frequency between the upper limit frequency of the fifth-order formant component h5 and the lower limit frequency of the sixth-order formant component h6, or a frequency between the upper limit frequency of the sixth-order formant component h6 and the lower limit frequency of the seventh-order formant component h7. In the low-pass filter 132 according to this embodiment, the second frequency is set to 5 kHz or 7 kHz. Low-pass filter 132 then outputs first speech signal x1" from which the fifth, sixth, or higher frequency components have been removed, i.e., a speech signal including only second- to fifth-order formant components h2 to h5 or only second- to sixth-order formant components h2 to h6. Higher-order formant components become noise that does not contribute to language understanding. For this reason, by low-pass filter 132 removing the second frequency component, it is possible to improve the signal-to-noise ratio in synthesized speech signal x', which will be described later.

[0093] However, when frequency components are removed from the first speech signal, a phase delay inevitably occurs. If a first speech signal x1, whose phase is delayed by time T, is combined with a second speech signal x2 without correction, the voltage V1 of the first speech signal x1 becomes −V1 due to the phase delay, as shown in FIG. 9. Therefore, the voltage V of the combined speech signal becomes V2 + (−V1) because the first speech signal x1 and the second speech signal x2 cancel each other out. As a result, as shown in FIG. 10, a dead zone (deep valley) called a dip occurs in a certain frequency range, causing extremely low sound pressures for the first-order formant component h1 and the second-order formant component h2, which are particularly important for speech comprehension, resulting in problems that significantly hinder improvements in clarity.

[0094] For this reason, the filter 13A has a function of correcting the phase delay. Specifically, as shown in FIG. 7, an operational amplifier 132a is provided in the low-pass filter 132 together with other elements 132b and 132c. The operational amplifier 132a amplifies the first audio signal x1. The operational amplifier 132a is configured so that the first audio signal is input to its inverting input terminal (-). The operational amplifier 132a receives a phase correction signal p at its non-inverting input terminal (+), and outputs a first audio signal x1' whose phase has been corrected to approach the phase of the second audio signal x2. Specifically, the operational amplifier 132a compares the voltage of the first audio signal input to its inverting input terminal (-) and the voltage of the phase correction signal p input to its non-inverting input terminal (+). If the voltage of the phase correction signal p is higher, the operational amplifier 132a outputs the first audio signal x1' whose phase has been advanced in accordance with the voltage difference. On the other hand, if the voltage of the phase correction signal p is lower, the operational amplifier 132a outputs a first audio signal x1' whose phase has been delayed in accordance with the voltage difference. As described above, the voltage of the phase correction signal p varies in accordance with the phase correction start frequency fs determined by the capacitor 121 and resistor 122 of the correction signal generation unit 12A. Therefore, the degree to which the operational amplifier 132a advances or delays the phase varies within a range of 0 to -180° depending on the phase correction start frequency fs, as shown in FIG. 8, for example. As described above, the phase correction start frequency fs is adjusted in advance. Therefore, a phase correction signal p having a voltage corresponding to the adjusted phase correction start frequency fs is input to the non-inverting input terminal (+) of the operational amplifier 132a. The operational amplifier 132a then outputs the first audio signal x1' corrected so that the difference in phase with the second audio signal x2 is minimized.

[0095] The filter 13A may be configured such that the low-pass filter 132 first removes the second frequency component from the first audio signal, and the high-pass filter 131 removes the first frequency component from the first audio signal x1 from which the second frequency component has been removed. In this case, the phase correction signal p is input to the non-inverting input terminal (+) of the operational amplifier included in the high-pass filter 131.

[0096] (adjustment section) The adjustment unit 17 adjusts the sound pressure of the phase-corrected first audio signal x1'. Then, the adjustment unit 17 supplies the adjusted first audio signal x1' to the synthesis unit 14A. The adjustment unit 17 adjusts the first audio signal based on the user's setting operation according to the usage environment and the degree of hearing loss. Therefore, the adjustment unit 17 can adjust only the first to fifth formant components h1 to h5, which are important for language comprehension, to a level that exceeds the masking line MK', which is an audible level. In addition, the adjustment unit 17 can freely select the output level curve of the synthesized audio signal x' to be "minimum," "maximum," or "intermediate," as shown in FIG. 11.

[0097] (Synthesis section) The synthesis unit 14A generates a synthetic audio signal x' by additively synthesizing the first audio signal x1', from which the first frequency component and the second frequency component have been removed and whose phase has been corrected, with the second audio signal x2 supplied from the audio signal generation unit 11A. As described above, the audio processor 10A according to this embodiment includes the adjustment unit 17. Therefore, the synthesis unit 14A according to this embodiment additively synthesizes the first audio signal x1, from which the sound pressure has been adjusted, with the second audio signal x2. The synthesis unit 14A then supplies the generated synthetic audio signal x' to the second port 18. As shown in FIG. 12, the phase of the corrected first audio signal x1' in the synthesis unit 14A leads the phase of the uncorrected first audio signal x1 by T'. Therefore, the corrected first audio signal x1' is in phase with the second audio signal x2. As a result, the frequency characteristics of the synthetic audio signal x' output by the synthesis unit 14A become as shown in FIG. 13.

[0098] (2nd port) The second port 18 outputs the synthesized voice signal x′ supplied from the synthesis unit 14A to a device connected to the second port 18.

[0099] (Power supply part) The power supply unit supplies power to each unit requiring power in the audio processor 10 A. The form of the power supply unit is not limited, and it may be an AC / DC converter such as an AC adapter, or a battery.

[0100] (Detection unit) The detection unit is configured to detect whether power is being supplied from the power supply unit to the audio processing unit. If the power supply unit is functioning normally, the detection unit supplies detection information indicating "Yes" to the control unit. On the other hand, if the power supply unit is not functioning normally, the detection unit supplies detection information indicating "No" to the control unit. If the power supply unit is an AC / DC converter, an example of a state in which the power supply unit is not functioning normally is a power outage. If the power supply unit is a battery, an example of a state in which the power supply unit is not functioning normally is a dead battery.

[0101] (Selection section) The selection unit selects whether or not to perform audio processing on the audio signal x under the control of the control unit. The selection unit is provided, for example, between the low-pass filter 132 and the synthesis unit 14A, and is configured by a switch that switches the circuit between ON and OFF. When the selection unit is ON, the first audio signal x1 that has passed through the high-pass filter 131 and the low-pass filter 132 is supplied to the synthesis unit 14A, where it is synthesized with the second audio signal x2. On the other hand, when the selection unit is OFF, the first audio signal x1 is not supplied to the synthesis unit 14A, and the second audio signal x2 is output from the synthesis unit 14A. Note that the selection unit may be provided between the audio signal generation unit 11A and the high-pass filter 131, or between the high-pass filter 131 and the low-pass filter 132.

[0102] (Control unit) The control unit controls the selection unit to apply or not apply audio processing to the audio signal x according to the detection result of the detection unit. If the detection result of the detection unit is "Yes", the control unit controls the selection unit to apply audio processing to the audio signal x (turn ON). On the other hand, if the detection result of the detection unit is "No", the control unit controls the selection unit to not apply audio processing to the audio signal x (turn OFF).

[0103] [Variations of the voice processing device] In the above embodiment, the high-pass filter 131, the low-pass filter 132, the correction signal generator 12A, and the synthesizer 14A of the audio processor 10A are configured as analog circuits. However, at least one of these may be configured as a digital circuit.

[0104] Furthermore, the audio processor 10A according to the above embodiment includes the high-pass filter 131 and the low-pass filter 132. However, instead of including these, the audio processor 10A may include a band-pass filter that simultaneously removes the first frequency component and the second frequency component.

[0105] Furthermore, the audio processor 10A may further include a polarity switch and an amplifier interposed between the first port 15A and the audio signal generator 11A.

[0106] The audio processor 10A also includes a selector, an amplifier, and a selector located between the synthesizer 14A and the second port 18. The amplifier may further include a circuit isolator, a transmitter power regulator, and a polarity switch.

[0107] The audio processor 10A may also include an attenuator between the synthesis unit 14A and the second port 18. The attenuator adjusts the output to suit the input audio source signal. By including such an attenuator, it is possible to provide a system that accommodates hearing loss symptoms due to hearing impairments by incorporating it into the microphone lines of existing wired or wireless audio transmission devices, or into the signal lines of televisions, radios, and other entertainment playback devices.

[0108] [Effects of the voice processing device] The audio processor 10A described above corrects the phase delay that occurs when the low-pass filter 132 removes the first and second frequency components based on the phase correction signal p. Therefore, the audio processor 10A can make audio even clearer than before.

[0109] [Modification] The speech processors 10 and 10A may be realized to include or be combined with the configurations described in the patent publications associated with the following patent numbers, for example. The technologies described in Japanese Patent Nos. 3731179 and 6548938 are technologies that perform speech processing for speech clarity by focusing on first-order (or lower-order) formant components and higher-order formant components (see, for example, Figure 1 of Japanese Patent No. 3731179 and Figure 4 of Japanese Patent No. 6548938). Therefore, instead of the configuration shown in the block diagram of Figure 5, the configurations shown in Figures 2, 4, and 6 of Japanese Patent No. 3731179 or the configuration shown in Figure 5 of Japanese Patent No. 6548938 may be used to realize the speech processor 10. The audio processor 10 is required only to perform audio processing on the audio signal representing the voice of the speaking party supplied from the microphone 21 to clarify the audio, and to supply the processed audio signal to the speaker 22.

[0110] Moreover, instead of the audio processor 10 shown in FIG. 5, the configuration shown in FIG. 4 of Japanese Patent No. 7105756 can also be adopted.

[0111] [Additional Notes] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in the description for implementing the invention are also included in the technical scope of the present invention. [Explanation of symbols]

[0112] 1. Audio processing device 10 Audio Processor 11 Branch 12 Phase corrector 13 Filter (Bandpass Filter) 14 Amplifier 15 Regulator 16 Additive combiner P11 port (first port) P12 port (second port) 21. Mike 22 speakers 1A Audio processing device 10A Audio Processor 10, 10A Audio Processor 11A Audio signal generation unit 12A correction signal generation section 121 Capacitor 122 resistor 13A filter 131 High-pass filter 132 Low-pass filter 132a operational amplifier 132b element 14A Synthesis section 15A Port 1 16A input buffer 17 Adjustment part 18 Second Port 21. Mike 22 speakers 24 Support member

Claims

1. 1. A portable audio processing device, comprising: an audio processor that applies audio processing to the audio signal supplied from the microphone to make the audio clearer; a speaker that converts the processed audio signal into audio; a housing that houses the audio processor and the speaker, The housing is separate from the microphone, The audio processor an audio signal generation unit that generates a first audio signal and a second audio signal having the same spectrum from an audio signal representing a speech; a filter that removes at least a first frequency component that is equal to or lower than a predetermined first frequency from the first audio signal; a synthesis unit that generates a synthesized audio signal by additively synthesizing the first audio signal from which the first frequency component has been removed and the second audio signal; a correction signal generation unit that generates a phase correction signal based on the second audio signal; Equipped with the filter includes an operational amplifier that amplifies the first audio signal; The operational amplifier comprises: The first audio signal, from which a phase delay occurs due to the removal of the first frequency component, is input to an inverting input terminal; When the phase correction signal is input to a non-inverting input terminal, the first audio signal is output, the phase of which is corrected so as to approach the phase of the second audio signal.

1. A voice processing device comprising:

2. Further comprising a microphone for converting voice into an audio signal.

2. The audio processing device according to claim 1, wherein:

Citation Information

Patent Citations

  • Loudspeaker

    JP1995336786A

  • Voice processor, voice clearing device, and voice processing method

    JP2016110050A

  • Acoustic device and information management system

    WO2020004363A1