Sound input device, sound input system, and input sound processing method

By using a multi-microphone system to detect sound pressure and set response levels, an appropriate output sound signal is generated, solving the problem of low recognition rate of AI assistants under ambient sound interference and achieving efficient speech recognition in noisy environments.

CN115462095BActive Publication Date: 2026-05-01JVC KENWOOD CORP
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JVC KENWOOD CORP
Filing Date
2021-05-27
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

When there is a lot of noise in the surroundings, existing voice input devices have difficulty effectively recognizing the speaker's voice, resulting in a decrease in the recognition rate of AI assistants.

Method used

A multi-microphone system is employed, including a first microphone that picks up ambient sound outside the external auditory canal, a second microphone that picks up speech sound outside the external auditory canal near the mouth, and a third microphone that picks up bone conduction sound inside the external auditory canal. By detecting sound pressure and setting the response level, an output sound signal is generated, and appropriate input sound signals are selected or mixed to improve the recognition rate.

Benefits of technology

Even in noisy environments, it can improve the AI ​​assistant's recognition rate of the speaker's voice, maintain the stability and clarity of the output sound signal, and reduce the impact of ambient noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115462095B_ABST
    Figure CN115462095B_ABST
Patent Text Reader

Abstract

The sound input device (91) includes a first microphone (M1), a second microphone (M2), a third microphone (M3), a control unit (3), and a communication unit (4). The first microphone (M1) picks up sound at a first position outside the speaker's (H) ear canal (E1) and sends out a first input sound signal. The second microphone (M2) picks up sound at a second position outside the ear canal (E1) that is closer to the mouth than the first position and sends out a second input sound signal (SN2). The third microphone (M3) picks up sound inside the speaker's (H) ear canal (E1) and sends out a third input sound signal (SN3). The control unit (3) detects the sound pressure (Va) of the first input sound signal (SN1), sets the response level (RF1) of the second input sound signal (SN2) and the response level (RF2) of the third input sound signal (SN3) based on the detected sound pressure (Va), and generates an output sound signal (SNt) including at least one of the second input sound signal (SN2) and the third input sound signal (SN3) based on the response levels (RF1, RF2). The communication unit (4) transmits the output sound signal (SNt) to the outside.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a sound input device, a sound input system, and a method for processing input sound. Background Technology

[0002] Patent Document 1 describes a wireless head-mounted assembly, including a microphone and earphones, used as a sound input device to transmit the user's voice to an AI assistant. Patent Document 2 describes a wireless earphone with a microphone used as a sound input device.

[0003] Existing technical documents

[0004] Patent documents

[0005] Patent Document 1: Japanese Patent Application Publication No. 2020-030780;

[0006] Patent document 2: Japanese Patent Application Publication No. 2019-195179. Summary of the Invention

[0007] When a speaker transmits their voice to an AI assistant through a voice input device such as a wireless headset with a microphone in a noisy environment, the loud ambient noise, along with the speaker's voice, is picked up by the microphone of the voice input device and transmitted to the AI ​​assistant. Therefore, the AI ​​assistant may be unable to recognize the user's voice and thus fail to provide an appropriate response.

[0008] The purpose of this invention is to provide a voice input device, a voice input system, and a voice input processing method that can achieve a high recognition rate of the speaker's voice by an AI assistant even in loud ambient noise.

[0009] This embodiment provides a sound input device, comprising: a first microphone for picking up sound at a first position outside a speaker's external auditory canal and transmitting a first input sound signal; a second microphone for picking up sound at a second position outside the speaker's external auditory canal, closer to the speaker's mouth than the first position, and transmitting a second input sound signal; a third microphone for picking up sound inside the speaker's external auditory canal and transmitting a third input sound signal; a control unit for detecting the sound pressure of the first input sound signal, setting a first response level representing the response level of the second input sound signal and a second response level representing the response level of the third input sound signal based on the detected sound pressure, and generating an output sound signal containing at least one of the second input sound signal and the third input sound signal based on the first response level and the second response level; and a communication unit for transmitting the output sound signal to an external location.

[0010] This embodiment provides an input sound processing method, wherein: a sound picked up at a first position outside the speaker's external auditory canal is acquired as a first input sound signal; the sound pressure of the first input sound signal is detected; a sound picked up at a second position outside the speaker's external auditory canal, closer to the speaker's mouth than the first position, is acquired as a second input sound signal; a sound picked up inside the speaker's external auditory canal is acquired as a third input sound signal; based on the sound pressure of the first input sound signal, a first response level representing the response level of the second input sound signal and a second response level representing the response level of the third input sound signal are set; based on the first response level and the second response level, an output sound signal comprising at least one of the second input sound signal and the third input sound signal is generated; and the output sound signal is transmitted externally.

[0011] This embodiment provides a sound input system, including: a first sound input device; and a second sound input device capable of communicating with the first sound input device; the first sound input device and the second sound input device each include: a first microphone, which picks up sound at a first position outside the speaker's external auditory canal and sends out a first input sound signal; a second microphone, which picks up sound at a second position outside the speaker's external auditory canal, closer to the speaker's mouth than the first position, and sends out a second input sound signal; a third microphone, which picks up sound inside the speaker's external auditory canal and sends out a third input sound signal; and a control unit, which detects the sound pressure of the first input sound signal and determines the input sound signal based on the detected sound pressure. The system sets a first response level representing the response level of the second input sound signal and a second response level representing the response level of the third input sound signal. Based on the first response level and the second response level, it generates an output sound signal that includes at least one of the second input sound signal and the third input sound signal. The system also includes a communication unit that transmits the output sound signal to the outside. The control unit of the first sound input device determines the magnitude between the sound pressure of the first input sound signal in the first sound input device and the sound pressure of the first input sound signal in the second sound input device, and sets the output sound signal transmitted from the communication unit to the outside based on the determination result.

[0012] This embodiment provides an input sound processing method, which acquires a sound picked up at a first position outside the external auditory canal of a speaker's left ear as a first input sound signal, detects the sound pressure of the first input sound signal, acquires a sound picked up at a second position outside the external auditory canal of the speaker's left ear, closer to the speaker's mouth than the first position, as a second input sound signal, acquires a sound picked up inside the external auditory canal of the speaker's left ear as a third input sound signal, sets a first response level representing the response level of the second input sound signal and a second response level representing the response level of the third input sound signal based on the sound pressure of the first input sound signal, generates a left-side output sound signal containing at least one of the second and third input sound signals based on the first and second response levels, acquires a sound picked up at a first position outside the external auditory canal of the speaker's right ear as a fourth input sound signal, detects the sound pressure of the fourth input sound signal, acquires a sound picked up at a second position outside the external auditory canal of the speaker's right ear, closer to the speaker's mouth than the first position, as a fifth input sound signal, and acquires a sound picked up inside the external auditory canal of the speaker's right ear. This invention acquires the sound of sound as the sixth input sound signal. Based on the sound pressure of the fourth input sound signal, it sets a third response level representing the response level of the fifth input sound signal and a fourth response level representing the response level of the sixth input sound signal. Based on the third response level and the fourth response level, it generates a right-side output sound signal that includes at least one of the fifth and sixth input sound signals. It determines the magnitude of the sound pressure of the first input sound signal and the sound pressure of the fourth input sound signal. Based on the determination result, it sets at least one of the left-side and right-side output sound signals as the output sound signal to be sent externally.

[0013] According to this embodiment, even if the surrounding noise is loud, the AI ​​assistant can improve its recognition rate of the speaker's voice. Attached Figure Description

[0014] Figure 1 This is a schematic cross-sectional view showing the earphone 91 as a sound input device in the first embodiment.

[0015] Figure 2 This is a block diagram of the Headphone 91.

[0016] Figure 3 This is a diagram illustrating the actions of the headphone 91.

[0017] Figure 4 This is a block diagram of a headphone 91A, which is a first modified example of a sound input device according to the first embodiment.

[0018] Figure 5 This is a diagram illustrating the operation of the 91A headphone.

[0019] Figure 6 This is a diagram showing the operation of the earphone 91B, a second variation of the sound input device as described in the first embodiment.

[0020] Figure 7 This is a diagram showing the operation of the headphones 91C, a third variation of the sound input device according to the first embodiment.

[0021] Figure 8 This is a diagram showing the operation of the headphones 91D, a fourth variation of the sound input device as described in the first embodiment.

[0022] Figure 9 This is a block diagram of the headphone system 91ST, which is the sound input system of the second embodiment.

[0023] Figure 10 This is a table representing the operation of the 91ST headphone system.

[0024] Figure 11 This is a schematic cross-sectional view showing an example of the configuration position when the third microphone M3 is a bone conduction microphone. Detailed Implementation

[0025] (First Implementation)

[0026] Using headphones 91, refer to Figure 1 and Figure 2 The sound input device of this embodiment will be described.

[0027] Figure 1 This is a longitudinal cross-sectional view of the headphone 91. Figure 1 The image shows the earphone 91 being worn on the ear E of the speaker H. Figure 2 This is a block diagram of the Headphone 91.

[0028] The earphone 91 has a main body 1 and an insertion part 2 that protrudes from the main body 1 and is inserted into the external auditory canal E1. The main body 1 includes a first microphone M1, a second microphone M2, a control unit 3, a communication unit 4, a driver unit 5, and a speaker unit 6. The insertion part 2 has a third microphone M3. The control unit 3 has a sound pressure detection unit 3a and an input selection unit 3b.

[0029] The main body 1 has an air chamber 1a on the sound emitting side of the speaker unit 6. The insertion part 2 has a sound emitting passage 2a communicating with the air chamber 1a. The sound emitting passage 2a has an open front end. When the earphone 91 is in use, the sound output from the speaker unit 6 by the operation of the drive unit 5 is sent into the external auditory canal E1 through the air chamber 1a and the sound emitting passage 2a. Thus, the communication unit 4 of the earphone 91 can receive sound signals wirelessly transmitted from an external sound reproduction device and reproduce them by the speaker unit 6 through the control unit 3 and the drive unit 5.

[0030] When the headset 91 is in use, the first microphone M1 is positioned in the main body 1 at a first position, away from the speaker H's mouth, and picks up sounds from the surrounding area of ​​the main body 1. When the headset 91 is in use, the second microphone M2 is positioned in the main body 1 at a second position, closer to the speaker H's mouth, and primarily picks up the sound emitted by the speaker H as airborne sound. That is, when the headset 91 is in use, the second microphone M2 is located closer to the speaker H's mouth than the first microphone M1.

[0031] Hereinafter, the sound around the main body 1 will be referred to simply as ambient sound. The third microphone M3 is an air conduction microphone, positioned at a third position facing the sound emission path 2a of the insertion part 2. When the earphone 91 is in use, the third microphone M3 picks up the air conduction sound generated by the sound emitted by the speaker H as bone conduction sound reaching the external auditory canal E1 and reverberating within the external auditory canal E1 and the internal space Ev of the sound emission path 2a. That is, the first position of the first microphone M1 is located outside the external auditory canal of the speaker H. The second position of the second microphone M2 is located outside the external auditory canal of the speaker H, closer to the speaker H's mouth than the first position. The third microphone M3 is located inside the external auditory canal of the speaker H.

[0032] The sound pressure detection unit 3a of the control unit 3 detects the sound pressure of the first input sound signal SN1, which is input from the first microphone M1, and outputs it as a detection sound signal SN1a. The sound pressure of the input sound signal SN1 is detected as the equivalent noise level (LAeq). Hereinafter, the sound pressure of the detection sound signal SN1a, which is detected as the equivalent noise level (LAeq) by the sound pressure detection unit 3a, will be referred to as sound pressure Va. As described above, since the first microphone M1 mainly picks up ambient sound, the sound pressure Va can be regarded as the sound pressure of the ambient sound.

[0033] like Figure 2As shown, the input selection unit 3b of the control unit 3 receives a second input sound signal (SN2) from the second microphone M2, a third input sound signal (SN3) from the third microphone M3, and a detection sound signal (SN1a) from the sound pressure detection unit 3a. The input selection unit 3b generates an output sound signal (SNt) and outputs it to the communication unit 4. At this time, the input selection unit 3b sets the response level RF1 of the input sound signal SN2 and the response level RF2 of the input sound signal SN3 in the output sound signal (SNt) based on the sound pressure Va of the detection sound signal (SN1a). The response level RF1 and the response level RF2 are indicators that represent the degree of response of the input sound signal SN2 and the input sound signal SN3 to the output sound signal (SNt), respectively. The indicator is, for example, the magnitude of the sound pressure. The response level RF1 and the response level RF2 are also referred to as the first response level and the second response level, respectively.

[0034] Specifically, the sound pressure detection unit 3a acquires the sound picked up at a first position outside the speaker H's ear canal as a first input sound signal and detects the sound pressure of the first input sound signal. The input selection unit 3b acquires the sound picked up at a second position outside the speaker H's ear canal as a second input sound signal. The input selection unit 3b acquires the sound picked up inside the speaker H's ear canal as a third input sound signal. Based on the sound pressure of the first input sound signal, the input selection unit 3b sets a first response level RF1, representing the response level of the second input sound signal, and a second response level RF2, representing the response level of the third input sound signal. Based on the first response level and the second response level, the input selection unit 3b generates an output sound signal that includes at least one of the second and third input sound signals. The input selection unit 3b sends the generated output sound signal to the outside.

[0035] As an example, the input selection unit 3b sets the response level RF1 and response level RF2 to a binary choice, where one response is made but the other is not. That is, the input selection unit 3b selects one of the input sound signals SN2 and SN3 in a binary choice manner corresponding to the sound pressure Va of the detected sound signal SN1a, and sets the selected input sound signal as the output sound signal SNt, thereby generating the output sound signal SNt. Therefore, the output sound signal SNt includes at least one of the input sound signals SN2 and SN3.

[0036] The communication unit 4 wirelessly transmits the output audio signal SNt from the input selection unit 3b to the outside of the headset 91. The wireless transmission is performed, for example, via Bluetooth.

[0037] Next, refer to Figure 3A detailed description of the input sound processing method based on the action of the input selection unit 3b is provided. Figure 3 This is a graph where the horizontal axis is set to sound pressure level Va, and the vertical axis is set to the input sound signal SN2 of the second microphone M2 and the input sound signal SN3 of the third microphone M3, which are selected as output sound signals SNt. Using the value of sound pressure level Va, a switching lower sound pressure level Va1 and a switching upper sound pressure level Va2, which is greater than the switching lower sound pressure level Va1, are preset.

[0038] When the sound pressure level Va is less than the switching sound pressure level Va1, the input selection unit 3b selects the input sound signal SN2 and sets the selected input sound signal SN2 as the output sound signal SNt. Conversely, when the sound pressure level Va exceeds the switching sound pressure level Va2, the input sound signal SN3 is selected and set as the output sound signal SNt.

[0039] When the input selection unit 3b sets the input sound signal SN2 as the output sound signal SNt, if the sound pressure Va increases and exceeds the switching sound pressure Va2, then the input sound signal SN2 is switched to the input sound signal SN3, and the input sound signal SN3 is set as the output sound signal SNt. When the input selection unit 3b sets the input sound signal SN3 as the output sound signal SNt, if the sound pressure Va decreases and is less than the switching sound pressure Va1, then the input sound signal SN3 is switched back to the input sound signal SN2, and the input sound signal SN2 is set as the output sound signal SNt.

[0040] That is, when the ambient sound is low, the earphone 91 outputs the speaker H's voice, which is picked up by the second microphone M2 as air conduction outside the external auditory canal E1, as an output sound signal SNt. Conversely, when the ambient sound is high, the earphone 91 outputs the speaker H's voice, which is picked up by the third microphone M3 as bone conduction air conduction inside the external auditory canal E1, as an output sound signal SNt.

[0041] Within the external auditory canal E1, the speaker H's voice, picked up as bone conduction sound or through air conduction, while having lower clarity compared to the sound picked up through air conduction outside the external auditory canal E1, is almost unaffected by ambient noise and can be obtained at a stable sound pressure level. Therefore, even in noisy environments, the headphones 91 can deliver a high-sound-pressure output sound signal SNt from the speaker H's voice without being drowned out by ambient noise. Conversely, in quieter environments, when the speaker H's voice is picked up through air conduction outside the external auditory canal E1, the higher sound pressure level of the speaker H's voice results in a clearer output sound signal SNt.

[0042] like Figure 3As shown, the input selection unit 3b of the headphone 91 sets the upper sound pressure level Va2 (switching from input sound signal SN2 to input sound signal SN3) and the lower sound pressure level Va1 (switching from input sound signal SN3 to input sound signal SN2) to different values. Specifically, the upper sound pressure level Va2 is set to be higher than the lower sound pressure level Va1.

[0043] By making the value of the switching sound pressure Va2 different from the value of the switching sound pressure Va1, when the sound pressure Va of the surrounding sound picked up by the first microphone M1 frequently fluctuates across the switching sound pressure Va1 or the switching sound pressure Va2, the phenomenon that the sound pressure or sound quality of the output sound signal SNt becomes unstable due to the frequent switching between the input sound signal SN2 and the input sound signal SN3 can be avoided. Therefore, the decrease in voice recognition of the AI ​​assistant 81 caused by the sound pressure fluctuations of the surrounding sound accompanying the headphones 91 can be prevented.

[0044] In addition, by setting the upper sound pressure level Va2 to be greater than the lower sound pressure level Va1, the problem of being unable to switch to the correct input sound signal is prevented when the increase or decrease of sound pressure Va between the lower sound pressure level Va1 and the upper sound pressure level Va2 is reversed.

[0045] The values ​​of switching down sound pressure level Va1 and switching up sound pressure level Va2 are appropriately set by the manufacturer based on the usage environment of the headphones 91, so as to maintain a high recognition rate for the AI ​​assistant 81. Alternatively, the speaker H may be able to adjust the values ​​of switching down sound pressure level Va1 and switching up sound pressure level Va2 according to the usage environment of the headphones 91.

[0046] As described above, regardless of the ambient sound level around the main body 1, the speaker H's voice is maintained at a relatively high sound pressure level in the output sound signal SNt generated by the control unit 3 and sent from the communication unit 4. Therefore, the AI ​​assistant 81, upon receiving the output sound signal SNt, improves its recognition rate of the speaker H's voice.

[0047] The earphone 91 described above is not limited to the structure and steps described above, and can be modified within the scope of the spirit of the present invention.

[0048] (First variation)

[0049] Figure 4 This is a block diagram of the headphone 91A, which is a first modified example of the sound input device in this embodiment. Figure 5 This is a diagram illustrating the operation of the 91A headphone. (For example...) Figure 4 As shown, the headphone 91A replaces the input selection unit 3b with the input mixing unit 3c in the headphone 91, but the structure is otherwise the same.

[0050] The input sound signal SN2 from the second microphone M2, the input sound signal SN3 from the third microphone M3, and the detection sound signal SN1a from the sound pressure detection unit 3a are input to the input mixing unit 3c of the control unit 3. The input mixing unit 3c mixes the input sound signals SN2 and SN3 at a sound pressure ratio corresponding to the sound pressure Va of the detection sound signal SN1a, and outputs the mixture as an output sound signal SNt to the communication unit 4. The input mixing unit 3c sets the response level RF1 and the response level RF2 of the input sound signal SN2 in the output sound signal SNt according to their respective sound pressure ratios. The sound pressure ratio is the ratio of the sound pressure of the input sound signal SN2 to the sound pressure of the input sound signal SN3 contained in the output sound signal SNt.

[0051] Reference Figure 5 This describes an input sound processing method based on the actions of the input mixing unit 3c. Figure 5 In this diagram, the horizontal axis is defined as the linear axis of sound pressure Va, the left vertical axis is defined as the linear axis of the combined sound pressure V of the input sound signals SN2 and SN3, and the right vertical axis is defined as the linear axis of the total sound pressure Vt of the output sound signal SNt. The total sound pressure Vt is the sound pressure of the sound signal after mixing the input sound signals SN2 and SN3, and also includes the case where either the input sound signal SN2 or the input sound signal SN3 is 0 (zero).

[0052] like Figure 5 As shown, using the sound pressure level Va, a lower mixing limit sound pressure level Va3 and a higher mixing limit sound pressure level Va4, which is greater than the lower mixing limit Va3, are preset. Hereinafter, the range of sound pressure level Va above the lower mixing limit Va3 and below the upper mixing limit Va4 is also referred to as the mixing range R in sound pressure level Va. Furthermore, for input sound signals SN2 and SN3, a minimum mixing sound pressure level, i.e., the minimum mixing sound pressure level Vmin, and a maximum mixing sound pressure level, i.e., the maximum mixing sound pressure level Vmax, are preset. The value of the minimum mixing sound pressure level Vmin can be 0 (zero).

[0053] When the sound pressure Va is less than the lower mixing limit sound pressure Va3, the input mixing unit 3c sets the input sound signal SN2 to the maximum mixing sound pressure Vmax and the input sound signal SN3 to the minimum mixing sound pressure Vmin. When the sound pressure Va exceeds the upper mixing limit sound pressure Va4, the input mixing unit 3c sets the input sound signal SN2 to the minimum mixing sound pressure Vmin and the input sound signal SN3 to the maximum mixing sound pressure Vmax. Within the mixing range R of sound pressure Va, the input mixing unit 3c, with respect to the input sound signal SN2, reduces the mixed sound pressure V as the sound pressure Va increases, and with respect to the input sound signal SN3, increases the mixed sound pressure V as the sound pressure Va increases. That is, the greater the sound pressure Va, the less the response RF1 of the input sound signal SN2 decreases, and the more the response RF2 of the input sound signal SN3 increases. Within the mixing range R of sound pressure Va, the input mixing unit 3c makes the mixed sound pressure V increase or decrease linearly relative to the sound pressure Va, for example.

[0054] Therefore, the input mixing unit 3c mixes the input sound signal SN2 and the input sound signal SN3 with the mixing sound pressure V2x and the mixing sound pressure V3x corresponding to the sound pressure Va in any sound pressure Va mixing range R, and generates an output sound signal SNt, and outputs the generated output sound signal SNt to the communication unit 4.

[0055] Through the operation of the input mixing unit 3c described above, the total sound pressure Vt of the output sound signal SNt becomes a constant total sound pressure Vtc, independent of the magnitude of the sound pressure Va.

[0056] The values ​​of the mixed lower sound pressure level Va3, mixed upper sound pressure level Va4, mixed minimum sound pressure level Vmin, and mixed maximum sound pressure level Vmax are appropriately set by the manufacturer based on the usage environment of the headphones 91A, in order to maintain a high voice recognition rate for the AI ​​assistant 81. The values ​​of the mixed lower sound pressure level Va3, mixed upper sound pressure level Va4, mixed minimum sound pressure level Vmin, and mixed maximum sound pressure level Vmax can also be adjusted by the speaker H.

[0057] According to the headphone 91A, when the ambient sound sound pressure Va is within the mixing range R between the lower mixing limit sound pressure Va3 and the upper mixing limit sound pressure Va4, the input sound signal SN2 and the input sound signal SN3 are mixed with a sound pressure ratio of response levels RF1 and RF2 corresponding to the sound pressure Va. The ratio of the mixed sound pressure gradually changes linearly according to the increase or decrease of the ambient sound sound pressure of the main body 1. For example, the response level RF1 in the output sound signal SNt is represented by Vmax / Vmin when the sound pressure Va is Va3, by V2x / V3x when the sound pressure Va is Vax, and by Vmin / Vmax when the sound pressure Va is Va4. In addition, the response level RF2 in the output sound signal SNt is represented by Vmin / Vmax when the sound pressure Va is Va3, by V3x / V2x when the sound pressure Va is Vax, and by Vmax / Vmin when the sound pressure Va is Va4. Therefore, the change in sound quality of the output sound signal SNt corresponding to the increase or decrease of ambient sound becomes gradual and smooth, and the recognition rate of the AI ​​assistant 81 for the voice emitted by the speaker H is maintained at a high level regardless of the sound pressure level of the ambient sound around the main body 1. In addition, the total sound pressure level Vt of the output sound signal SNt of the earphone 91A remains constant and does not change abruptly regardless of the increase or decrease of ambient sound, so the recognition rate of the AI ​​assistant 81 for the voice emitted by the speaker H is maintained at an even higher level.

[0058] (Second variation)

[0059] The headphones 91A can also replace the input mixing unit 3c, which makes the total sound pressure Vt of the output sound signal SNt constant regardless of the sound pressure Va, and such as Figure 6 As shown, the input mixing section 3cB is configured to change the total sound pressure Vt according to the sound pressure Va (refer to...). Figure 4 The earphone 91B (see reference ) is a second variation of the sound input device in this embodiment. Figure 4 ).

[0060] The input mixing section 3cB, for example, within the mixing range R of sound pressure Va, causes the total sound pressure Vt to increase as sound pressure Va increases. More specifically, as... Figure 6 As shown, the input mixing unit 3cB performs a mixing operation by using different values ​​for the maximum mixing sound pressure level V2max of the input sound signal SN2 and the maximum mixing sound pressure level V3max of the input sound signal SN3. For example, the maximum mixing sound pressure level V3max is made greater than the maximum mixing sound pressure level V2max. As a result, the sound pressure level of the output sound signal SNt increases or decreases between the total sound pressure level Vt1 of the lower mixing limit sound pressure level Va3 and the total sound pressure level Vt2, which is greater than the total sound pressure level Vt1 of the upper mixing limit sound pressure level Va4.

[0061] With the total sound pressure Vt set constant, when the sound pressure Va is high (i.e., the ambient sound is loud), the sound pressure ratio of ambient sound, which is included to some extent as background noise in the input sound signal SN2, increases. Therefore, the sound pressure ratio of ambient sound in the total sound pressure Vt of the output sound signal SNt relatively increases. Conversely, in the earphone 91B, within the mixing range R of the sound pressure Va, as the sound pressure Va increases, the mixing ratio of the sound pressure of the input sound signal SN3 relative to the input sound signal SN2 increases. Therefore, the increase in the sound pressure ratio of ambient sound in the total sound pressure Vt of the output sound signal SNt is suppressed. Thus, the sound recognition rate of the output sound signal SNt received by the AI ​​assistant 81 is stably maintained.

[0062] (Third variation)

[0063] like Figure 7 As shown, the earphone 91A can also be replaced by an input mixing unit 3cC that non-linearly decreases and increases (see reference). Figure 4 The headphone 91C (see reference ) is a third variation of the sound input device in this embodiment. Figure 4 ).

[0064] like Figure 7 As shown, in the input mixing unit 3cC, within the mixing range R of sound pressure Va, the sound pressure Va5, which mixes the input sound signal SN2 and the input sound signal SN3 with equal sound pressure when the sound pressure Va decreases over time, is set closer to the lower mixing limit sound pressure Va3 than the midpoint between the lower mixing limit sound pressure Va3 and the upper mixing limit sound pressure Va4. That is, the input mixing unit 3cC performs the mixing of the input sound signal SN2 and the input sound signal SN3 when the sound pressure Va decreases based on nonlinear characteristic lines LN2b and LN3b.

[0065] On the other hand, in the input mixing unit 3cC, the sound pressure Va6, which is the sound pressure of the input sound signal SN2 and the input sound signal SN3 mixed with equal sound pressure as the sound pressure Va increases over time, is set closer to the upper mixing sound pressure Va4 than the midpoint between the lower mixing limit sound pressure Va3 and the upper mixing limit sound pressure Va4. That is, the input mixing unit 3cC performs the mixing of the input sound signal SN2 and the input sound signal SN3 as the sound pressure Va increases based on the nonlinear characteristic lines LN2a and LN3a.

[0066] Furthermore, when the sound pressure Va increases but does not reach the upper mixing limit Va4 and then decreases, the input mixing unit 3cC changes the mixing ratio on characteristic lines LN2a and LN3a. Conversely, when the sound pressure Va decreases but does not reach the lower mixing limit Va3 and then increases, the input mixing unit 3cC changes the mixing ratio on characteristic lines LN3b and LN2b.

[0067] In addition, the input mixing unit 3cC controls the mixing ratio of the input sound signal SN2 and the input sound signal SN3 so that the total sound pressure Vt of the output sound signal SNt becomes a constant total sound pressure Vtc regardless of the magnitude of the sound pressure Va. Figure 7 The non-linear characteristics of the input audio signals SN2 and SN3 are preset by the manufacturer of the headphones 91, or can be set by the speaker H.

[0068] In the headphone 91C, when the ambient sound sound pressure level Va is kept relatively low and close to the lower mixing limit Va3 of the mixing range R, the output sound signal SNt is generated by mixing the input sound signal SN2 with a higher ratio to the input sound signal SN3, prioritizing sound clarity. Conversely, when the ambient sound sound pressure level Va is kept relatively high and close to the upper mixing limit Va4 of the mixing range R, the headphone 91C generates the output sound signal SNt by mixing the input sound signal SN3 with a higher ratio to the input sound signal SN2, prioritizing high sound pressure level.

[0069] In this way, the earphone 91C generates an output sound signal SNt that is more suitable for sound recognition and corresponds to the magnitude of the sound pressure Va of the surrounding sound. Therefore, the recognition rate of the AI ​​assistant 81 for the voice emitted by the speaker H is maintained at a higher level.

[0070] (Fourth variation)

[0071] like Figure 8 As shown, the headphone 91C can also be replaced by an input mixing unit 3cD that changes the total sound pressure Vt according to the sound pressure Va (see reference). Figure 4 The earphone 91D (see reference ) is a fourth variation of the sound input device in this embodiment. Figure 4 ).

[0072] The input mixing section 3cD, for example, within the mixing range R of sound pressure Va, increases the total sound pressure Vt as sound pressure Va increases. Specifically, as... Figure 8As shown, the input mixing unit 3cD performs a mixing operation by setting the maximum mixing sound pressure level V2max of the input sound signal SN2 and the maximum mixing sound pressure level V3max of the input sound signal SN3 to different values. For example, the maximum mixing sound pressure level V3max is made greater than the maximum mixing sound pressure level V2max. As a result, the sound pressure level of the output sound signal SNt increases or decreases between the total sound pressure level Vt1 of the lower mixing limit sound pressure level Va3 and the total sound pressure level Vt2, which is greater than the total sound pressure level Vt1 of the upper mixing limit sound pressure level Va4.

[0073] Therefore, similar to the second variation, the greater the sound pressure Va, the greater the mixing ratio of the sound pressure of the input sound signal SN3 with that of the input sound signal SN2. Thus, the increase in the sound pressure ratio of ambient sounds in the total sound pressure Vt of the output sound signal SNt is suppressed, and the voice recognition rate of the AI ​​assistant 81 receiving the output sound signal SNt is stably maintained.

[0074] Furthermore, when selling headphones 91, 91A to 91D as products, they are not limited to single items; they can also be sold in sets of two or more.

[0075] If the earphones 91, 91A to 91D are designed to be worn on both ears, they can be sold as sets of earphones 91, 91A, 91A, 91B, 91B, 91C, 91C, 91D, and 91D. Alternatively, earphones 91, 91A to 91D, as single-ear earphones with microphones worn by multiple staff members in large stores, can be sold in sets of three or more.

[0076] (Second Implementation)

[0077] Using the 91ST headphone system, mainly referring to... Figure 1 , Figure 9 and Figure 10 The sound input system of this embodiment will be described. Figure 9 This is a block diagram of the 91ST headphone system. Figure 10 This is a table representing the operation of the 91ST headphone system.

[0078] like Figure 9 As shown, the headphone system 91ST is configured as a pair of headphones 91L as a first sound input device and headphones 91R as a second sound input device. Headphone 91L is worn in the left ear of the speaker H, and headphones 91R are worn in the right ear of the speaker H.

[0079] like Figure 1As shown, earphone 91L has a main body 1L and an insertion part 2, and earphone 91R has a main body 1R and an insertion part 2. The structure and arrangement of the first to third microphones M1 to M3, the driver part 5, and the speaker unit 6 in earphones 91L and 91R are the same as those in earphone 91 in the first embodiment. Hereinafter, the same reference numerals will be used to label the same parts as in earphone 91, and L and R will be used to distinguish different parts.

[0080] like Figure 1 as well as Figure 9 As shown, the earphones 91L and 91R have control units 3L and 3R respectively instead of the control unit 3 of the earphones 91, and communication units 4L and 4R respectively instead of the communication unit 4 of the earphones 91.

[0081] In the earphone 91L, the main body 1L includes a first microphone M1, a second microphone M2, a control unit 3L, a communication unit 4L, a driver unit 5, and a speaker unit 6. The insertion part 2 includes a third microphone M3. In the earphone 91R, the main body 1R includes a first microphone M1, a second microphone M2, a control unit 3R, a communication unit 4R, a driver unit 5, and a speaker unit 6. The insertion part 2 includes a third microphone M3.

[0082] like Figure 1 As shown, the main body 1L and main body 1R have air chambers 1a and 1a on the sound emitting side of the speaker units 6 and 6. Insertion parts 2 and 2 have sound emitting paths 2a and 2a communicating with the air chambers 1a and 1a. The sound emitting paths 2a and 2a have open front ends. When the earphones 91L and 91R are in use, the sound output from the speaker units 6 and 6 by the operation of the drive units 5 and 5 is transmitted through the air chambers 1a and 1a and the sound emitting paths 2a and 2a into the external auditory canals E1 and E1 of the left and right ears. Thus, the earphones 91L and 91R can receive sound signals wirelessly transmitted from an external sound reproduction device by the communication units 4L and 4R, and reproduced by the speaker units 6 and 6 through the control units 3L and 3R and the drive units 5 and 5. The earphones 91L and 91R can communicate with each other between the communication units 4L and 4R.

[0083] When the earphones 91L and 91R are in use, the first microphones M1 and M2 included in the main body 1L and 1R are positioned at the first position, away from the mouth of the speaker H, to pick up ambient sound from the main body 1L and 1R. When the earphones 91L and 91R are in use, the second microphones M2 and M2 included in the main body 1L and 1R are positioned at the second position, closer to the mouth of the speaker H. That is, when the earphones 91L and 91R are in use, the second microphones M2 and M2 are located closer to the mouth of the speaker H than the first microphones M1 and M1. The third microphones M3 and M3 are air conduction microphones and are positioned at the third position facing the sound transmission paths 2a and 2a of the insertion parts 2 and 2. When headphones 91L and 91R are in use, the third microphones M3 and M3 pick up the airborne sound generated by the sound emitted by the speaker H, which is transmitted through bone conduction and reverberates within the external auditory canals E1 and E1 and the internal spaces Ev and Ev of the sound transmission pathways 2a and 2a. Specifically, the first microphone M1 is positioned outside the speaker H's external auditory canal. The second microphone M2 is positioned outside the speaker H's external auditory canal, closer to the speaker H's mouth than the first position. The third microphone M3 is positioned inside the speaker H's external auditory canal.

[0084] like Figure 9 As shown, the control unit 3L of the earphone 91L includes a sound pressure detection unit 3aL, an input selection unit 3bL, and a sound pressure difference evaluation unit 3d. The control unit 3R of the earphone 91R includes a sound pressure detection unit 3aR, an input selection unit 3bR, and an output control unit 3e.

[0085] In the earphone 91L, the sound pressure detection unit 3aL detects the sound pressure of the input sound signal SN1L input from the first microphone M1 and outputs it as the detection sound signal SNL to both the input selection unit 3bL and the sound pressure difference evaluation unit 3d. In the earphone 91R, the sound pressure detection unit 3aR detects the sound pressure of the input sound signal SN1R input from the first microphone M1 and outputs it as the detection sound signal SNR to both the input selection unit 3bR and the output control unit 3e. Furthermore, the input sound signal SN1L and the input sound signal SN1R are also referred to as the first input sound signal. The input sound signal SN1L and the input sound signal SN1R can also be referred to as the first input sound signal and the fourth input sound signal, respectively.

[0086] The sound pressure of the input sound signals SN1L and SN1R is detected as the equivalent noise level (LAeq). Hereinafter, the sound pressure of the detected sound signals SNL and SNR, which are detected as the equivalent noise level (LAeq) by the sound pressure detection units 3aL and 3aR, will be referred to as sound pressure VL and VR, respectively.

[0087] The first microphone M1 included in the main body 1L picks up ambient sound from the main body 1L. The first microphone M1 included in the main body 1R picks up ambient sound from the main body 1R. Therefore, the sound pressure VL can be considered as the sound pressure of the ambient sound of the left earphone 91L. In addition, the sound pressure VR can be considered as the sound pressure of the ambient sound of the right earphone 91R.

[0088] The output control unit 3e of the headphone 91R sends sound pressure information JR1, which includes the sound pressure VR of the input detected sound signal SNR, and communication control information JR2 (described in detail later) to the communication unit 4R.

[0089] The input sound processing method based on the operation of the input selection units 3bL and 3bR of the headphones 91L and 91R is the same as that of the input selection unit 3b of the headphones 91 in the first embodiment.

[0090] like Figure 9 As shown, the input selection unit 3bL of the control unit 3L receives a second input sound signal (SN2L) from the second microphone M2, a third input sound signal (SN3L) from the third microphone M3, and a detection sound signal (SNL) from the sound pressure detection unit 3aL. The input selection unit 3bL generates an output sound signal (SNtL) and outputs it to the communication unit 4L. At this time, the input selection unit 3bL sets the response level RF1L and the response level RF2L of the input sound signal (SN2L) in the output sound signal (SNtL) based on the sound pressure VL of the detection sound signal (SNL). The response level RF1L and the response level RF2L are indicators that represent the degree of response of the input sound signal (SN2L) and the input sound signal (SN3L) to the output sound signal (SNtL), respectively. The indicator is, for example, the sound pressure level. The response level RF1L and the response level RF2L are also referred to as the first response level and the second response level, respectively. The output sound signal (SNtL) is also referred to as the left output sound signal.

[0091] Specifically, the sound pressure detection unit 3aL acquires the sound picked up at a first position outside the external auditory canal of the speaker H's left ear as a first input sound signal, and detects the sound pressure of the first input sound signal. The input selection unit 3bL acquires the sound picked up at a second position outside the external auditory canal of the speaker H's left ear as a second input sound signal. The input selection unit 3bL acquires the sound picked up inside the external auditory canal of the speaker H's left ear as a third input sound signal. Based on the sound pressure of the first input sound signal, the input selection unit 3bL sets a first response level representing the response level of the second input sound signal RF1L and a second response level representing the response level of the third input sound signal RF2L. Based on the first and second response levels, the input selection unit 3bL generates an output sound signal SNtL that includes at least one of the second and third input sound signals.

[0092] Similarly, the input sound signal SN2R, the fifth input sound signal from the second microphone M2, the input sound signal SN3R, the sixth input sound signal from the third microphone M3, and the detected sound signal SNR from the sound pressure detection unit 3aR are input to the input selection unit 3bR of the control unit 3R. The input selection unit 3bR generates an output sound signal SNtR and outputs it to the communication unit 4R. At this time, the input selection unit 3bR sets the response level RF1R and the response level RF2R of the input sound signal SN3R in the output sound signal SNtR based on the sound pressure VR of the detected sound signal SNR. The response level RF1R and the response level RF2R are indicators that represent the degree of response of the input sound signals SN2R and SN3R to the output sound signal SNtR, respectively. The indicator is, for example, the sound pressure level. The response level RF1R and the response level RF2R are also referred to as the third response level and the fourth response level, respectively. The output sound signal SNtR is also referred to as the right-side output sound signal.

[0093] Specifically, the sound pressure detection unit 3aR acquires the sound picked up at a first position outside the external auditory canal of the speaker H's right ear as the fourth input sound signal, and detects the sound pressure of the fourth input sound signal. The input selection unit 3bR acquires the sound picked up at a second position outside the external auditory canal of the speaker H's right ear as the fifth input sound signal. The input selection unit 3bR acquires the sound picked up inside the external auditory canal of the speaker H's right ear as the sixth input sound signal. Based on the sound pressure of the fourth input sound signal, the input selection unit 3bR sets a third response level (RF1R) representing the response level of the fifth input sound signal and a fourth response level (RF2R) representing the response level of the sixth input sound signal. Based on the third and fourth response levels, the input selection unit 3bR generates an output sound signal SNtR that includes at least one of the fifth and sixth input sound signals.

[0094] As an example, such as Figure 3 As shown, when the sound pressure VL of the detected sound signal SNL output from the sound pressure detection unit 3aL is less than the preset switching down sound pressure Va1, the input selection unit 3bL of the earphone 91L selects the input sound signal SN2L and sets the selected input sound signal SN2L as the output sound signal SNtL. Conversely, when the sound pressure VR exceeds the switching up sound pressure Va2, the input sound signal SN3L is selected and set as the output sound signal SNtL.

[0095] The input selection unit 3bL outputs the output audio signal SNtL, as set as described above, to the communication unit 4L. In this way, the input selection unit 3bL sets the response levels RF1L and RF2L of the input audio signal SN2L and input audio signal SN3L in the output audio signal SNtL based on the detected sound pressure level VL of the audio signal SNL. In this example, the input selection unit 3bL sets the response level RF1L and response level RF2L to a choice between responding to one and not the other.

[0096] Furthermore, when the sound pressure VR of the detected sound signal SNR output from the sound pressure detection unit 3aR is less than the preset switching down-sound pressure Va1, the input selection unit 3bR of the earphone 91R selects the input sound signal SN2R and sets the selected input sound signal SN2R as the output sound signal SNtR. Conversely, when the sound pressure VR exceeds the switching up-sound pressure Va2, the input sound signal SN3R is selected and set as the output sound signal SNtR.

[0097] The input selection unit 3bR outputs the output audio signal SNtR, as set as described above, to the communication unit 4R. In this way, the input selection unit 3bR sets the response levels RF1R and RF2R of the input audio signal SN2R and input audio signal SN3R in the output audio signal SNtR based on the sound pressure level VR detected by the audio signal SNR. In this example, the input selection unit 3bR sets the response level RF1R and response level RF2R to a choice between responding to one and not the other.

[0098] The communication unit 4R wirelessly transmits the sound pressure information JR1 input from the output control unit 3e to the outside of the headset 91R. The wireless transmission method is, for example, Bluetooth (registered trademark). Here, the presence or absence of wireless transmission of the output audio signal SNtR output from the input selection unit 3bR in the communication unit 4R is controlled by the communication control information JR2 input from the output control unit 3e. That is, the communication control information JR2 contains either a permission or a prohibition instruction for the wireless transmission of the output audio signal SNtR. The communication unit 4R controls the wireless transmission of the output audio signal SNtR according to this instruction.

[0099] The communication unit 4L receives the sound pressure information JR1 wirelessly transmitted from the communication unit 4R of the earphone 91R and sends it to the sound pressure difference evaluation unit 3d. The sound pressure difference evaluation unit 3d obtains the sound pressure VR based on the sound pressure information JR1 sent from the communication unit 4L and compares the sound pressure VR with the sound pressure VL of the detection sound signal SNL obtained from the sound pressure detection unit 3aL.

[0100] The sound pressure difference evaluation unit 3d, corresponding to the relationship between sound pressure VL and sound pressure VR, sets at least one of the output sound signal SNtL and the output sound signal SNtR as the output sound signal SNst wirelessly transmitted to the outside by the headphone system 91ST. That is, the sound pressure difference evaluation unit 3d determines the magnitudes of the sound pressure VL of the first input sound signal and the sound pressure VR of the fourth input signal, and based on the determination result, sets at least one of the left output sound signal and the right output sound signal as the output sound signal transmitted to the outside.

[0101] Next, the sound pressure difference evaluation unit 3d outputs communication control information JL2 to the communication unit 4L to determine the signal set as the output sound signal SNst. The communication unit 4L wirelessly transmits the communication control information JL2 to the communication unit 4R of the earphone 91R. When the communication unit 4R receives the communication control information JL2, it sends the received communication control information JL2 to the output control unit 3e.

[0102] Reference Figure 10 The operation of the 3D sound pressure difference evaluation unit is described in detail. Figure 10 This is a table showing the relationship between the magnitudes of sound pressure level VL and sound pressure level VR, and the output sound signal SNst wirelessly transmitted by the headphone system 91ST. For example... Figure 10 As shown, when it is determined that the sound pressure VR is greater than the sound pressure VL, the sound pressure difference evaluation unit 3d sets the output sound signal SNtL to the output sound signal SNst wirelessly transmitted by the headphone system 91ST.

[0103] Furthermore, the sound pressure difference evaluation unit 3d includes an instruction to wirelessly transmit the output sound signal SNtL in the communication control information JL2 and sends it to the communication unit 4L. If the sound pressure difference evaluation unit 3d determines that the sound pressure VR is less than the sound pressure VL, it includes an instruction to stop wirelessly transmitting the output sound signal SNtL in the communication control information JL2 and sends it to the communication unit 4L.

[0104] Communication unit 4L sends communication control information JL2 to communication unit 4R, and based on the instructions to communication unit 4L contained in communication control information JL2, executes or stops the wireless transmission of the output voice signal SNtL. On the other hand, communication unit 4R receives communication control information JL2 sent from communication unit 4L and sends it to output control unit 3e.

[0105] When communication control information JL2 contains an instruction to wirelessly transmit the output audio signal SNtL, output control unit 3e includes an instruction to stop wireless transmission of the output audio signal SNtR in communication control information JR2 and sends it to communication unit 4R. Alternatively, when communication control information JL2 contains an instruction to stop wireless transmission of the output audio signal SNtL, output control unit 3e includes an instruction to wirelessly transmit the output audio signal SNtR in communication control information JR2 and sends it to communication unit 4R. Communication unit 4R, based on the communication control information JR2 sent from output control unit 3e, executes or stops wireless transmission of the output audio signal SNtR.

[0106] The headphone system 91ST selectively chooses the output sound signal from the earphone 91L, 91R where the ambient sound is quieter and transmits it wirelessly to the outside. As a result, the AI ​​assistant 81 improves its recognition rate of the speaker H's voice within the headphone system 91ST.

[0107] The headphone system 91ST described above is not limited to the structure and sequence described, and may also be a modified example that is modified within the scope of the present invention.

[0108] Similar to the earphone 91A in the first variation of the first embodiment, earphones 91L and 91R can also replace input selection units 3bL and 3bR, and have input mixing units 3cL and 3cR that perform the same operation as input mixing unit 3c (see reference). Figure 9For example, in the input mixing unit 3cL, the input sound signal SN2L and the input sound signal SN3L are mixed with a sound pressure ratio of response levels RF1L and RF2L corresponding to the sound pressure VL of the detected sound signal SNL, and the sound pressure ratio gradually changes linearly according to the increase or decrease of the sound pressure of the surrounding sound of the main body unit 1. The response levels RF1L and RF2L are indicators representing the degree of response of the input sound signal SN2L and the input sound signal SN3L to the output sound signal SNtL, respectively. The indicator is, for example, the magnitude of the sound pressure. Therefore, the sound pressure ratio is the ratio of the sound pressure of the input sound signal SN2L to the sound pressure of the input sound signal SN3L contained in the output sound signal SNtL.

[0109] For example, such as Figure 5 As shown, the response level RF1L in the output sound signal SNtL is represented by Vmax / Vmin when the sound pressure VL is Va3, by V2x / V3x when the sound pressure VL is Vax, and by Vmin / Vmax when the sound pressure VL is Va4. Similarly, the response level RF2 in the output sound signal SNtL is represented by Vmin / Vmax when the sound pressure VL is Va3, by V3x / V2x when the sound pressure VL is Vax, and by Vmax / Vmin when the sound pressure VL is Va4. Headphones 91L and 91R have input mixing units 3cL and 3cR respectively, replacing input selection units 3bL and 3bR. This results in a smoother and more gradual change in the sound quality of the output sound signal SNst in response to changes in ambient sound. Therefore, regardless of the sound pressure level of the surrounding sound in the main body 1L or the main body 1R, the AI ​​assistant 81 maintains a high recognition rate for the voice emitted by the speaker H. Furthermore, since the earphones 91L and 91R each have input mixing sections 3cL and 3cR respectively, the total sound pressure level of the output sound signal SNst remains constant and does not change abruptly, regardless of the increase or decrease of the surrounding sound. Therefore, the AI ​​assistant 81 maintains an even higher recognition rate for the voice emitted by the speaker H.

[0110] The headphone system 91ST may also have an input mixing unit that performs the same operation as each of the input mixing units 3cB, 3cC, and 3cD, in the same way as each of the second to fourth variations of the first embodiment, where the headphone 91L and headphone 91R have input mixing units 3cL and 3cR respectively.

[0111] The wireless communication methods of Communication Units 4, 4L, and 4R are not limited to Bluetooth (registered trademark) as described above, and various methods can be applied. Furthermore, Communication Units 4, 4L, and 4R are not limited to wireless communication with external systems; they can also be used in a wired manner.

[0112] In the first embodiment, the first to fourth variations of the first embodiment, and the second embodiment, the sound input device, namely the headphones 91, 91A to 91D, 91L, and 91R, the third microphone M3 is not limited to the air conduction microphone described above, but may also be a bone conduction microphone that picks up bone conduction sound. Figure 11 This diagram illustrates the configuration of the third microphone, M3, when it is a bone conduction microphone. (See diagram for example.) Figure 11 As shown, the third microphone M3 is a bone conduction microphone. When the insertion part 2 is inserted into the external auditory canal E1, it is located in the third position that is in close contact with the inner surface of the external auditory canal E1, and picks up the bone conduction sound of the speaker H.

[0113] In the sound input system 91ST, the use of the earphone 91L as the first sound input device and the earphone 91R as the second sound input device is not limited to wearing them on one ear and the other ear of the speaker H. For example, it is also possible to wear the earphone 91L on the first speaker's ear and the earphone 91R on the ear of a second speaker who is different from the first speaker.

[0114] The indicators of the degree of reflection RF1L, RF2L, RF1, and RF2 are not limited to sound pressure level; they can also be physical quantities related to sound quality.

[0115] This application claims priority based on Japan Patent Application Nos. 2020-094795 and 2020-094797, filed on May 29, 2020, the entire disclosure of which is set forth herein by reference.

[0116] Industrial applicability

[0117] According to the voice input device, voice input system, and input voice processing method of this embodiment, even in loud ambient noise, the recognition rate of the AI ​​assistant's voice can be improved.

Claims

1. A sound input device, comprising: The first microphone picks up sound at a first position outside the speaker's ear canal and sends out the first input sound signal; The second microphone picks up sound at a second position outside the speaker's ear canal, closer to the speaker's mouth than the first position, and sends out a second input sound signal; The third microphone picks up sound from inside the speaker's ear canal and sends out the third input sound signal; The control unit detects the sound pressure of the first input sound signal, sets a first response level representing the response level of the second input sound signal and a second response level representing the response level of the third input sound signal based on the detected sound pressure, and generates an output sound signal containing at least one of the second input sound signal and the third input sound signal based on the first response level and the second response level. as well as The communication unit transmits the output audio signal to the outside. The control unit sets the first response level and the second response level to either reflect one without reflecting the other. In a two-choice manner corresponding to the detected sound pressure, one of the second input sound signal and the third input sound signal is set as the output sound signal.

2. The sound input device as described in claim 1, characterized in that, The control unit If the sound pressure of the detected first input sound signal is less than the preset first sound pressure, the second input sound signal is set as the output sound signal; If the sound pressure of the first input sound signal exceeds a second sound pressure that is greater than the first sound pressure, the third input sound signal is set as the output sound signal.

3. A sound input device, characterized in that, include: The first microphone picks up sound at a first position outside the speaker's ear canal and sends out the first input sound signal; The second microphone picks up sound at a second position outside the speaker's ear canal, closer to the speaker's mouth than the first position, and sends out a second input sound signal; The third microphone picks up sound from inside the speaker's ear canal and sends out the third input sound signal; The control unit detects the sound pressure of the first input sound signal, sets a first response level representing the response level of the second input sound signal and a second response level representing the response level of the third input sound signal based on the detected sound pressure, and generates an output sound signal containing at least one of the second input sound signal and the third input sound signal based on the first response level and the second response level. as well as The communication unit transmits the output audio signal to the outside. The control unit The first and second response levels are respectively set as sound pressure ratios. The second input sound signal and the third input sound signal are mixed and set as the output sound signal according to the sound pressure ratio corresponding to the detected sound pressure. If the sound pressure level of the detected first input sound signal is lower than a preset first sound pressure level, the sound pressure level of the second input sound signal is mixed with the sound pressure level of the third input sound signal to a greater extent. If the sound pressure of the detected first input sound signal exceeds a second sound pressure that is greater than the first sound pressure, the sound pressure of the third input sound signal is mixed with the sound pressure of the second input sound signal to a greater extent. When the sound pressure of the detected first input sound signal is in the range of above the first sound pressure and below the second sound pressure, the sound signals are mixed in such a way that the ratio of the sound pressure of the third input sound signal to the sound pressure of the second input sound signal increases as the sound pressure increases.

4. An input sound processing method, characterized in that, The sound picked up at a first position outside the speaker's external auditory canal is used as the first input sound signal; Detect the sound pressure level of the first input sound signal; The sound picked up at a second position outside the speaker's external auditory canal, which is closer to the speaker's mouth than the first position, is used as the second input sound signal; The sound picked up in the speaker's external auditory canal is used as the third input sound signal; Based on the sound pressure of the first input sound signal, a first degree of response representing the degree of response of the second input sound signal and a second degree of response representing the degree of response of the third input sound signal are set; Based on the first and second response levels, one of the second and third input sound signals is selectively set as the output sound signal and sent to the outside.

5. An input sound processing method, wherein, Acquire the sound picked up at a first position outside the speaker's external auditory canal as the first input sound signal, and detect the first input sound signal; The sound picked up at a second position outside the speaker's external auditory canal, which is closer to the speaker's mouth than the first position, is used as the second input sound signal; Acquire the sound picked up in the speaker's external auditory canal as a third input sound signal; Regarding the second input sound signal and the third input sound signal If the sound pressure level of the detected first input sound signal is lower than a preset first sound pressure level, the sound pressure level of the second input sound signal is mixed with the sound pressure level of the third input sound signal to a greater extent. If the sound pressure of the detected first input sound signal exceeds a second sound pressure that is greater than the first sound pressure, the sound pressure of the third input sound signal is mixed with the sound pressure of the second input sound signal to a greater extent. If the detected sound pressure level of the first input sound signal is above the first sound pressure level and below the second sound pressure level, the sound signals are mixed in such a way that the ratio of the sound pressure level of the third input sound signal to the sound pressure level of the second input sound signal increases as the sound pressure level increases. It is then sent to the outside as an output sound signal.

6. A voice input system, comprising: First sound input device; as well as The second audio input device is capable of communicating with the first audio input device. The first audio input device and the second audio input device each include: The first microphone picks up sound at a first position outside the speaker's ear canal and sends out the first input sound signal; The second microphone picks up sound at a second position outside the speaker's ear canal, closer to the speaker's mouth than the first position, and sends out a second input sound signal; The third microphone picks up sound from inside the speaker's ear canal and sends out the third input sound signal; The control unit detects the sound pressure of the first input sound signal, sets a first response level representing the response level of the second input sound signal and a second response level representing the response level of the third input sound signal based on the detected sound pressure, and generates an output sound signal containing at least one of the second input sound signal and the third input sound signal based on the first response level and the second response level; and The communication unit transmits the output audio signal to the outside. The control unit of the first sound input device determines the difference between the sound pressure of the first input sound signal in the first sound input device and the sound pressure of the first input sound signal in the second sound input device, and sets the output sound signal to be sent from the communication unit to the outside based on the determination result.

7. The sound input system as described in claim 6, characterized in that, In each of the first and second audio input devices, the control unit The first level of responsiveness and the second level of responsiveness are set to be either one that reflects one without reflecting the other, and In a two-choice manner corresponding to the detected sound pressure, one of the second input sound signal and the third input sound signal is set as the output sound signal.

8. The sound input system as described in claim 6, characterized in that, In each of the first and second audio input devices, the control unit The first and second response levels are respectively set as sound pressure ratios, and The second input sound signal and the third input sound signal are mixed and set as the output sound signal according to the sound pressure ratio corresponding to the detected sound pressure.

9. An input sound processing method, characterized in that, The sound picked up at a first position outside the external auditory canal of the speaker's left ear is used as the first input sound signal; Detect the sound pressure level of the first input sound signal; The sound picked up at a second position, which is closer to the speaker’s mouth than the first position, outside the external auditory canal of the speaker’s left ear, is used as the second input sound signal; The sound picked up in the external auditory canal of the speaker's left ear is used as the third input sound signal; Based on the sound pressure of the first input sound signal, a first degree of response representing the degree of response of the second input sound signal and a second degree of response representing the degree of response of the third input sound signal are set; Based on the first level of response and the second level of response, a left-side output sound signal is generated, which includes at least one of the second input sound signal and the third input sound signal; The sound picked up at a first position outside the external auditory canal of the speaker's right ear is used as the fourth input sound signal; Detect the sound pressure level of the fourth input sound signal; The sound picked up at a second position, which is closer to the speaker’s mouth than the first position, outside the external auditory canal of the speaker’s right ear, is used as the fifth input sound signal; The sound picked up in the external auditory canal of the speaker's right ear is used as the sixth input sound signal; Based on the sound pressure of the fourth input sound signal, a third degree of response representing the degree of response of the fifth input sound signal and a fourth degree of response representing the degree of response of the sixth input sound signal are set; Based on the third and fourth response levels, a right-side output sound signal containing at least one of the fifth and sixth input sound signals is generated; Determine the magnitude of the sound pressure level between the first input sound signal and the fourth input sound signal; as well as Based on the determination result, at least one of the left output sound signal and the right output sound signal is set as an output sound signal to be sent to the outside.

Citation Information

Patent Citations

  • Acoustic reproduction system

    JP2019195179A

  • Hands-free nursing and care recording system utilizing ai assistant

    JP2020030780A

  • Device for discharging explosive charge and military equipment with such device loaded

    JP2020094795A

  • Device management system

    JP2020094797A

  • Earphone signal processing method, earphone signal processing system and earphone

    CN111131947A