Ear-worn device, and play method

The ear-worn device distinguishes between direct speech and announcement sounds using signal processing, enhancing user experience by selectively attenuating or emphasizing these sounds for improved clarity.

JP2025160509APending Publication Date: 2025-10-22PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025134053
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-12-25
Filing Date
2025-08-12
Publication Date
2025-10-22

AI Technical Summary

Technical Problem

Existing ear-worn devices fail to effectively distinguish between direct human speech and announcement sounds, leading to inconsistent sound attenuation and user experience.

Method used

An ear-worn device equipped with a microphone, DSP, and speaker, which utilizes signal processing to differentiate between direct speech and announcement sounds, applying phase inversion or equalization to attenuate or emphasize specific frequency components based on the type of sound detected.

Benefits of technology

The device enhances user experience by selectively attenuating or emphasizing direct speech or announcement sounds, improving clarity and reducing noise interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025160509000001_ABST
    Figure 2025160509000001_ABST
Patent Text Reader

Abstract

To provide an ear-worn device capable of attenuating sounds while considering whether the voice included in the sound is an announcement sound.SOLUTION: An ear-worn device 20 includes: a microphone 21 that acquires sounds and outputs a sound signal of the acquired sound; a DSP 22 that determines whether the audio included in the above sound is an announcement sound, and based on the result of determining whether the audio included in the sound is an announcement sound, outputs a second sound signal in which a series of second signal processing, including phase inversion that is performed on the above sound signal; a speaker 28 that plays sound based on the output second sound signal; and a housing 29 that houses the microphone 21, the DSP 22, and the speaker 28.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an ear-worn device and a playback method. [Background technology]

[0002] Various technologies have been proposed for ear-mounted devices such as earphones and headphones. Patent Document 1 discloses a technology for canal-type earphones. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2012-249184 Summary of the Invention [Problem to be solved by the invention]

[0004] The present disclosure provides an ear-worn device that can attenuate a sound by taking into account whether the sound contained in the sound is an announcement sound. [Means for solving the problem]

[0005] An ear-worn device according to one embodiment of the present disclosure includes a microphone that acquires sound and outputs a sound signal of the acquired sound; a signal processing circuit that determines whether a voice contained in the sound is an announcement sound and, based on the determination result of whether the voice contained in the sound is an announcement sound, outputs a second sound signal that has been subjected to second signal processing including phase inversion processing on the sound signal; a speaker that reproduces sound based on the output second sound signal; and a housing that contains the microphone, the signal processing circuit, and the speaker. [Effects of the Invention]

[0006] An ear-worn device according to one aspect of the present disclosure can attenuate a sound by taking into consideration whether the sound contained in the sound is an announcement sound. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 1 is an external view of devices constituting a sound signal processing system according to an embodiment. [Figure 2] FIG. 2 is a block diagram showing a functional configuration of the sound signal processing system according to the embodiment. [Figure 3] FIG. 3 is a sequence diagram of the operation for setting the operation mode. [Figure 4] FIG. 4 is a diagram showing an example of the operation mode selection screen. [Figure 5] FIG. 5 is a flowchart of an example of operation in the announcement mode. [Figure 6] FIG. 6 is a flowchart of an example of operation in the interactive mode. [Figure 7] FIG. 7 is a flowchart of an example of operation in the voice detection mode. [Figure 8] FIG. 8 is a diagram for explaining the onset time. [Figure 9] FIG. 9 is a diagram showing an example of onset information of human speech sounds that reach directly. [Figure 10] FIG. 10 is a diagram showing an example of onset information of the announcement sound. [Figure 11] FIG. 11 is a diagram showing the power spectrum of human speech sounds that reach the listener directly. [Figure 12] FIG. 12 is a diagram showing the power spectrum of reverberation sounds contained in human speech sounds that reach the listener directly. [Figure 13] FIG. 13 is a diagram showing the power spectrum of an attack sound contained in a person's speech sound that reaches the listener directly. [Figure 14] FIG. 14 is a diagram showing the power spectrum of the announcement sound. [Figure 15] FIG. 15 is a diagram showing the power spectrum of the reverberation sound included in the announcement sound. [Figure 16] FIG. 16 is a diagram showing the power spectrum of an attack sound included in an announcement sound. DETAILED DESCRIPTION OF THE INVENTION

[0008] Hereinafter, the embodiments will be described in detail with reference to the drawings. Note that the embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, component placement and connection forms, steps, and step order shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components not recited in independent claims will be described as optional components.

[0009] It should be noted that the drawings are schematic diagrams and are not necessarily strict illustrations. In addition, in the drawings, substantially the same components are denoted by the same reference numerals, and overlapping descriptions may be omitted or simplified.

[0010] (Embodiment) [composition] First, the configuration of a sound signal processing system according to an embodiment will be described. Fig. 1 is an external view of devices constituting the sound signal processing system according to an embodiment. Fig. 2 is a block diagram showing the functional configuration of the sound signal processing system according to an embodiment.

[0011] As shown in FIGS. 1 and 2, a sound signal processing system 10 according to an embodiment includes an ear-worn device 20 and a mobile terminal 30.

[0012] First, the ear-worn device 20 will be described. The ear-worn device 20 is an earphone-type device that plays a third sound signal provided from the mobile terminal 30. The third sound signal is, for example, a sound signal of music content. The ear-worn device 20 has a noise cancellation function that reduces environmental sound (noise) around the user wearing the ear-worn device 20 while the third sound signal (music content) is being played. The ear-worn device 20 also has an ambient sound capture function that captures sounds around the user while the third sound signal is being played. Furthermore, the ear-worn device 20 can distinguish whether the human voice is a speech sound that directly reaches the user (a sound that the user hears when someone speaks to them) or an announcement sound, and selectively apply the ambient sound capture function to either the speech sound or the announcement sound that directly reaches the user.

[0013] Speech sounds that reach the user directly are sounds in which the direct sound component is relatively strong compared to the indirect sound component and the reverberation is small. Announcement sounds are human voices that are output from the speaker and reach the ear-worn device 20, and are sounds in which the indirect sound component is relatively strong compared to the direct sound component and the reverberation is large. Announcement sounds are specifically sounds that are output for guidance purposes at airports, train stations, on trains, etc.

[0014] Direct sound refers to sound that arrives directly without being reflected from the sound source, while indirect sound refers to sound that arrives after being reflected from the sound source one or more times by an object. When sound from the same sound source reaches a listener as a direct sound and one or more indirect sounds, the frequency characteristics and phase of the sound change depending on the path. For this reason, when a listener hears a sound in which these sounds are superimposed, if the direct sound is relatively strong, they perceive the reverberation to be small, and if the direct sound is relatively weak, they perceive the reverberation to be large. For example, when a person speaks directly to a listener, the reverberation is small, and when an announcement is made (a special sound heard from a speaker), the reverberation is large. The reverberation is felt to be strong (not in the usual situation, but in general).

[0015] The ear-worn device 20 can selectively apply the external sound capture function to either the speech sound or the announcement sound that reaches the user directly by estimating whether the sound is an announcement sound or a sound from a person speaking directly to the user based on the level of reverberation.

[0016] Reverberation refers to the sensation that indirect sounds reflected from walls, ceilings, etc., are heard together with the direct sound within a few to several hundred milliseconds after the direct sound is heard, as if they were a single stream of sound. In other words, reverberant sound is a sound in which the direct sound is superimposed with indirect sounds that arrive with delay from multiple directions. Sound without reverberation is a sound in which the direct sound is dominant and the superimposed indirect sounds are small to the ear, or are suppressed to a negligible level.

[0017] Specifically, the ear-worn device 20 includes a microphone 21, a DSP 22, a communication module 27, and a speaker 28. The microphone 21, the DSP 22, the communication module 27, and the speaker 28 are housed in a housing 29 (shown in FIG. 1).

[0018] The microphone 21 is a sound collection device that acquires sounds around the ear-worn device 20 and outputs sound signals of the acquired sounds. Specifically, the microphone 21 is a condenser microphone, a dynamic microphone, a MEMS (Micro Electro Mechanical Systems) microphone, or the like, but is not particularly limited thereto. The microphone 21 may be omnidirectional or directional.

[0019] The DSP 22 performs signal processing on the sound signal output from the microphone 21 to realize a noise cancellation function and an ambient sound capture function. The noise cancellation function reduces noise by inverting the phase of the sound signal and reproducing it through the speaker 28. The ambient sound capture function, for example, emphasizes specific frequency components of the sound (e.g., frequency components between 100 Hz and 2 kHz) by performing equalization processing on the sound signal and reproducing the signal through the speaker 28. In the ear-worn device 20, the ambient sound capture function is used to emphasize human voice or announcement sounds. Note that the ambient sound capture function may also be a function that allows the user to hear the sound represented by the sound signal by reproducing the sound signal substantially as is through the speaker 28; equalization processing is not essential. The DSP 22 is an example of a signal processing circuit. The DSP 22 includes a filter unit 23, a signal processing unit 24, a neural network unit 25, and a memory unit 26. Hereinafter, the neural network unit 25 will also be referred to as the NN (Neural Network) unit 25.

[0020] The filter unit 23 includes a high-pass filter 23a, a low-pass filter 23b, and a band-pass filter 23c. The high-pass filter 23a attenuates components in a band of 200 Hz or less that are included in the sound signal output from the microphone 21. The low-pass filter 23b attenuates components in a band of 500 Hz or more that are included in the sound signal output from the microphone 21. The band-pass filter 23c attenuates components in a band of 200 Hz or less and a band of 5 kHz or more that are included in the sound signal output from the microphone 21. Note that these cutoff frequencies are merely examples, and the cutoff frequencies may be determined empirically or experimentally.

[0021] The signal processing unit 24 includes, as functional components, a reverberation detection unit 24a, a noise detection unit 24b, a speech detection unit 24c, and a switching unit 24d. The functions of the reverberation detection unit 24a, the noise detection unit 24b, the speech detection unit 24c, and the switching unit 24d are realized, for example, by a circuit corresponding to the signal processing unit 24 executing a computer program stored in the storage unit 26. The functions of the reverberation detection unit 24a, the noise detection unit 24b, the speech detection unit 24c, and the switching unit 24d will be described in detail below.

[0022] The NN unit 25 includes, as functional components, a speech determination unit 25a and a reverberation determination unit 25b. The functions of the speech determination unit 25a and the reverberation determination unit 25b are realized, for example, by a circuit equivalent to the NN unit 25 executing a computer program stored in the storage unit 26. The functions of the speech determination unit 25a and the reverberation determination unit 25b will be described in detail later.

[0023] The storage unit 26 is a storage device that stores a computer program executed by a circuit corresponding to the signal processing unit 24, a computer program executed by a circuit corresponding to the NN unit 25, and various information required to implement the noise cancellation function and the ambient sound capture function. The storage unit 26 is realized by a semiconductor memory or the like. Note that the storage unit 26 may be realized as an external memory to the DSP 22 instead of being an internal memory of the DSP 22.

[0024] The communication module 27 receives the third sound signal from the mobile terminal 30, mixes the received third sound signal with a processed sound signal (a first sound signal or a second sound signal, which will be described later) output by the DSP 22, and outputs the result to the speaker 28. The communication module 27 is realized by, for example, an SoC (System-on-a-Chip). The communication module 27 includes a communication circuit 27a and a mixing circuit 27b.

[0025] The communication circuit 27a receives the third sound signal from the mobile terminal 30. The communication circuit 27a is, for example, a wireless communication circuit, and communicates with the mobile terminal 30 based on a communication standard such as Bluetooth (registered trademark) or BLE (Bluetooth (registered trademark) Low Energy).

[0026] The mixing circuit 27b mixes one of the first and second sound signals output by the DSP 22 with the third sound signal received by the communication circuit 27a, and outputs the result to the speaker .

[0027] The speaker 28 reproduces sound based on the mixed sound signal obtained from the mixing circuit 27b. The speaker 28 is a speaker that emits sound waves toward the ear canal (eardrum) of the user wearing the ear-worn device 20, but may also be a bone conduction speaker.

[0028] Next, the mobile terminal 30 will be described. The mobile terminal 30 is an information terminal that functions as a user interface device in the sound signal processing system 10 by installing a predetermined application program. The mobile terminal 30 also functions as a sound source that provides a third sound signal (music content) to the ear worn device 20. Specifically, by operating the mobile terminal 30, the user can select music content to be played by the speaker 28 and switch the operation mode of the ear worn device 20. The mobile terminal 30 includes a UI (User Interface) unit 31, a communication circuit 32, an information processing unit 33, and a storage unit 34.

[0029] The UI unit 31 is a user interface device that accepts user operations and presents images to the user. The UI unit 31 is realized by an operation acceptance unit such as a touch panel and a display unit such as a display panel.

[0030] The communication circuit 32 transmits a third sound signal, which is a sound signal of the music content selected by the user, to the ear worn device 20. The communication circuit 32 is, for example, a wireless communication circuit, and communicates with the ear worn device 20 based on a communication standard such as Bluetooth (registered trademark) or BLT.

[0031] The information processing unit 33 performs information processing related to displaying an image on the display unit and transmitting the third sound signal using the communication circuit 32. The information processing unit 33 is realized by, for example, a microcomputer, but may also be realized by a processor. The image display function and the third sound signal transmission function are realized by the microcomputer or the like constituting the information processing unit 33 executing a computer program stored in the storage unit 34.

[0032] The storage unit 34 is a storage device that stores various information required for the information processing unit 33 to perform information processing, computer programs executed by the information processing unit 33, the third sound signal (music content), etc. The storage unit 34 is realized by, for example, a semiconductor memory.

[0033] [Operation mode setting operation] The ear-worn device 20 is provided with three operation modes, and the user can set one of the three operation modes to the ear-worn device 20. The operation of setting such an operation mode will be described below. Figure 3 is a sequence diagram of the operation of setting the operation mode.

[0034] First, the information processing unit 33 of the mobile terminal 30 displays an operation mode selection screen on the UI unit 31 (display unit) (S11). FIG. 4 is a diagram showing an example of the operation mode selection screen. As shown in FIG. 4, the operation modes include three modes: an announcement mode, a dialogue mode, and a voice detection mode. The announcement mode is an operation mode for selectively emphasizing announcement sounds to assist a user in hearing the announcement sounds. The dialogue mode is an operation mode for selectively emphasizing speech sounds that reach the user directly to assist a user in having a dialogue with another user. The voice detection mode is an operation mode for emphasizing a person's voice regardless of whether the person's voice is speech sounds that reach the user directly or an announcement sound to assist a user in hearing the person's voice. Details of the operation in each operation mode will be described later.

[0035] When this selection screen is displayed, the user operates the UI unit 31 of the mobile terminal 30 to select an operation mode, and the UI unit 31 accepts this operation (S12). When this operation is accepted by the UI unit 31, the information processing unit 33 transmits a setting command to the ear worn device 20 using the communication circuit 32 to set the selected operation mode to the ear worn device 20 (S13).

[0036] The communication circuit 27a of the ear worn device 20 receives the setting command. When the communication circuit 27a receives the setting command, the setting command is transferred from the communication module 27 to the DSP 22, and the operation mode selected by the user in step S12 is set in the DSP 22 (S14). Specifically, the setting value stored in the memory unit 26 of the DSP 22 is set to the value specified in the setting command (a value indicating one of the above three modes).

[0037] [Announcement mode operation example] Next, an example of operation of the ear worn device 20 set to the announcement mode will be described. Fig. 5 is a flowchart of an example of operation in the announcement mode of the ear worn device 20. The announcement mode is an example of the first mode, and is an operation mode for selectively emphasizing the announcement sound to help the user hear the announcement sound.

[0038] The microphone 21 acquires sound and outputs a sound signal of the acquired sound (S21). The reverberation detection unit 24a calculates an acoustic feature of the sound signal by performing signal processing on the sound signal output from the microphone 21, which has been filtered through the high-pass filter 23a (S22). The acoustic feature here is an acoustic feature for determining whether or not the human voice included in the sound acquired by the microphone 21 has a sense of reverberation. Specific examples of the acoustic feature will be described later. The detected acoustic feature is output to the reverberation determination unit 25b.

[0039] The noise detection unit 24b calculates the ZCR (Zero-Crossing Rate) of the sound signal by performing signal processing on the sound signal output from the microphone 21 and having been filtered by the low-pass filter 23b (S23). The ZCR is an acoustic feature for calculating whether the sound indicated by the sound signal is close to noise, and indicates the number of times the sound signal crosses zero or the number of times the sign of the sound signal is changed. The calculated ZCR is output to the voice determination unit 25a. Note that in step S23, other acoustic features for estimating noise, such as flatness (signal flatness ratio), may be calculated, and the other acoustic feature may be used instead of the ZCR in steps S24 and onward.

[0040] The voice detection unit 24c performs signal processing on the sound signal output from the microphone 21 and having been filtered through the band-pass filter 23c, thereby calculating MFCCs (Mel-Frequency Cepstral Coefficients) (S24). MFCCs are cepstral coefficients used as features in voice recognition and the like, and are obtained by converting a power spectrum compressed using a Mel filter bank into a logarithmic power spectrum and applying an inverse discrete cosine transform to the logarithmic power spectrum. The calculated MFCCs are output to the voice determination unit 25a.

[0041] The voice determination unit 25a determines whether or not the sound acquired by the microphone 21 includes a human voice, based on the ZCR output from the noise detection unit 24b and the MFCC output from the voice detection unit 24c (S25). The voice determination unit 25a includes a first machine learning model (neural network) that receives the ZCR and MFCC as input and outputs a determination result as to whether or not the sound includes a human voice, and can determine whether or not the sound acquired by the microphone 21 includes a human voice using this first machine learning model. The determination result is output to the reverberation determination unit 25b. Note that it is not essential that the determination be made based on both the ZCR and the MFCC; it is sufficient that the determination be made based on at least one of the ZCR and the MFCC. In other words, one of the noise detection unit 24b and the voice detection unit 24c may be omitted.

[0042] If the determination result output from the voice determination unit 25a indicates that the sound acquired by the microphone 21 includes human voice (Yes in S25), the reverberation determination unit 25b determines whether the human voice included in the sound acquired by the microphone 21 has a sense of reverberation based on the acoustic features output from the reverberation detection unit 24a (S26). In this embodiment, determining whether the voice has a sense of reverberation does not mean in the strict sense, but rather means determining the degree (large or small) of the sense of reverberation that the human voice has. Whether the human voice has a sense of reverberation can be rephrased as whether the sense of reverberation included in the human voice is strong, whether the reverberation sound components included in the human voice are greater than a predetermined amount, etc.

[0043] Specifically, the reverberation determination unit 25b inputs the acoustic feature values ​​output from the reverberation detection unit 24a to a second machine learning model (neural network) included in the reverberation determination unit 25b. This second machine learning model receives the acoustic feature values ​​as input and outputs a determination result as to whether or not the human voice has a sense of reverberation. In other words, the reverberation determination unit 25b can use this second machine learning model to determine whether or not the human voice contained in the sound acquired by the microphone 21 has a sense of reverberation. The reverberation determination unit 25b outputs the determination result to the switching unit 24d.

[0044] The switching unit 24d switches between performing equalization processing (an example of first signal processing) or phase inversion processing (an example of second signal processing) on ​​the sound signal output by the microphone 21 based on the judgment result output from the voice judgment unit 25a and the judgment result output from the reverberation judgment unit 25b.

[0045] When the determination result output from the reverberation determination unit 25b indicates that the human voice included in the sound acquired by the microphone 21 has a sense of reverberation (Yes in S26), in other words, an announcement sound has been acquired by the microphone 21. In such a case, the switching unit 24d performs an equalization process on the sound signal to emphasize a specific frequency component, and outputs the result as a first sound signal (S27). The specific frequency component is, for example, a frequency component between 100 Hz and 2 kHz.

[0046] The mixing circuit 27b mixes the first sound signal with the third sound signal (music content) received by the communication circuit 27a and outputs the mixed sound to the speaker 28 (S29), and the speaker 28 reproduces sound based on the first sound signal mixed with the third sound signal (S30). As a result of the processing of step S27, the announcement sound is emphasized, making it easier for the user of the ear worn device 20 to hear the announcement sound.

[0047] On the other hand, when the determination result output from the voice determination unit 25a indicates that the sound acquired by the microphone 21 does not include human voice (No in S25), and when the determination result output from the reverberation determination unit 25b indicates that the human voice included in the sound acquired by the microphone 21 does not have a sense of reverberation (has little sense of reverberation) (No in S26), this is, in other words, when a sound other than the announcement sound has been acquired by the microphone 21. In such cases, the switching unit 24d performs phase inversion processing on the sound signal and outputs it as a second sound signal (S28).

[0048] The mixing circuit 27b mixes the second sound signal with the third sound signal (music content) received by the communication circuit 27a and outputs the mixed sound to the speaker 28 (S29), and the speaker 28 reproduces sound based on the second sound signal mixed with the third sound signal (S30). As a result of the processing of step S28, the user of the ear worn device 20 perceives the sounds around the ear worn device 20 as attenuated, and the user can clearly hear the music content.

[0049] As described above, during operation in the announcement mode, the DSP 22 determines whether the human voice contained in the sound acquired by the microphone 21 has a reverberant sound, and outputs a first sound signal if it determines that the human voice contained in the sound has a reverberant sound, and outputs a second sound signal if it determines that the human voice contained in the sound does not have a reverberant sound. The first sound signal is a sound signal obtained by performing equalization processing on the sound signal output from the microphone 21 to emphasize specific frequency components of the sound, and the second sound signal is a sound signal obtained by performing phase inversion processing on the sound signal output from the microphone 21.

[0050] This allows the ear worn device 20 operating in the announcement mode to attenuate sounds other than the announcement sound while assisting the user in hearing the announcement sound.

[0051] [Interactive mode example] Next, an example of operation of the ear worn device 20 set to the dialogue mode will be described. Fig. 6 is a flowchart of an example of operation in the dialogue mode of the ear worn device 20. The dialogue mode is an example of the second mode, and is an operation mode for supporting a user in dialogue with another user by selectively emphasizing speech sounds that reach the user directly.

[0052] The processing of steps S31 to S35 is the same as steps S21 to S25 in the operation example of the announcement mode. When the determination result output from the voice determination unit 25a indicates that the sound acquired by the microphone 21 includes a human voice (Yes in S35), the reverberation determination unit 25b determines whether the human voice included in the sound acquired by the microphone 21 has a sense of reverberation, based on the acoustic feature output from the reverberation detection unit 24a (S36).

[0053] After step S36, the switching unit 24d switches between performing equalization processing or phase inversion processing on the sound signal output by the microphone 21 based on the judgment result output from the voice judgment unit 25a and the judgment result output from the reverberation judgment unit 25b.

[0054] When the determination result output from the reverberation determination unit 25b indicates that the human voice contained in the sound acquired by the microphone 21 does not have a sense of reverberation (has little sense of reverberation) (No in S36), in other words, the microphone 21 has acquired a speech sound that directly reaches the user. In such a case, the switching unit 24d performs an equalization process on the sound signal to emphasize a specific frequency component, and outputs the result as a first sound signal (S37). The specific frequency component is, for example, a frequency component between 100 Hz and 2 kHz.

[0055] The mixing circuit 27b mixes the first sound signal with the third sound signal (music content) received by the communication circuit 27a and outputs the mixed sound to the speaker 28 (S39), and the speaker 28 reproduces sound based on the first sound signal mixed with the third sound signal (S40). As a result of the processing of step S37, the speech sound that reaches the user directly is emphasized, making it easier for the user of the ear worn device 20 to hear the speech sound that reaches the user directly.

[0056] On the other hand, when the determination result output from the voice determination unit 25a indicates that the sound acquired by the microphone 21 does not include human voice (No in S35), and when the determination result output from the reverberation determination unit 25b indicates that the human voice included in the sound acquired by the microphone 21 has a reverberant sound (Yes in S36), this is, in other words, when a sound other than speech sound that directly reaches the user is acquired by the microphone 21. In such a case, the switching unit 24d performs phase inversion processing on the sound signal and outputs it as a second sound signal (S38).

[0057] The mixing circuit 27b mixes the second sound signal with the third sound signal (music content) received by the communication circuit 27a and outputs the mixed sound to the speaker 28 (S39), and the speaker 28 reproduces sound based on the second sound signal mixed with the third sound signal (S40). As a result of the processing of step S38, the user of the ear worn device 20 perceives the sounds around the ear worn device 20 as attenuated, and the user can clearly hear the music content.

[0058] As described above, during operation in the interactive mode, the DSP 22 determines whether the human voice contained in the sound acquired by the microphone 21 has a reverberant sound, and outputs a first sound signal if it determines that the human voice contained in the sound does not have a reverberant sound, and outputs a second sound signal if it determines that the human voice contained in the sound has a reverberant sound. The first sound signal is a sound signal obtained by subjecting the sound signal output from the microphone 21 to equalization processing in order to emphasize specific frequency components of the sound, and the second sound signal is a sound signal obtained by subjecting the sound signal output from the microphone 21 to phase inversion processing.

[0059] This allows the ear-worn device 20 operating in the interactive mode to assist the user in interacting with other users while attenuating sounds other than speech sounds that reach the user directly.

[0060] [Example of voice detection mode operation] Next, an operation example of the ear worn device 20 set to the voice detection mode will be described. Fig. 7 is a flowchart of an operation example of the voice detection mode of the ear worn device 20. The voice detection mode is an example of the third mode, and is an operation mode for emphasizing a person's voice regardless of whether the person's voice is a speech sound that reaches the user directly or an announcement sound, thereby assisting the user in hearing the person's voice.

[0061] The microphone 21 acquires a sound and outputs a sound signal of the acquired sound (S41). The noise detection unit 24b calculates the ZCR of the sound signal by performing signal processing on the sound signal output from the microphone 21 and having been filtered by the low-pass filter 23b (S42). The calculated ZCR is output to the voice determination unit 25a.

[0062] The voice detection unit 24c calculates MFCCs by performing signal processing on the sound signal output from the microphone 21 and having been filtered by the band-pass filter 23c (S43). The calculated MFCCs are output to the voice determination unit 25a.

[0063] The voice determination unit 25a determines whether or not the sound acquired by the microphone 21 includes a human voice, based on the ZCR output from the noise detection unit 24b and the MFCC output from the voice detection unit 24c (S44). The specific processing in step S44 is the same as that in steps S25 and S35.

[0064] The switching unit 24d switches between performing equalization processing or phase inversion processing on the sound signal output by the microphone 21 based on the determination result output from the sound determination unit 25a.

[0065] If the determination result output from the voice determination unit 25a indicates that the sound acquired by the microphone 21 includes a human voice (Yes in S44), the switching unit 24d performs equalization processing on the sound signal to emphasize a specific frequency component, and outputs the result as a first sound signal (S45). The specific frequency component is, for example, a frequency component between 100 Hz and 2 kHz.

[0066] The mixing circuit 27b mixes the first sound signal with the third sound signal (music content) received by the communication circuit 27a and outputs the mixed sound to the speaker 28 (S47), and the speaker 28 reproduces sound based on the first sound signal mixed with the third sound signal (S48). As a result of the processing of step S45, the sound is emphasized, making it easier for the user of the ear worn device 20 to hear the sound.

[0067] On the other hand, if the judgment result output from the voice judgment unit 25a indicates that the sound acquired by the microphone 21 does not contain human voice (No in S44), the switching unit 24d performs phase inversion processing on the sound signal and outputs it as a second sound signal (S46).

[0068] The mixing circuit 27b mixes the second sound signal with the third sound signal (music content) received by the communication circuit 27a and outputs the mixed sound to the speaker 28 (S47), and the speaker 28 reproduces sound based on the second sound signal mixed with the third sound signal (S48). As a result of the processing of step S46, the user of the ear worn device 20 perceives the sounds around the ear worn device 20 as attenuated, and therefore the user can clearly hear the music content.

[0069] As described above, the DSP 22 operating in the voice detection mode determines whether or not the sound acquired by the microphone 21 includes human voice, and outputs a first sound signal if it determines that the sound includes human voice, and outputs a second sound signal if it determines that the sound does not include human voice. The first sound signal is a sound signal obtained by performing equalization processing on the sound signal output from the microphone 21 to emphasize specific frequency components of the sound, and the second sound signal is a sound signal obtained by performing phase inversion processing on the sound signal output from the microphone 21.

[0070] This allows the ear-worn device 20 operating in the voice detection mode to assist the user in hearing human voices while attenuating sounds other than human voices.

[0071] [Example of acoustic features 1] Next, a first example of an acoustic feature calculated by the reverberation detection unit 24a will be described. For example, onset information indicating the relationship between the change in sound pressure level of a sound signal over time and the onset time is used as the acoustic feature. The onset information includes a waveform indicating the change in sound pressure level over time and the position of the onset time in the waveform. Fig. 8 is a diagram for explaining the onset time, where (a) of Fig. 8 shows the change in the waveform of a sound signal over time, and (b) of Fig. 8 shows the change in sound power over time. More specifically, (b) of Fig. 8 is a diagram in which the waveform of (a) of Fig. 8 is frequency-decomposed to calculate a mel spectrogram, and the calculated mel spectrogram is superimposed to create an envelope in the time direction. As shown in Fig. 8, the onset time refers to the time when a sound starts to be produced.

[0072] Fig. 9 is a diagram showing an example of onset information for a person's speech sound that reaches directly, and Fig. 10 is a diagram showing an example of onset information for an announcement sound. Fig. 9 shows onset information obtained when a person's voice is directly picked up by a microphone, and Fig. 10 shows onset information obtained when the same person's voice is indirectly picked up by the same microphone via a speaker. In other words, the only difference between the onset information in Fig. 9 and the onset information in Fig. 10 is the presence or absence of reverberation (the degree of reverberation).

[0073] 9 and 10, the solid lines indicate the change in overall sound pressure level over time, which is obtained by extracting the sound pressure level at each frequency by frequency analysis of the sound signal of the human voice (specifically, frequency decomposition and calculation of a time-series envelope from a mel spectrogram) and superimposing the extracted sound pressure levels. In FIGS. 9 and 10, the dashed lines indicate the onset time. The onset time in FIGS. 9 and 10 is determined based on the change in sound pressure level at the frequency with the highest sound pressure level, which is determined by extracting the sound pressure level at each frequency by frequency analysis of the sound signal of the human voice.

[0074] In this way, the onset information is information including a waveform showing a change in sound pressure level over time and the position of the onset time in the waveform. In steps S22 and S32, the reverberation detection unit 24a calculates such onset information as an acoustic feature and outputs it to the reverberation determination unit 25b.

[0075] The second machine learning model included in the reverberation determination unit 25b is constructed in advance by learning sets of onset information (i.e., sets of onset information that differ only in the presence or absence of reverberation) as shown in Figures 9 and 10. During learning, the presence or absence of reverberation is assigned (annotated) as a label to the onset information.

[0076] In this way, the DSP 22 calculates onset information from the sound signal and, based on the calculated onset information, can determine whether or not the human voice contained in the sound acquired by the microphone 21 has a reverberant sound.

[0077] [Acoustic feature example 2] Next, a second example of the acoustic feature calculated by the reverberation detection unit 24a will be described. For example, the power spectrum of reverberant sound is used as the acoustic feature. FIG. 11 is a diagram showing the power spectrum of a speech sound that directly reaches the user. FIG. 12 is a diagram showing the power spectrum of a reverberant sound included in the speech sound that directly reaches the user. FIG. 13 is a diagram showing the power spectrum of an attack sound included in the speech sound that directly reaches the user. FIG. 14 is a diagram showing the power spectrum of an announcement sound. FIG. 15 is a diagram showing the power spectrum of a reverberant sound included in the announcement sound. FIG. 16 is a diagram showing the power spectrum of an attack sound included in the announcement sound. In FIGS. 11 to 16, the whiter the color, the higher the power value, and the darker the color, the lower the power value. The speech sound that directly reaches the user, which is the source of FIGS. 11 to 13, and the announcement sound that is the source of FIGS. 14 to 16 differ only in the presence or absence of reverberation (the degree of reverberation).

[0078] The power spectrum of reverberation sound is a partial power spectrum excluding the attack portion in Figure 8(b). The power spectrum of reverberation sound is a power spectrum obtained by extracting a continuous section in the time domain. Specifically, the power spectrum of reverberation sound is matrix information in which each element indicates a power value. The attack portion is the portion that corresponds to the point from when sound is generated to when sound pressure peaks when a continuous section in the frequency domain (a state in which sound is produced over a wide frequency band) is captured on the time axis, and the power spectrum of attack sound is a power spectrum obtained by extracting a continuous section in the frequency domain.

[0079] In steps S22 and S32, the reverberation detection unit 24a calculates the power spectrum of the reverberation sound as an acoustic feature and outputs it to the reverberation determination unit 25b. Any existing method may be used to calculate the power spectrum of the reverberation sound. Here, HPSS (Hermonic / Percussive Source Separation) is used, modified for reverberation detection.

[0080] The second machine learning model included in the reverberation determination unit 25b is constructed in advance by learning pairs of power spectra of reverberant sounds (i.e., pairs of power spectra of reverberant sounds that differ only in the presence or absence of reverberation) as shown in Figures 12 and 15. During learning, the presence or absence of reverberation is assigned (annotated) as a label to the power spectrum of the reverberant sound.

[0081] In this way, the DSP 22 can calculate the power spectrum of the reverberation sound from the sound signal and determine whether or not the human voice has a sense of reverberation based on the calculated power spectrum of the reverberation sound.

[0082] [Effects, etc.] As described above, the ear-worn device 20 includes the microphone 21 that acquires sound and outputs a sound signal of the acquired sound, the DSP 22 that performs signal processing on the sound signal to determine whether or not the sound contained in the sound has a reverberant sound, and outputs a first sound signal obtained by performing first signal processing on the sound signal based on the determination result, the speaker 28 that reproduces sound based on the output first sound signal, and a housing 29 that accommodates the microphone 21, the DSP 22, and the speaker 28. The DSP is an example of a signal processing circuit.

[0083] Such an ear-worn device 20 can perform signal processing by distinguishing between sound signals of speech sounds that reach the user directly and sound signals of announcement sounds.

[0084] Furthermore, for example, the DSP 22 selectively outputs the first sound signal and the second sound signal obtained by performing second signal processing on the sound signal, which is different from the first signal processing, based on the determination result. The speaker 28 reproduces sound based on one of the output first sound signal and the output second sound signal.

[0085] Such an ear-worn device 20 can perform different signal processing on the sound signals of speech sounds that reach the user directly and the sound signals of announcement sounds.

[0086] Also, for example, the first signal processing includes equalizing processing for emphasizing specific frequency components of the acquired sound, and the second signal processing includes phase inversion processing.

[0087] Such an ear-worn device 20 can emphasize one of the direct sound and the announcement sound and attenuate the other.

[0088] Also, for example, the DSP22 outputs a first sound signal when it determines that the sound contained in the sound has a reverberation feeling, and outputs a second sound signal when it determines that the sound contained in the sound does not have a reverberation feeling.

[0089] Such an ear-worn device 20 can enhance the announcement sound and attenuate the direct sound, thereby helping the user to hear the announcement sound.

[0090] Also, for example, DSP22 outputs a first sound signal when it determines that the sound contained in the sound does not have a reverberation feeling, and outputs a second sound signal when it determines that the sound contained in the sound has a reverberation feeling.

[0091] Such an ear-worn device 20 can emphasize speech sounds that reach the user directly and attenuate announcement sounds, and can help the user interact with other users who are speaking to the user.

[0092] Also, for example, the DSP22 selectively operates in an announcement mode and an interactive mode. When operating in the announcement mode, the DSP22 outputs a first sound signal if it determines that the voice included in the sound has a reverberation, and outputs a second sound signal if it determines that the voice included in the sound does not have a reverberation. When operating in the interactive mode, the DSP22 outputs the first sound signal if it determines that the voice included in the sound does not have a reverberation, and outputs the second sound signal if it determines that the voice included in the sound has a reverberation. The announcement mode is an example of a first mode, and the interactive mode is an example of a second mode.

[0093] Such an ear-worn device 20 can selectively operate in an announcement mode that emphasizes announcement sounds and attenuates speech sounds that reach the user directly, and an interactive mode that emphasizes speech sounds that reach the user directly and attenuates announcement sounds.

[0094] Furthermore, for example, the DSP22 selectively operates in an announcement mode, a dialogue mode, and a voice detection mode. During operation in the voice detection mode, the DSP22 performs signal processing on the sound signal to determine whether the sound contains voice, and outputs a first sound signal if it determines that the acquired sound contains voice, and outputs a second sound signal if it determines that the acquired sound does not contain voice. The voice detection mode is an example of a third mode.

[0095] Such an ear-worn device 20 can operate in an announcement mode, an interaction mode, and also in a voice detection mode that emphasizes human voices and attenuates noise.

[0096] Furthermore, for example, the DSP 22 performs signal processing on the sound signal to calculate the power spectrum of the reverberation sound contained in the sound, and determines whether or not the voice contained in the sound has a reverberation feeling based on the calculated power spectrum.

[0097] Such an ear-worn device 20 can determine whether or not a sound has a reverberant sound based on the power spectrum of the reverberant sound.

[0098] Furthermore, for example, the DSP22 performs signal processing on the sound signal to calculate onset information indicating the change over time in the sound pressure level of the sound signal and the onset time, and determines whether or not the sound contained in the sound has a reverberant feel based on the calculated onset information.

[0099] Such an ear-worn device 20 can determine whether or not a person's voice has a reverberant sound based on the onset information.

[0100] For example, the ear-worn device 20 further includes a mixing circuit 27b that mixes the output first sound signal with a third sound signal provided from the mobile terminal 30. The speaker 28 reproduces sound based on the first sound signal mixed with the third sound signal. The mobile terminal 30 is an example of a sound source.

[0101] Such an ear-worn device can operate in an announcement mode or the like while the third sound signal is being played.

[0102] In addition, the playback method executed by a computer such as the ear-worn device 20 includes a determination step S26 in which signal processing is performed on the sound signal of the sound output by the microphone that acquires the sound to determine whether the voice contained in the sound has a reverberant feel, an output step S27 in which a first sound signal obtained by performing first signal processing on the sound signal is output based on the determination result in the determination step S26, and a playback step S30 in which sound is played based on the output first sound signal.

[0103] Such a reproduction method can perform signal processing by distinguishing between the sound signal of the speech sound that reaches the user directly and the sound signal of the announcement sound.

[0104] (Other embodiments) Although the embodiments have been described above, the present disclosure is not limited to the above-described embodiments.

[0105] For example, in the above embodiment, the ear-worn device is described as an earphone-type device, but it may also be a headphone-type device. Also, in the above embodiment, the ear-worn device selectively operates in three operation modes, but it may be a device having at least one of the three operation modes, or a device specialized in any one of the three operation modes.

[0106] In the above embodiment, the ear-worn device has a function of playing music content, but it does not have to have a function of playing music content (communication module). For example, the ear-worn device may be earplugs with a noise cancellation function and an external sound capture function.

[0107] In the above embodiment, the determination of whether or not the sound captured by the microphone contains speech is performed using a machine learning model, but the determination may be performed based on other algorithms that do not use a machine learning model. The same applies to the determination of whether or not speech has reverberation.

[0108] The configuration of the ear-worn device according to the above embodiment is an example, and the ear-worn device may include components not shown, such as a D / A converter, a filter, a power amplifier, or an A / D converter.

[0109] Furthermore, in the above embodiment, the sound signal processing system is realized by multiple devices, but it may also be realized as a single device. When the sound signal processing system is realized by multiple devices, the functional components of the sound signal processing system may be distributed in any manner among the multiple devices. For example, in the above embodiment, some or all of the functional components of the ear-worn device may be provided in a mobile terminal.

[0110] Furthermore, the communication method between the devices in the above-described embodiments is not particularly limited. When two devices communicate with each other in the above-described embodiments, a relay device (not shown) may be interposed between the two devices.

[0111] The order of the processes described in the above embodiments is merely an example. The order of multiple processes may be changed, or multiple processes may be executed in parallel. Furthermore, a process executed by a specific processing unit may be executed by another processing unit. Furthermore, part of the digital signal processing described in the above embodiments may be realized by analog signal processing.

[0112] In the above-described embodiments, each component may be realized by executing a software program suitable for that component, or by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.

[0113] Furthermore, each component may be realized by hardware. For example, each component may be a circuit (or integrated circuit). These circuits may form a single circuit as a whole, or each may be a separate circuit. Furthermore, each of these circuits may be a general-purpose circuit or a dedicated circuit.

[0114] Furthermore, the general or specific aspects of the present disclosure may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM. Also, any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium may be realized. For example, the present disclosure may be implemented as a playback method executed by a computer such as an ear-worn device or a mobile terminal, or may be realized as a program for causing a computer to execute such a playback method. Also, the present disclosure may be realized as a computer-readable non-transitory recording medium on which such a program is recorded. Note that the program here includes an application program for causing a general-purpose mobile terminal to function as the mobile terminal of the above-mentioned embodiment.

[0115] In addition, this disclosure also includes forms obtained by applying various modifications to each embodiment that a person skilled in the art would think of, or forms realized by arbitrarily combining the components and functions of each embodiment within the scope that does not deviate from the intent of this disclosure. [Industrial Applicability]

[0116] The ear-worn device of the present disclosure can perform signal processing by distinguishing between sound signals of sounds with relatively strong direct sound components and sound signals of sounds with relatively strong indirect sound components. [Explanation of symbols]

[0117] 10 Sound signal processing system 20 Ear-worn devices 21 Microphone 22 DSP 23 Filter section 23a High-pass filter 23b Low-pass filter 23c Bandpass Filter 24 Signal Processing Section 24a Reverberation detection section 24b Noise detection section 24c Voice detection section 24d Switching section 25 Neural Network Department 25a Audio detection unit 25b Reverberation determination section 26 Memory section 27 Communication Module 27a Communication circuit 27b Mixing circuit 28 Speaker 29 Housing 30 Mobile Devices 31 UI section 32 Communication Circuit 33 Information Processing Department 34 Storage section

Claims

1. a microphone that acquires sound and outputs a sound signal of the acquired sound; a signal processing circuit that determines whether or not a voice included in the sound is an announcement sound, and outputs a second sound signal obtained by performing second signal processing including phase inversion processing on the sound signal based on the determination result of whether or not the voice included in the sound is an announcement sound; a speaker that reproduces sound based on the output second sound signal; a housing that accommodates the microphone, the signal processing circuit, and the speaker; Ear-worn device.

2. The announcement sound is a sound other than a sound that directly reaches the microphone. The ear-worn device of claim 1 .

3. The announcement sound is a sound in which an indirect sound component is relatively strong compared to a direct sound component, The direct sound is a sound that reaches the microphone directly without being reflected from a sound source, and the indirect sound is a sound that reaches the microphone after being reflected from the sound source one or more times.

3. The ear-worn device according to claim 1 or 2.

4. When the signal processing circuit determines that the voice is an announcement sound, it outputs the second sound signal. The ear-worn device of claim 1 .

5. When the signal processing circuit determines that the sound is not an announcement sound, the signal processing circuit outputs a first sound signal obtained by performing first signal processing on the sound signal, the first signal processing including equalization processing for emphasizing a frequency component of the sound; The speaker reproduces a sound based on one of the output first sound signal and the output second sound signal. The ear-worn device of claim 4.

6. When the signal processing circuit determines that the voice is not an announcement sound, it outputs the second sound signal. The ear-worn device of claim 1 .

7. When the signal processing circuit determines that the sound is an announcement sound, the signal processing circuit outputs a first sound signal obtained by performing first signal processing on the sound signal, the first signal processing including equalization processing for emphasizing a frequency component of the sound; The speaker reproduces a sound based on one of the output first sound signal and the output second sound signal. The ear-worn device of claim 6.

8. The signal processing circuit performs signal processing on the sound signal to calculate a power spectrum of reverberation sound included in the sound, and determines whether the sound is an announcement sound based on the calculated power spectrum. The ear-worn device according to any one of claims 1 to 7.

9. The signal processing circuit processes the sound signal to calculate onset information indicating a time-dependent change in the sound pressure level of the sound signal and an onset time, and determines whether the sound is an announcement sound based on the calculated onset information. The ear-worn device according to any one of claims 1 to 7.

10. further comprising a mixing circuit that mixes the output second sound signal with a third sound signal provided from a sound source; The speaker reproduces sound based on the second sound signal mixed with the third sound signal. The ear-worn device according to any one of claims 1 to 9.

11. a determining step of determining whether or not a voice included in the sound acquired by the microphone is an announcement sound; an output step of outputting a second sound signal obtained by performing second signal processing including phase inversion processing on the sound signal of the sound based on a determination result of whether or not the voice included in the sound is an announcement sound in the determination step; and a reproduction step of reproducing sound based on the output second sound signal. How to play.

12. A program for causing a computer to execute the reproducing method according to claim 11.

Citation Information

Patent Citations

  • Earphone

    JP2012249184A