Wearable device, reproduction method, and program

The wearable device accurately distinguishes and processes direct speech sounds by detecting sound direction, ensuring proper reproduction and preventing misclassification.

WO2026009783A1PCT designated stage Publication Date: 2026-01-08PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/022812
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-05
Filing Date
2025-06-25
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Existing wearable devices fail to accurately distinguish between speech sounds that reach the user directly and those that do not, leading to potential misclassification and inappropriate signal processing.

Method used

A wearable device equipped with a microphone, signal processing circuit, and speaker that detects the direction of sound arrival and applies appropriate signal processing to emphasize or remove frequency components based on this detection, ensuring accurate reproduction of direct speech sounds.

Benefits of technology

Enables accurate reproduction of direct speech sounds by emphasizing or removing frequency components as needed, preventing misclassification and improving sound quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025022812_08012026_PF_FP_ABST
    Figure JP2025022812_08012026_PF_FP_ABST
Patent Text Reader

Abstract

A wearable device 20 includes: a microphone 21 that acquires sound and outputs a sound signal of the acquired sound; a signal processing circuit 22 that detects the incoming direction of the sound, determines whether or not the sound has directly reached the microphone 21 on the basis of the detection result, and, when it is determined that the sound has directly reached the microphone 21, outputs a first sound signal obtained by subjecting the sound to first signal processing including equalizing processing for emphasizing a frequency component of a voice included in the sound; and a speaker 28 that reproduces a reproduction sound expressed by the output first sound signal.
Need to check novelty before this filing date? Find Prior Art

Description

Wearable device, playback method, and program

[0001] The present disclosure relates to techniques for reproducing sound in a wearable device.

[0002] Patent Document 1 discloses a technology for determining whether or not the sound contained in a speech sound has a reverberant sound, and, based on the determination result, performing signal processing by distinguishing between the sound signal of the speech sound that reaches the user directly and the sound signal of the announcement sound.

[0003] However, in Patent Document 1, the direction from which the speech sound comes is not detected to distinguish whether the speech sound is reaching the user directly or not, so there is a risk that a sound signal of a speech sound without a sense of reverberation, such as the speech sound of two people talking facing each other near the user, may be processed as a sound signal of a speech sound reaching the user directly.

[0004] International Publication No. 2022 / 137806

[0005] An object of the present disclosure is to provide a wearable device, a playback method, and a program that can appropriately play back speech sounds that reach the user directly.

[0006] A wearable device in one aspect of the present disclosure includes a microphone that acquires sound and outputs a sound signal of the acquired sound; a signal processing circuit that detects the direction from which the sound is coming and determines whether the sound has arrived directly at the microphone based on the detection result, and if it is determined that the sound has arrived directly at the microphone, outputs a first sound signal that has been subjected to first signal processing on the sound signal, including equalization processing to emphasize frequency components of the audio contained in the sound; and a speaker that reproduces a sound indicated by the output first sound signal.

[0007] FIG. 1 is a block diagram showing the functional configuration of a sound signal processing system according to an embodiment; FIG. 2 is a diagram showing an example of two sounds arriving from different directions; FIG. 3 is a diagram showing an example of a MUSIC spectrum; FIG. 4 is a diagram showing an example of a speech sound reaching a user; FIG. 5 is a diagram showing another example of a MUSIC spectrum; FIG. 6 is a diagram showing an example of a MUSIC spectrum used in machine learning for constructing a determination model; FIG. 7 is a diagram showing a first example of an image showing a recurrence plot of sound; FIG. 8 is a diagram showing a second example of an image showing a recurrence plot of sound; and FIG. 9 is a flowchart showing sound playback processing in a wearable device.

[0008] (Background to One Aspect of the Present Disclosure) The present inventors have been studying improvements to the ambient sound capture function of a wearable device, which enhances specific frequency components of sounds contained in ambient sounds, such as speech sounds and announcements, that reach the user directly. In the course of their studies, the present inventors proposed a technology in Patent Literature 1 that determines whether or not the sounds contained in speech sounds have reverberation, and, based on the determination result, performs signal processing by distinguishing between sound signals of speech sounds that reach the user directly (sounds heard when someone speaks to the user) and sound signals of announcements. However, this technology does not detect the direction from which the speech sounds arrive. Therefore, there is a risk that sound signals of speech sounds that do not have reverberation, such as the sound signals of two people talking face to face near the user, may be processed as sound signals of speech sounds that reach the user directly.

[0009] Based on the above findings, the inventors have conducted extensive research into technologies for appropriately reproducing speech sounds that reach the user directly in a wearable device, and have come up with the following aspects of the present disclosure.

[0010] (1) A wearable device according to one aspect of the present disclosure includes a microphone that acquires sound and outputs a sound signal of the acquired sound; a signal processing circuit that detects the direction from which the sound is coming and determines whether the sound has arrived directly at the microphone based on the detection result; and, if it is determined that the sound has arrived directly at the microphone, outputs a first sound signal that has been subjected to first signal processing on the sound signal, including equalization processing to emphasize frequency components of the audio contained in the sound; and a speaker that reproduces a sound indicated by the output first sound signal.

[0011] In this configuration, whether a sound acquired by the microphone has arrived directly at the microphone is determined based on the detection result of the direction from which the sound has arrived. Therefore, when the microphone acquires a speech sound that reaches the user directly, it can be appropriately determined that the speech sound has arrived directly at the microphone. Furthermore, with this configuration, a reproduced sound indicated by a first sound signal in which the frequency components of the sound contained in the speech sound have been emphasized by the first signal processing is reproduced, so that the speech sound can be appropriately reproduced.

[0012] (2) In another aspect of the present disclosure, a wearable device includes a microphone that acquires sound and outputs a sound signal of the acquired sound; a signal processing circuit that detects the direction from which the sound is coming and determines whether the sound has arrived directly at the microphone based on the detection result; if it is determined that the sound has arrived directly at the microphone, outputs the sound signal as a first sound signal; and if it is determined that the sound has not arrived directly at the microphone, outputs a second sound signal that has been obtained by performing second signal processing on the sound signal, including removing frequency components of the audio contained in the sound; and a speaker that reproduces a sound indicated by one of the output first sound signal and the second sound signal.

[0013] In this configuration, whether a sound acquired by the microphone has arrived directly at the microphone is determined based on the detection result of the direction from which the sound has arrived. Therefore, when the microphone acquires a speech sound that reaches the user directly, it can be appropriately determined that the speech sound has arrived directly at the microphone. In this case, a playback sound indicated by a first sound signal, which is a sound signal of the speech sound, is reproduced, so that the speech sound can be appropriately reproduced.

[0014] On the other hand, when the microphone acquires a sound that is not a speech sound that directly reaches the user, such as the speech sound of two people talking face to face near the user, it is possible to appropriately determine that the sound is not directly reaching the microphone. In this case, the reproduced sound indicated by the second sound signal from which the frequency components of the sound contained in the sound have been removed by the second signal processing is reproduced, thereby preventing the sound contained in the sound from being reproduced.

[0015] (3) In the wearable device described in (1) above, when the signal processing circuit determines that the sound is not arriving directly at the microphone, it may output a second sound signal obtained by performing second signal processing on the sound signal, including removing frequency components of the voice contained in the sound, and the speaker may reproduce a playback sound indicated by one of the output first sound signal and the output second sound signal.

[0016] According to this configuration, when the microphone acquires a sound that is not a speech sound that directly reaches the user, such as the speech sound of two people talking face to face near the user, it is possible to appropriately determine that the sound did not directly reach the microphone. In this case, a reproduced sound indicated by the second sound signal from which the frequency components of the sound contained in the sound have been removed by the second signal processing is reproduced, thereby preventing the sound contained in the sound from being reproduced.

[0017] (4) In the wearable device described in (2) or (3) above, the second signal processing may include phase inversion processing.

[0018] In this configuration, when the microphone picks up a sound that is not a speech sound that directly reaches the user, a playback sound represented by a second sound signal in which the frequency components of the sound contained in the picked-up sound are cancelled out by phase inversion processing is reproduced, thereby making it possible to suppress the playback of the sound contained in the picked-up sound.

[0019] (5) In the wearable device described in any one of (1) to (4) above, the signal processing circuit may perform signal processing on the sound signal to detect the MUSIC spectrum of the sound as information indicating the direction from which the sound is coming, and determine whether the sound has arrived directly at the microphone based on the MUSIC spectrum of the sound.

[0020] According to this configuration, the MUSIC spectrum of sound, which is known as information that can detect the direction from which the sound is coming, can be used to appropriately determine whether the sound captured by the microphone arrived directly at the microphone.

[0021] (6) In the wearable device described in (5) above, the signal processing circuit may determine that the sound has arrived directly at the microphone if the peak width in the MUSIC spectrum of the sound is less than a predetermined width, and may determine that the sound has not arrived directly at the microphone if the peak width is greater than or equal to the predetermined width.

[0022] According to this configuration, it is possible to appropriately determine whether or not a sound has arrived directly at the microphone depending on whether or not the peak width in the MUSIC spectrum of the sound acquired by the microphone is less than a predetermined width.

[0023] (7) In the wearable device described in (5) above, the signal processing circuit may determine whether the sound has arrived directly at the microphone based on a predetermined range of spectrum including a peak portion of the MUSIC spectrum of the sound.

[0024] In this configuration, whether or not a sound has directly reached the microphone is determined based on a predetermined range of the spectrum including the peak portion of the MUSIC spectrum of the sound, rather than the entire MUSIC spectrum of the sound captured by the microphone, which reduces the processing load for determining whether or not a sound has directly reached the microphone.

[0025] (8) In the wearable device described in (5) or (7) above, the signal processing circuit may determine whether the sound arrived directly at the microphone using a learning model trained based on the MUSIC spectrum of the sound.

[0026] According to this configuration, it is possible to appropriately determine whether or not a sound has arrived directly at the microphone using the learning model.

[0027] (9) In the wearable device described in any one of (1) to (8) above, the signal processing circuit may further perform signal processing on the sound signal to determine whether the sound contains voice, and output the first sound signal if it determines that the sound has arrived directly at the microphone and if it determines that the sound contains voice.

[0028] In this configuration, when the sound that directly reaches the microphone includes speech, the playback sound indicated by the first sound signal is reproduced. Therefore, when a speech sound that directly reaches the user reaches the microphone, the speech sound can be appropriately reproduced.

[0029] (10) In the wearable device described in (9) above, the signal processing circuit may further determine whether the audio contained in the sound includes audio indicating a predetermined word, and output the first sound signal if it is determined that the audio contained in the sound includes audio indicating the predetermined word.

[0030] In this configuration, when the sound that directly reaches the microphone includes a voice representing a predetermined word, the reproduced sound indicated by the first sound signal is reproduced. Therefore, when a speech sound that directly reaches the user and includes a voice representing a predetermined word reaches the microphone, the speech sound can be reproduced appropriately.

[0031] (11) In the wearable device described in (10) above, the signal processing circuit may further perform signal processing on the sound signal to calculate the periodicity of the sound contained in the sound, and determine whether the sound contained in the sound includes a sound indicating the specified word based on the periodicity of the sound contained in the sound.

[0032] According to this configuration, it is possible to appropriately determine whether the audio contained in the sound that arrives directly at the microphone includes audio that indicates a specified word, based on the periodicity of the audio contained in the sound.

[0033] (12) In the wearable device described in (9) above, the signal processing circuit may further perform signal processing on the sound signal to calculate the periodicity of the sound, and determine whether the sound includes speech based on the periodicity of the sound.

[0034] According to this configuration, it is possible to appropriately determine whether or not a sound that directly reaches a microphone includes speech, based on the periodicity of the sound.

[0035] (13) In the wearable device described in (11) above, the periodicity of the sound contained in the sound may be represented by a recurrence plot, and the signal processing circuit may use a learning model trained based on the recurrence plot to determine whether the sound contained in the sound includes a sound indicating the specified word.

[0036] According to this configuration, it is possible to use the learning model to appropriately determine whether or not the speech contained in the sound that directly arrives at the microphone includes speech indicating a predetermined word.

[0037] (14) In the wearable device described in any one of (9) to (13) above, the signal processing circuit may output the first sound signal if the point at which it determines that the sound did not arrive directly at the microphone or that the sound does not contain speech is the point at which it most recently determined that the sound arrived directly at the microphone and a predetermined time has not elapsed since it determined that the sound contained speech.

[0038] In this configuration, the sound indicated by the first sound signal is reproduced for a predetermined time from the time when it is determined that the sound has arrived directly at the microphone and that the sound includes speech, so that the speech sound can be reproduced continuously as long as the speech sound that arrives directly at the user is not interrupted for a predetermined time or longer.

[0039] (15) The wearable device according to any one of (1) to (14) above may further include a mixing circuit that mixes the output first sound signal with a third sound signal provided from a sound source, and the speaker may reproduce a playback sound indicated by the first sound signal mixed with the third sound signal.

[0040] In this configuration, when the microphone captures a speech sound that directly reaches the user, a playback sound represented by the first sound signal mixed with the third sound signal provided from the sound source is reproduced, so that the user can listen to the speech that directly reaches the user while listening to the sound represented by the third sound signal.

[0041] The present disclosure can be realized not only as a wearable device having a characteristic configuration corresponding to the characteristic processing described above, but also as a playback method for executing the characteristic processing described above. Furthermore, it can also be realized as a computer program that causes a computer to execute the characteristic processing included in such a playback method. Therefore, the same effects as those of the above-described wearable device can be achieved in the following other aspects.

[0042] (16) In another aspect of the present disclosure, a playback method is a playback method in a wearable device that includes a microphone that acquires sound and outputs a sound signal of the acquired sound, a signal processing circuit that performs signal processing on the sound signal, and a speaker, wherein the processing performed by the signal processing circuit includes detecting the direction from which the sound is coming indicated by the sound signal output by the microphone, determining whether the sound has arrived directly at the microphone based on the detection result, and if it is determined that the sound has arrived directly at the microphone, outputting a first sound signal obtained by performing first signal processing on the sound signal, including equalizing processing to emphasize frequency components of the audio contained in the sound, and causing the speaker to reproduce the playback sound indicated by the output first sound signal.

[0043] (17) In another aspect of the present disclosure, a program for a wearable device includes a microphone that acquires sound and outputs a sound signal of the acquired sound, a signal processing circuit that performs signal processing on the sound signal, and a speaker, and causes the signal processing circuit to: detect the direction from which the sound is coming indicated by the sound signal output by the microphone; determine whether the sound has arrived directly at the microphone based on the detection result; and, if it is determined that the sound has arrived directly at the microphone, output a first sound signal obtained by performing first signal processing on the sound signal, including equalizing processing to emphasize the frequency components of the audio contained in the sound; and cause the speaker to reproduce the playback sound indicated by the output first sound signal.

[0044] The present disclosure can also be realized as an information processing system operated by such a program. Needless to say, such a program can be distributed on a computer-readable non-transitory recording medium such as a CD-ROM or via a communication network such as the Internet.

[0045] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. Note that each of the embodiments described below represents a specific example of the present disclosure. The numerical values, shapes, components, steps, order of steps, etc. shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components that are not described in an independent claim that represents a superordinate concept will be described as optional components. Furthermore, in all embodiments, the respective contents can be combined.

[0046] (Embodiment) The configuration of a sound signal processing system 10 according to an embodiment will be described. Fig. 1 is a block diagram showing the functional configuration of the sound signal processing system 10 according to the embodiment. The sound signal processing system 10 according to the embodiment includes a wearable device 20 and an information terminal 30. The wearable device 20 is an earphone-type device that reproduces a sound signal provided from the information terminal 30. The sound signal is, for example, a sound signal of music content.

[0047] The wearable device 20 has a noise cancellation function that reduces environmental sounds (noises) around the user wearing the wearable device 20. The wearable device 20 also has an ambient function that emphasizes speech sounds that reach the user directly. Here, speech sounds that reach the user directly refer to sounds that the user hears when someone speaks to them directly. Therefore, speech sounds that the user hears when two people are talking face to face near the user are not speech sounds that reach the user directly.

[0048] The wearable device 20 is not limited to an earphone-type device, but may be a headphone-type or goggle-type device worn on the ear. Furthermore, the wearable device 20 may be a body-worn device such as a neck speaker-type device worn on the shoulder, a goggle-type device worn on the head, or a wearable type device worn on clothing or any part of the body, as long as it has at least an ambient function.

[0049] Specifically, the wearable device 20 includes a plurality of microphones 21, a DSP 22 (signal processing circuit), a communication module 27, and a speaker 28. The plurality of microphones 21, the DSP 22, the communication module 27, and the speaker 28 are accommodated in a housing (not shown).

[0050] The microphone 21 acquires sounds around the wearable device 20 and outputs sound signals of the acquired sounds. The multiple microphones 21 are arranged in a row inside the housing. Specifically, the microphones 21 include condenser microphones, dynamic microphones, MEMS (Micro Electro Mechanical Systems) microphones, etc. However, the microphones 21 are not limited to these.

[0051] The DSP 22 performs signal processing on the sound signal output from the microphone 21 to realize a noise cancellation function and an ambient function. The noise cancellation function is a function for reducing noise around the wearable device 20. Specifically, the noise cancellation function is realized by performing a phase inversion process (second signal processing) that inverts the phase of the sound signal, and outputting the phase-inverted sound signal (second sound signal) from the speaker 28.

[0052] The ambient function is a function that emphasizes speech sounds that reach the user directly. Speech sounds that reach the user directly are speech sounds that are spoken directly to the user and arrive directly at the user. Specifically, the ambient function is realized by performing an equalization process to emphasize specific frequency components on the sound signal output from the microphone 21, and outputting the equalized sound signal from the speaker 28. The specific frequency components can be frequency components of human voice (for example, frequency components between 100 Hz and 2 kHz). Note that the equalization process is not essential for the ambient function. The ambient function may also be realized by outputting the sound signal of the speech sounds that reach the user directly from the speaker 28 substantially as is.

[0053] Specifically, the DSP 22 includes a storage unit 26 , a signal processing unit 24 , and a determination unit 25 .

[0054] The storage unit 26 stores a computer program executed by a circuit corresponding to the signal processing unit 24 and a computer program executed by a circuit corresponding to the determination unit 25. The storage unit 26 stores various information necessary for implementing the noise cancellation function and the ambient function, such as a determination model 261. The storage unit 26 is realized by a semiconductor memory or the like. The storage unit 26 may also be realized as an external memory of the DSP 22.

[0055] The determination model 261 is a learning model used by the determination unit 25 to determine whether a sound has arrived directly at the microphone 21 and whether the sound contains speech.

[0056] Specifically, the judgment model 261 is constructed by machine learning using a predetermined algorithm such as CNN on the relationship between the MUSIC (Multiple Signal Classification) spectrum, MFCC (Mel-Frequency Cepstral Coefficient), and recurrence plot of the sound and whether the sound has arrived directly at the microphone 21, and the relationship between the MUSIC spectrum, MFCC, and recurrence plot of the sound and whether the sound contains speech.

[0057] The following describes the MUSIC spectrum, MFCC, and recurrence plot of a sound used in machine learning to construct the determination model 261. The MUSIC spectrum of a sound is information indicating the relationship between the direction of arrival of the sound and the intensity of the sound signal of the sound. The MUSIC spectrum of a sound is obtained by signal processing the sound signal of the sound using the MUSIC method, which is one of the signal analysis processing techniques.

[0058] FIG. 2 is a diagram showing an example of two sounds arriving from different directions. FIG. 2 shows an example in which a sound output from a sound source V21 facing two microphones 21 arrives at the two microphones 21. FIG. 2 also shows an example in which a sound output from a sound source V22 located in a direction rotated clockwise by an angle θ from the front direction of the two microphones 21 arrives at the two microphones 21. FIG. 3 is a diagram showing an example of a MUSIC spectrum. The horizontal axis of FIG. 3 indicates the angle rotated clockwise from the front direction of the two microphones 21. The vertical axis of FIG. 3 indicates the intensity (spectrum) of each angular component of the sound signal of the sound acquired by the microphone 21. Reference symbol G31 in FIG. 3 indicates an example of a MUSIC spectrum of a sound arriving at the microphone 21 from the sound source V21 in FIG. 2. Reference symbol G32 in FIG. 3 indicates an example of a MUSIC spectrum of a sound arriving at the microphone 21 from the sound source V22 in FIG. 2.

[0059] In the MUSIC method, a predetermined signal processing is performed on a sound signal to calculate a MUSIC spectrum of the sound represented by the sound signal, and the direction corresponding to the angle of the peak of the MUSIC spectrum is detected as the arrival direction of the sound. In the example shown in Figures 2 and 3, the peak of the MUSIC spectrum G31 of the sound arriving at the microphone 21 from the sound source V21 occurs when the angle of clockwise rotation from the front direction of the microphone 21 is 0 degrees. Therefore, the front directions of the two microphones 21 are detected as the arrival direction of the sound. On the other hand, the peak of the MUSIC spectrum G32 of the sound arriving at the microphone 21 from the sound source V22 occurs when the angle of clockwise rotation from the front direction of the microphone 21 is θ. Therefore, the direction rotated by the angle θ clockwise from the front direction of the two microphones 21 is detected as the arrival direction of the sound.

[0060] Fig. 4 is a diagram showing an example of speech sound reaching the user 42. The left diagram of Fig. 4 shows an example in which speech sound including a sound indicating the words "Hello" spoken to the user 42 by a speaker 41 around the user 42 directly reaches the wearable device 20 worn by the user 42. In other words, the left diagram of Fig. 4 shows an example of speech sound directly reaching the user 42.

[0061] The right diagram of Fig. 4 shows an example in which speech sound including a sound indicating the word "Hello" spoken by a speaker 41 near a user 42 toward the direction of an object 43 different from the user 42 arrives at the wearable device 20 worn by the user 42, both directly and after being reflected by the object 43. Hereinafter, speech sound that is spoken indirectly to the user 42 and reaches the user 42 directly and indirectly, as shown in the right diagram of Fig. 4, will be referred to as speech sound that reaches the user 42 indirectly.

[0062] Fig. 5 is a diagram showing another example of a MUSIC spectrum. Reference symbol G51 in Fig. 5 indicates the MUSIC spectrum of the speech sound that directly reaches the user 42 shown in the left diagram of Fig. 4. Reference symbol G52 in Fig. 5 indicates the MUSIC spectrum of the speech sound that indirectly reaches the user 42 shown in the right diagram of Fig. 4.

[0063] The peaks of MUSIC spectrum G51 and MUSIC spectrum G52 both occur at an angle of 0 degrees clockwise rotation from the front direction of microphone 21. However, the peak width of MUSIC spectrum G51 is narrower than the peak width of MUSIC spectrum G52. The peak width of the MUSIC spectrum is the difference between two angles corresponding to an intensity in the MUSIC spectrum that is a predetermined intensity (e.g., 6 dB) lower than the peak.

[0064] Thus, simply detecting the direction of arrival of a speech sound from the MUSIC spectrum of the speech sound may not adequately determine whether the speech sound is a speech sound that reaches the user 42 directly or a speech sound that reaches the user 42 indirectly. However, by further understanding the peak width of the MUSIC spectrum of the speech sound, it is possible to determine whether the speech sound is a speech sound that reaches the user 42 directly.

[0065] 6 is a diagram showing an example of a MUSIC spectrum used in machine learning to construct the determination model 261. Therefore, for example, as shown in FIG. 6 , a predetermined range of spectrum 61 including a peak portion of the MUSIC spectrum of the speech sound is used in the machine learning to construct the determination model 261. The predetermined range of spectrum 61 is a spectrum corresponding to an angle range ΔD in the MUSIC spectrum from an angle that is a predetermined angle smaller than the angle indicating the peak to an angle that is the predetermined angle larger than the angle indicating the peak. This allows the processing load of the machine learning to construct the determination model 261 to be reduced compared to when the entire MUSIC spectrum is used for the machine learning.

[0066] The MUSIC spectrum used in the machine learning to construct the determination model 261 may be an image showing the MUSIC spectrum as shown in FIG. 6, or may be two-dimensional array data showing the MUSIC spectrum.

[0067] The MFCC of a sound is information obtained by signal processing the sound signal of the sound, and is obtained by quantifying the timbre of the sound based on the Mel frequency scale. The Mel frequency scale is a frequency scale designed based on the characteristics of human hearing. The MFCC of a sound used in machine learning to construct the determination model 261 may be an image showing the time-series changes of each Mel frequency component of the sound, or may be array data of numerical values ​​showing the time-series changes.

[0068] A sound recurrence plot is information obtained by signal processing the sound signal of the sound, visualizing the periodicity of the sound in a two-dimensional image with the time axis as the second axis. FIG. 7 is a diagram showing a first example of an image showing a sound recurrence plot, and FIG. 8 is a diagram showing a second example of an image showing a sound recurrence plot. The recurrence plot shown in FIG. 7 is more uniform than the recurrence plot shown in FIG. 8, and many diagonal lines indicating periodic behavior are observed. Furthermore, the recurrence plot shown in FIG. 7 has black dots distributed more evenly throughout and at a higher density than the recurrence plot shown in FIG. 8. These facts suggest that the sound corresponding to the recurrence plot shown in FIG. 7 has stronger periodicity than the sound corresponding to the recurrence plot shown in FIG. 8.

[0069] For example, a recurrence plot of speech sounds that are spoken directly to a user, such as repeating the user's name, will be a recurrence plot that suggests the presence of strong periodicity, as shown in Figure 7. On the other hand, a recurrence plot of speech sounds that are not spoken directly to the user, such as conversations taking place around the user, will be a recurrence plot that shows many complex and irregular patterns, as shown in Figure 8.

[0070] The sound recurrence plot used for machine learning to construct the judgment model 261 may be a two-dimensional image that visualizes the periodicity of the sound, as shown in Figures 7 and 8, or it may be two-dimensional array data that shows the image.

[0071] The MFCC and recurrence plot of a sound can be used as information that uniquely identifies the sound. For example, the MFCC and recurrence plot of a sound will be different for a sound that includes a sound indicating a word (a wake word (a predetermined word)) used when speaking directly to a user, such as the word "Hello," and a sound that does not include the sound. Therefore, by using the MFCC and recurrence plot of a sound in machine learning to build the determination model 261, it is possible to build a determination model 261 with higher determination accuracy than when the MFCC and recurrence plot of a sound are not used.

[0072] When the judgment model 261 receives a spectrum of a predetermined range including the peak portion of the MUSIC spectrum of the sound, the MFCC of the sound, and a recurrence plot of the sound as input, it outputs information indicating whether the sound arrived directly at the microphone 21 and information indicating whether the sound contains speech.

[0073] The signal processing unit 24 includes an utterance detection unit 241 and a switching unit 242. The functions of the utterance detection unit 241 and the switching unit 242 are realized, for example, by a circuit corresponding to the signal processing unit 24 executing a computer program stored in the storage unit 26.

[0074] The speech detection unit 241 detects a sound signal indicating a speech sound (hereinafter, referred to as a speech sound signal) from the sound signal output by the microphone 21 , and outputs the detected speech sound signal to the determination unit 25 .

[0075] Specifically, the speech detection unit 241 detects a speech sound signal by performing signal-to-noise ratio (SN ratio) based voice activity detection (VAD) processing on the sound signal output by the microphone 21, and outputs the detected speech sound signal.

[0076] As a result, a speech sound signal from which noise different from speech sounds has been removed is output to the determination unit 25. Note that the speech detection unit 241 is not limited to signal-to-noise ratio-based voice activity detection processing, and may detect a speech sound signal from the sound signal output by the microphone 21 by other signal processing.

[0077] The switching unit 242 switches the operation mode between the ambient mode and the noise cancellation mode based on the determination result by the determination unit 25. The ambient mode is an operation mode that enables an ambient function to emphasize speech sounds that reach the user directly. The noise cancellation mode is an operation mode that enables the noise cancellation function to reduce ambient noise.

[0078] Specifically, when the determination unit 25 determines that sound has directly reached the microphone 21 and that the sound includes speech, the switching unit 242 switches the operation mode to the ambient mode. When the operation mode is switched to the ambient mode, the switching unit 242 performs equalization processing (first signal processing) to emphasize specific frequency components on the sound signal output from the microphone 21. The switching unit 242 outputs the sound signal after the equalization processing (hereinafter referred to as the first sound signal) to the mixing circuit 272. The specific frequency components may be frequency components of human speech (for example, frequency components between 100 Hz and 2 kHz).

[0079] On the other hand, if the determination unit 25 determines that sound has not directly reached the microphone 21, the switching unit 242 switches the operation mode to the noise cancellation mode. Furthermore, if the determination unit 25 determines that sound has directly reached the microphone 21 but that the sound includes speech, the switching unit 242 switches the operation mode to the noise cancellation mode. When the operation mode is switched to the noise cancellation mode, the switching unit 242 performs a phase inversion process on the sound signal output by the microphone 21. The switching unit 242 outputs the sound signal after the phase inversion process (hereinafter, referred to as the second sound signal) to the mixing circuit 272.

[0080] The determination unit 25 includes a direction determination unit 251 and a sound determination unit 252. The functions of the direction determination unit 251 and the sound determination unit 252 are realized, for example, by a circuit corresponding to the determination unit 25 executing a computer program stored in the storage unit 26.

[0081] The direction determination unit 251 performs signal processing on the speech sound signal output by the speech detection unit 241 to detect the direction from which the sound indicated by the speech sound signal is coming, and determines whether the sound has arrived directly at the microphone 21 based on the detection result.

[0082] Specifically, the direction determination unit 251 performs signal processing on the speech sound signal using the MUSIC method to calculate (detect) the MUSIC spectrum of the sound indicated by the speech sound signal as information indicating the direction from which the sound arrived. Based on the calculated MUSIC spectrum of the sound, the direction determination unit 251 determines whether the sound arrived directly at the microphone 21.

[0083] In detail, the direction determination unit 251 further calculates the MFCC and recurrence plot of the sound indicated by the speech sound signal by performing signal processing on the speech sound signal input from the speech detection unit 241. The direction determination unit 251 acquires a determination model 261 stored in the storage unit 26. The direction determination unit 251 inputs, into the acquired determination model 261, a spectrum of a predetermined range including a peak portion of the MUSIC spectrum of the calculated sound, and the MFCC and recurrence plot of the sound indicated by the speech sound signal.

[0084] As a result, when the determination model 261 outputs information indicating that sound has arrived directly at the microphone 21, the direction determination unit 251 determines that the sound indicated by the speech sound signal has arrived directly at the microphone 21. On the other hand, when the determination model 261 outputs information indicating that sound has not arrived directly at the microphone 21, the direction determination unit 251 determines that the sound indicated by the speech sound signal has not arrived directly at the microphone 21.

[0085] The voice determination unit 252 performs signal processing on the speech sound signal output by the speech detection unit 241 to determine whether or not the sound indicated by the speech sound signal includes speech.

[0086] Specifically, similar to the direction determination unit 251, the voice determination unit 252 calculates the MUSIC spectrum, MFCC, and recurrence plot of the sound indicated by the speech sound signal, and acquires the determination model 261 stored in the storage unit 26. The voice determination unit 252 inputs, into the acquired determination model 261, a spectrum of a predetermined range including a peak portion of the calculated MUSIC spectrum of the sound, and the MFCC and recurrence plot of the sound indicated by the speech sound signal.

[0087] As a result, when the determination model 261 outputs information indicating that the sound contains speech, the speech determination unit 252 determines that the sound indicated by the speech sound signal contains speech. On the other hand, when the determination model 261 outputs information indicating that the sound does not contain speech, the speech determination unit 252 determines that the sound indicated by the speech sound signal does not contain speech.

[0088] The communication module 27 receives a sound signal from the information terminal 30. The communication module 27 mixes the received sound signal with a processed sound signal output by the DSP 22, and outputs the mixed signal to the speaker .

[0089] Specifically, the communication module 27 includes a communication circuit 271 and a mixing circuit 272. The communication circuit 271 receives an audio signal from the information terminal 30. The communication circuit 271 is, for example, a wireless communication circuit, and communicates with the information terminal 30 based on a communication standard such as Bluetooth (registered trademark) or BLE (Bluetooth (registered trademark) Low Energy). The mixing circuit 272 mixes the audio signal received by the communication circuit 271 with the audio signal output by the DSP 22, and outputs the mixed audio signal to the speaker 28.

[0090] The speaker 28 reproduces the playback sound indicated by the mixed sound signal obtained from the mixing circuit 272. The speaker 28 is a speaker that emits sound waves toward the ear canal (eardrum) of the user wearing the wearable device 20. The speaker 28 is not limited to this, and may also be a bone conduction speaker.

[0091] By installing a predetermined application program, the information terminal 30 functions as a user interface device in the sound signal processing system 10. The information terminal 30 also functions as a sound source that provides sound signals to the wearable device 20.

[0092] Specifically, the user can operate the information terminal 30 to select music content to be played by the speaker 28. The information terminal 30 includes a UI (User Interface) unit 31, a communication circuit 32, a control unit 33, and a storage unit 34.

[0093] The UI unit 31 is a user interface device that receives user operations and presents images to the user, and is configured with an operation device such as a touch panel and a display device such as a display panel.

[0094] The communication circuit 32 transmits a sound signal of a sound source (hereinafter referred to as a third sound signal) such as music content selected by the user to the wearable device 20. The communication circuit 32 is, for example, a wireless communication circuit, and communicates with the wearable device 20 based on a communication standard such as Bluetooth (registered trademark) or BLT.

[0095] The control unit 33 performs information processing related to the display of images on the display device provided in the UI unit 31 and the transmission of the third sound signal using the communication circuit 32. The control unit 33 is realized by, for example, a microcomputer, but may also be realized by a processor. The image display function and the third sound signal transmission function are realized by the microcomputer or the like constituting the control unit 33 executing a computer program stored in the storage unit 34.

[0096] The storage unit 34 is a storage device that stores various information required for the control unit 33 to perform information processing, computer programs executed by the control unit 33, sound signals of music content, etc. The storage unit 34 is configured by, for example, a semiconductor memory.

[0097] Next, a method for reproducing sound in the wearable device 20 will be described. Fig. 9 is a flowchart showing the sound reproduction process in the wearable device 20. The reproduction process shown in Fig. 9 starts when an operation to start the wearable device 20 is performed, and is repeated until an operation to stop the wearable device 20 is performed. Note that the reproduction process shown in Fig. 9 may also start when the communication circuit 271 receives a third sound signal, which is a sound signal of a sound source such as music content, from the information terminal 30 after an operation to start the wearable device 20 is performed.

[0098] In step S1, the microphone 21 picks up a sound and outputs a sound signal of the picked up sound.

[0099] In step S2 , the speech detection unit 241 detects a speech sound signal indicating a speech sound from the sound signal output by the microphone 21 , and outputs the detected speech sound signal to the determination unit 25 .

[0100] In step S3 , the direction determination unit 251 and the voice determination unit 252 calculate the MUSIC spectrum, MFCC, and recurrence plot of the sound indicated by the speech sound signal output by the speech detection unit 241 .

[0101] In step S4, the direction determination unit 251 determines whether or not the sound indicated by the speech sound signal output by the speech detection unit 241 has directly arrived at the microphone 21. Specifically, the direction determination unit 251 acquires a determination model 261 from the storage unit 26. The direction determination unit 251 inputs to the determination model 261 a spectrum of a predetermined range including a peak portion of the MUSIC spectrum of the sound calculated in step S3, and the MFCC and recurrence plot of the sound calculated in step S3.

[0102] If the determination model 261 outputs information indicating that the sound has arrived directly at the microphone 21, the direction determination unit 251 determines that the sound has arrived directly at the microphone 21 (YES in step S4). In this case, the process proceeds to step S5. On the other hand, if the determination model 261 outputs information indicating that the sound has not arrived directly at the microphone 21, the direction determination unit 251 determines that the sound has not arrived directly at the microphone 21 (NO in step S4). In this case, the process proceeds to step S9.

[0103] In step S5, the voice determination unit 252 determines whether or not voice is included in the sound indicated by the speech sound signal output by the speech detection unit 241. Specifically, the voice determination unit 252 acquires a determination model 261 from the storage unit 26. The voice determination unit 252 inputs into the determination model 261 a spectrum of a predetermined range including a peak portion of the MUSIC spectrum of the sound calculated in step S3, and the MFCC and recurrence plot of the sound calculated in step S3.

[0104] If the determination model 261 outputs information indicating that the sound contains speech, the speech determination unit 252 determines that the sound contains speech (YES in step S5). In this case, the process proceeds to step S6. On the other hand, if the determination model 261 outputs information indicating that the sound does not contain speech, the speech determination unit 252 determines that the sound does not contain speech (NO in step S5). In this case, the process also proceeds to step S9.

[0105] In step S6, the switching unit 242 switches the operation mode to the ambient mode. When the operation mode is switched to the ambient mode, the switching unit 242 performs an equalizing process to emphasize specific frequency components on the sound signal output by the microphone 21. The switching unit 242 outputs a first sound signal, which is the sound signal after the equalizing process, to the mixing circuit 272.

[0106] In step S9, if a predetermined time has not elapsed since the most recent time at which it was determined in step S5 that the sound contained speech (YES in step S5) (NO in step S9), the switching unit 242 proceeds to step S6. In other words, in step S9, if a predetermined time has not elapsed since the most recent time at which it was determined in step S4 that the sound directly reached the microphone 21 (YES in step S4) and the most recent time at which it was determined in step S5 that the sound contained speech (YES in step S5) (NO in step S9), the switching unit 242 proceeds to step S6.

[0107] On the other hand, in step S9, if the switching unit 242 determines that sound arrived directly at the microphone 21 in step S4 (YES in step S4) and that the sound contains voice in step S5 (YES in step S5), and a predetermined time has passed since then (YES in step S9), the processing proceeds to step S10.

[0108] 9 starts, step S9 may be performed when it has not been determined in step S4 that sound has directly reached the microphone 21 (YES in step S4) and it has not been determined in step S5 that the sound contains speech (YES in step S5). In this case, the switching unit 242 proceeds to step S6. However, this is not limiting, and in this case, the switching unit 242 may proceed to step S10.

[0109] In step S7, the mixing circuit 272 mixes the sound signal output by the DSP 22 with a third sound signal, which is a sound signal of a sound source such as music content received by the communication circuit 271, and outputs the mixed sound signal to the speaker 28.

[0110] In step S7, which follows step S6, the mixing circuit 272 mixes the first sound signal output in step S6 with the third sound signal received by the communication circuit 271, and outputs the mixed sound to the speaker 28. In this case, in step S8, the speaker 28 reproduces the playback sound indicated by the first sound signal mixed with the third sound signal.

[0111] In step S10, the switching unit 242 switches the operation mode to the noise cancellation mode. When the operation mode is switched to the noise cancellation mode, the switching unit 242 performs phase inversion processing on the sound signal output by the microphone 21, and outputs a second sound signal, which is the sound signal after the phase inversion processing, to the mixing circuit 272.

[0112] In step S7, which follows step S10, the mixing circuit 272 mixes the second sound signal output in step S10 with the third sound signal received by the communication circuit 271, and outputs the result to the speaker 28. In this case, in step S8, the speaker 28 reproduces the playback sound indicated by the second sound signal mixed with the third sound signal.

[0113] In the above embodiment, when the sound that directly reaches the microphone 21 includes speech, the reproduced sound indicated by the first sound signal mixed with the third sound signal is reproduced, assuming that the speech sound that directly reaches the user has arrived at the microphone 21. In this case, the speech sound that directly reaches the user is emphasized, so the user can listen to the speech that directly reaches the user while listening to the sound indicated by the third sound signal.

[0114] On the other hand, suppose that a speech sound indirectly reaches the microphone 21. Or, suppose that a speech sound that does not include human voice, such as an announcement sound, reaches the microphone 21 directly. In these cases, the speech sound that reaches the user indirectly is treated as having arrived at the microphone 21, and the reproduced sound represented by the second sound signal mixed with the third sound signal is reproduced. This reduces ambient noise around the microphone 21, including the speech sound that reaches the user indirectly. This makes it easier for the user to hear the sound represented by the third sound signal.

[0115] In the above embodiment, the first sound signal is output for a predetermined time from the time when it is determined that a sound has directly reached the microphone 21 and that the sound contains speech, so that the user can continue to listen to the speech as long as the speech reaching the user directly is not interrupted for a predetermined time or longer.

[0116] The present disclosure can employ the following modifications.

[0117] (1) When the switching unit 242 switches the operating mode to the noise cancellation mode, the sound signal output by the microphone 21 may be subjected to signal processing (second signal processing) that removes frequency components of the audio contained in the sound indicated by the sound signal, such as equalization processing or filtering processing that is different from phase inversion processing.

[0118] (2) In the above embodiment, an example has been described in which a predetermined range of the spectrum including the peak portion of the MUSIC spectrum of the sound is used in the machine learning for constructing the determination model 261 used by the direction determination unit 251 and the voice determination unit 252. However, instead of a predetermined range of the spectrum including the peak portion of the MUSIC spectrum of the sound, the entire MUSIC spectrum of the sound may be used in the machine learning. In this case, the direction determination unit 251 and the voice determination unit 252 may input the entire MUSIC spectrum of the sound indicated by the speech sound signal to the determination model constructed by the machine learning.

[0119] Furthermore, only at least one of the MFCC and recurrence plot of the sound may be used in the machine learning for constructing the determination model 261. For example, only the MFCC of the sound may be used in the machine learning for constructing the determination model 261, without using the recurrence plot of the sound. In this case, the direction determination unit 251 and the voice determination unit 252 may input at least one of the MFCC and recurrence plot of the sound indicated by the speech sound signal, which were used in the machine learning, to the determination model constructed by the machine learning.

[0120] (3) In the above embodiment, an example has been described in which the direction determination unit 251 and the voice determination unit 252 perform determination using the same determination model 261. However, the learning model (hereinafter, “first determination model”) used by the direction determination unit 251 to determine whether or not a sound has directly reached the microphone 21 may be different from the learning model (hereinafter, “second determination model”) used by the voice determination unit 252 to determine whether or not a sound contains voice.

[0121] Specifically, the first determination model may be constructed by machine learning the relationship between the MUSIC spectrum, MFCC, and recurrence plot of a sound and whether or not the sound directly arrived at the microphone 21. On the other hand, the second determination model may be constructed by machine learning the relationship between the MUSIC spectrum, MFCC, and recurrence plot of a sound and whether or not the sound includes speech.

[0122] Alternatively, the second determination model may be constructed by machine learning the relationship between the MUSIC spectrum, MFCC, and recurrence plot of the sound contained in the sound and whether the sound contains a sound indicating a predetermined word. In this case, the sound determination unit 252 may use the second determination model to determine whether the sound contained in the sound indicated by the speech sound signal contains a sound indicating a predetermined word. The predetermined word may be, for example, a word spoken to the user, such as "Hello," which is used as a wake word.

[0123] Alternatively, the direction determination unit 251 may not use the first determination model, and may determine that the sound has arrived directly at the microphone 21 if the peak width in the MUSIC spectrum of the sound indicated by the speech sound signal is less than a predetermined width, and may determine that the sound has not arrived directly at the microphone 21 if the peak width is equal to or greater than the predetermined width.

[0124] (4) The determination unit 25 does not need to include the voice determination unit 252. In this case, step S5 ( FIG. 9 ) may be omitted. In this case, in step S9 ( FIG. 9 ), if the predetermined time has not elapsed since the switching unit 242 most recently determined in step S4 that a sound directly reached the microphone 21, the processing proceeds to step S6. If the predetermined time has elapsed since the switching unit 242 most recently determined in step S4 that the sound directly reached the microphone 21, the processing proceeds to step S10.

[0125] The wearable device of the present disclosure can appropriately reproduce speech sounds that reach the user directly, making it suitable for use when working while listening to music in a location where other people are nearby.

Claims

1. A wearable device comprising: a microphone that acquires sound and outputs a sound signal of the acquired sound; a signal processing circuit that detects the direction from which the sound is coming and determines whether the sound has arrived directly at the microphone based on the detection result, and if it is determined that the sound has arrived directly at the microphone, outputs a first sound signal obtained by performing first signal processing on the sound signal, including equalization processing to emphasize frequency components of the audio contained in the sound; and a speaker that reproduces a sound indicated by the output first sound signal.

2. A wearable device comprising: a microphone that acquires sound and outputs a sound signal of the acquired sound; a signal processing circuit that detects the direction from which the sound is coming and determines whether the sound has arrived directly at the microphone based on the detection result, and outputs the sound signal as a first sound signal if it is determined that the sound has arrived directly at the microphone, and outputs a second sound signal that is obtained by performing second signal processing on the first sound signal, which includes removing frequency components of audio contained in the sound, if it is determined that the sound has not arrived directly at the microphone; and a speaker that reproduces a sound indicated by one of the output first sound signal and the second sound signal.

3. The wearable device according to claim 1, wherein, when the signal processing circuit determines that the sound is not arriving directly at the microphone, it outputs a second sound signal obtained by performing second signal processing on the sound signal, including removing frequency components of the voice contained in the sound, and the speaker reproduces a reproduced sound indicated by one of the output first sound signal and the output second sound signal.

4. The wearable device according to claim 2 or 3, wherein the second signal processing includes phase inversion processing.

5. The wearable device according to claim 1 or 2, wherein the signal processing circuit performs signal processing on the sound signal to detect the MUSIC spectrum of the sound as information indicating the direction from which the sound is coming, and determines whether the sound is coming directly to the microphone based on the MUSIC spectrum of the sound.

6. The wearable device according to claim 5, wherein the signal processing circuit determines that the sound has arrived directly at the microphone if the peak width in the MUSIC spectrum of the sound is less than a predetermined width, and determines that the sound has not arrived directly at the microphone if the peak width is equal to or greater than the predetermined width.

7. The wearable device according to claim 5, wherein the signal processing circuit determines whether the sound has arrived directly at the microphone based on a predetermined range of the spectrum including a peak portion of the MUSIC spectrum of the sound.

8. The wearable device according to claim 5, wherein the signal processing circuit determines whether the sound arrived directly at the microphone using a learning model trained based on the MUSIC spectrum of the sound.

9. The wearable device according to claim 1 or 2, wherein the signal processing circuit further performs signal processing on the sound signal to determine whether the sound contains speech, and outputs the first sound signal when it determines that the sound has arrived directly at the microphone and when it determines that the sound contains speech.

10. The wearable device according to claim 9, wherein the signal processing circuit further determines whether the audio contained in the sound includes audio representing a predetermined word, and outputs the first sound signal if it determines that the audio contained in the sound includes audio representing the predetermined word.

11. The wearable device of claim 10, wherein the signal processing circuit further performs signal processing on the sound signal to calculate the periodicity of the sound contained in the sound, and determines whether the sound contained in the sound includes a sound indicating the specified word based on the periodicity of the sound contained in the sound.

12. The wearable device according to claim 9, wherein the signal processing circuit further performs signal processing on the sound signal to calculate the periodicity of the sound, and determines whether the sound includes speech based on the periodicity of the sound.

13. The wearable device of claim 11, wherein the periodicity of the speech contained in the sound is represented by a recurrence plot, and the signal processing circuit uses a learning model trained based on the recurrence plot to determine whether the speech contained in the sound includes speech indicating the specified word.

14. The wearable device of claim 9, wherein the signal processing circuit outputs the first sound signal when the point at which it determines that the sound did not arrive directly at the microphone or the point at which it determines that the sound does not contain speech is the point at which it determines that the sound arrived directly at the microphone and a predetermined time has not elapsed since it determined that the sound contained speech.

15. The wearable device according to claim 1 or 2, further comprising a mixing circuit that mixes the output first sound signal with a third sound signal provided from a sound source, and the speaker reproduces a sound represented by the first sound signal mixed with the third sound signal.

16. A playback method for a wearable device comprising a microphone that acquires sound and outputs a sound signal of the acquired sound, a signal processing circuit that performs signal processing on the sound signal, and a speaker, wherein the processing performed by the signal processing circuit includes: detecting the direction from which the sound is coming indicated by the sound signal output by the microphone; determining whether the sound has arrived directly at the microphone based on the detection result; if it is determined that the sound has arrived directly at the microphone, outputting a first sound signal obtained by performing first signal processing on the sound signal, including equalizing processing to emphasize the frequency components of the audio contained in the sound; and causing the speaker to reproduce the playback sound indicated by the output first sound signal.

17. A program for a wearable device comprising a microphone that acquires sound and outputs a sound signal of the acquired sound, a signal processing circuit that performs signal processing on the sound signal, and a speaker, the program causing the signal processing circuit to: detect the direction from which the sound is coming indicated by the sound signal output by the microphone; determine based on the detection result whether the sound has arrived directly at the microphone; if it is determined that the sound has arrived directly at the microphone, output a first sound signal obtained by performing first signal processing on the sound signal, including equalizing processing to emphasize the frequency components of the audio contained in the sound; and reproduce the reproduced sound indicated by the output first sound signal on the speaker.

Citation Information

Patent Citations

  • Voice processing device, voice processing method and program

    JP2018169473A

  • Interactive device, interactive method, and interactive computer program

    WO2018078885A1

  • Sound signal processing system and sound signal processing method

    WO2022054414A1

  • Ear-mounted type device and reproduction method

    WO2022137806A1