Sound processing system, sound processing device, sound processing method, and sound processing program

The sound processing system efficiently identifies and outputs attention-relevant sounds using dual microphones and analysis, allowing users to focus on important auditory events with ease.

WO2026004753A1PCT designated stage Publication Date: 2026-01-02KYOCERA CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/022221
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-26
Filing Date
2025-06-19
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing sound processing systems require complex operations for selecting target sounds, failing to efficiently identify and output attention-relevant sounds in real-time.

Method used

A sound processing system with extracorporeal and intracorporeal microphones to detect and output 'first sounds' (user's inaudible murmurs) and 'second sounds' (attention-relevant external sounds), using volume, timing, and frequency analysis to determine and emphasize attention events.

Benefits of technology

Enables users to selectively focus on relevant external sounds with simple operations, enhancing attention to important auditory events in real-time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025022221_02012026_PF_FP_ABST
    Figure JP2025022221_02012026_PF_FP_ABST
Patent Text Reader

Abstract

A sound processing system according to one embodiment comprises: a sound acquisition unit that acquires sounds including a first sound uttered by a user and having a duration of a first time period or less, and a second sound, which is a sound other than the sound uttered by the user; and a sound output unit that outputs the second sound when the sound acquisition unit has acquired the first sound.
Need to check novelty before this filing date? Find Prior Art

Description

SOUND PROCESSING SYSTEM, SOUND PROCESSING DEVICE, SOUND PROCESSING METHOD, AND SOUND PROCESSING PROGRAM Cross-reference to related applications

[0001] This application claims priority to Japanese Patent Application No. 2024-103150, filed on June 26, 2024, the entire disclosure of which is incorporated herein by reference.

[0002] The present disclosure relates to a sound processing system, a sound processing device, a sound processing method, and a sound processing program.

[0003] Conventionally, electronic devices that process a selected sound when the user selects a target sound are known. For example, Patent Literature 1 discloses an invention that includes a control device (smartphone) and that emphasizes or deletes a target sound by touching with a fingertip a symbol indicating the location of the target sound source on the display of the control device.

[0004] Patent Publication No. 2014-552983

[0005] A sound processing system according to one embodiment includes a sound acquisition unit that acquires a sound including a first sound, which is a sound uttered by a user and has a duration equal to or shorter than a first time, and a second sound, which is a sound other than the sound uttered by the user, and a sound output unit that outputs the second sound when the sound acquisition unit acquires the first sound.

[0006] Moreover, a sound processing device according to one embodiment includes a sound acquisition unit that acquires a first sound, which is a sound uttered by a user and has a duration equal to or shorter than a first time, and a second sound, which is a sound other than the sound uttered by the user; and a sound output unit that outputs the second sound when the sound acquisition unit acquires the first sound.

[0007] Moreover, a sound processing method according to one embodiment includes: a sound acquisition unit acquiring a sound including a first sound, which is a sound uttered by a user and has a duration equal to or shorter than a first time, and a second sound, which is a sound other than the sound uttered by the user; and a sound output unit outputting the second sound when the sound acquisition unit acquires the first sound.

[0008] In addition, a sound processing program according to one embodiment causes a sound acquisition unit to acquire a sound including a first sound, which is a sound uttered by a user and has a duration equal to or shorter than a first time, and a second sound, which is a sound other than the sound uttered by the user; and when the sound acquisition unit acquires the first sound, causes a sound output unit to output the second sound.

[0009] FIG. 1 is a diagram for explaining the configuration of a sound processing system 100 of a first embodiment. FIG. 2 is a flowchart for explaining a series of flows in a processing method of the sound processing system 100 of the first embodiment. FIG. 3 is a diagram for explaining a method for determining an attention event sound. FIG. 4 is a diagram for explaining a method for determining an attention event sound. FIG. 5 is a diagram for explaining a method for determining an attention event sound. FIG. 6 is a diagram for explaining a method for determining whether or not a sound is the first sound using a decision tree. FIG. 7 is a table showing an example of a method for determining whether or not a sound is the first sound using a score table. FIG. 1 is a diagram for explaining the configuration of a sound processing system 200 of a second embodiment. FIG. 1 is a flowchart for explaining a series of flows in a processing method of the sound processing system 200 of the second embodiment. FIG. 2 is a diagram for explaining the configurations of sound processing systems 300, 400 of a third embodiment. FIG. 3 is a flowchart for explaining a series of flows in a processing method of the sound processing system 300 of the third embodiment. FIG. 4 is a flowchart for explaining a series of flows in a processing method of the sound processing system 400 of the third embodiment.

[0010] According to one embodiment, it is possible to provide a sound processing system or the like that allows a user to select sounds with a simple operation. Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In each drawing, the same reference numerals indicate components having the same or equivalent functions. Note that the configurations, numerical values, processing flows, functions, elements, etc. described in the following embodiments are merely examples, and at least one of them may be freely modified and changed, and it is not intended that the technical scope of the present disclosure be limited to the following description.

[0011] First Embodiment Configuration of Sound Processing System 100 Fig. 1 is a diagram illustrating a schematic configuration of a sound processing system 100. The sound processing system 100 includes a sound processing device that is worn on a user's ear when used. The user is a person who uses at least one of the sound processing system 100 and the sound processing device, and is, for example, a human. The sound processing device may be, but is not limited to, an inner-ear earphone, an in-ear earphone, a binaural earphone, a monoaural earphone, or a monaural earphone. Furthermore, the sound processing device is not limited to earphones, and may also be headphones or a headset.

[0012] The sound processing system 100 includes a sound acquisition unit 110, a first sound detection unit 113, a storage unit 114, a sound determination unit 115, a sound output unit 116, and a control unit 117. The sound acquisition unit 110 may include at least one of an extracorporeal sound acquisition unit 111 and an intracorporeal sound acquisition unit 112. One or more components included in the sound processing system 100 may be included in the same device. For example, at least the sound output unit 116 and the sound acquisition unit 110 may be included in the same device to configure a sound processing device. The sound processing device may include the first sound detection unit 113 and the sound determination unit 115. Note that the sound processing system 100 does not need to be configured as a single sound processing device. Some of the components included in the sound processing system 100 may be provided outside the sound processing device and configured to be able to communicate with the sound processing device via a wired or wireless connection. For example, the storage unit 114 may be located in a remote server via a network system or the like. For example, the internal body sound acquisition section 112 included in the sound processing device may be configured to be able to communicate with some of the components included in the sound processing system 100.

[0013] The sound acquisition unit 110 includes one or more microphones that acquire sound. The sound acquisition unit 110 may include at least an extracorporeal sound acquisition unit 111. The sound acquisition unit 110 may include an internal body sound acquisition unit 112 in addition to the extracorporeal sound acquisition unit 111. The number of extracorporeal sound acquisition units 111 may be one or more. The number of extracorporeal sound acquisition units 111 and the internal body sound acquisition units 112 included in the sound acquisition unit 110 may be the same or different. For example, two extracorporeal sound acquisition units 111 and one internal body sound acquisition unit 112 may be provided. Providing a plurality of extracorporeal sound acquisition units 111 makes it easier to determine the direction from which the acquired sound is coming.

[0014] The extracorporeal sound acquisition unit 111 is a microphone that is arranged facing the outside of the user's body. The extracorporeal sound acquisition unit 111 acquires sounds that come from outside the user's body. The sounds that come from outside the user's body include sounds that come from sound sources around the user. The sounds that come from around the user include sounds uttered by the user, sounds uttered by people around the user, and sounds emitted from sound sources around the user. The sound sources around the user may include not only people, but also at least one of speakers and natural phenomena that generate some kind of sound.

[0015] The body sound acquisition unit 112 is a microphone that acquires sounds arriving through the user's body. The body sound acquisition unit 112 acquires sounds including sounds that propagate through the user's body and arrive. The body sound acquisition unit 112 may be arranged so as to contact the ear of a user wearing the sound processing system 100 or a part of a device included in the sound processing system 100. The body sound acquisition unit 112 may be arranged, for example, near the entrance of the user's ear canal and facing the inside of the user's body. The body sound acquisition unit 112 may use, for example, any one of a Non-Audible Murmur (NAM) microphone, an ear canal microphone, a bone conduction microphone, and an intraoral microphone, or multiple types of microphones. The body sound acquisition unit 112 is not limited to these microphones and may be any device that can acquire sounds arriving through the body.

[0016] The internal body sound acquisition unit 112 can acquire sounds uttered by the user without the intention of letting surrounding people hear them. Such sounds may be referred to as "first sounds" in this specification. The first sounds may be sounds uttered by the user when paying attention to a sound emitted from an external sound source. The first sounds may include, for example, at least one of an inaudible murmur, a murmur, a whisper, and a low voice. The first sounds may be sounds of a length equal to or less than a predetermined time (sometimes referred to as "first time" in this specification), for example, an utterance consisting of two or fewer moras. The first time may be the time it takes the user to utter two moras, for example, 0.5 seconds. The first sounds may include, for example, nasal sounds. The first sounds may be sounds registered in advance by the user, for example, short utterances such as "hmm," "eh," or "ah." By detecting the first sound, the sound processing system 100 can determine that the user has directed selective attention to a sound emitted from a sound source outside the user. Hereinafter, in this specification, "a sound emitted from a person other than the user or a sound source around the user to which the user has directed selective attention" may be referred to as an attention event sound.

[0017] The extracorporeal sound acquisition unit 111 and the intracorporeal sound acquisition unit 112 can acquire a first sound. The first sound acquired by the intracorporeal sound acquisition unit 112 includes sound acquired by propagating only inside the user's body (i.e., sound acquired by propagating directly inside the user's body without going outside the user's body) and sound acquired by propagating inside the user's body after being emitted outside the user's body. The first sound acquired by the extracorporeal sound acquisition unit 111 includes sound acquired by the extracorporeal sound acquisition unit 111 from outside the user's body.

[0018] The first sound detection unit 113 includes at least one processor. The processor may be, for example, a general-purpose processor such as a central processing unit (CPU) or a graphics processing unit (GPU), or a dedicated processor specialized for a specific process.

[0019] The first sound detection unit 113 determines whether the sound acquired by the sound acquisition unit 110 includes the first sound. The first sound detection unit 113 may separate the sound acquired by the sound acquisition unit 110 into sound source units and determine whether the separated sounds are the first sound. When the volume of a sound uttered by the user toward outside the body or a sound generated outside the body is high, at least a portion of this sound passes through the user's head and propagates into the ear canal and is acquired by the body sound acquisition unit 112. When the user speaks in everyday conversation, i.e., when the user intends to have others hear it, the sound acquired by the body sound acquisition unit 112 is quieter than the sound acquired by the extracorporeal sound acquisition unit 111. On the other hand, the volume of the first sound, i.e., a sound uttered without the intention of having others hear it, is low, and therefore the body sound acquisition unit 112 tends to acquire it at a louder volume than the extracorporeal sound acquisition unit 111. Therefore, the sound acquiring unit 110 can determine whether or not a sound is the first sound by, for example, comparing the volume of the sound acquired by the extracorporeal sound acquiring unit 111 with the volume of the sound acquired by the intracorporeal sound acquiring unit 112. Details of the method for determining whether or not the acquired sound includes the first sound will be described later.

[0020] The sound determination unit 115 includes at least one processor. The processor may be, for example, a general-purpose processor such as a CPU or a GPU, or a dedicated processor specialized for a specific process.

[0021] When the sound acquisition unit 110 acquires a first sound uttered by the user after the sound source around the user has emitted a sound (immediately after or after a predetermined time), the sound determination unit 115 determines whether the sound emitted from the sound source around the user is an attention event sound. A method for determining whether the sound emitted from the sound source around the user is an attention event sound will be described later. In this specification, a sound other than the sound uttered by the user and generated by a sound source around the user before the user uttered the first sound may be referred to as a second sound. The second sound may be acquired by the sound acquisition unit 110. The second sound may include a human voice, a sound played from a speaker, or a sound generated by a natural phenomenon.

[0022] The first sound detection unit 113 and the sound determination unit 115 may be configured as hardware or software. The first sound detection unit 113 and the sound determination unit 115 may be controlled by the control unit 117. The first sound detection unit 113 and the sound determination unit 115 may be included in the control unit 117.

[0023] The storage unit 114 is a storage medium including a ROM (Read Only Memory) or a RAM (Random Access Memory), etc. The storage unit 114 may store a program executed by the sound processing system 100. This program corresponds to a sound processing program executed by the sound processing system 100. The storage unit 114 may store the sound acquired by the sound acquisition unit 110 as a sound signal. The storage unit 114 may store sounds including the second sound. A part of the storage unit 114 may be configured as a ring buffer, and the sound acquired by the sound acquisition unit 110 may be temporarily stored as a sound signal. The storage unit 114 may store the sounds acquired by the sound acquisition unit 110 separately for each sound source. The storage unit 114 may store the sounds acquired by the extracorporeal sound acquisition unit 111 in association with the time of acquisition.

[0024] The sound output unit 116 is, for example, a speaker. The sound output unit 116 plays back the sound acquired by the sound acquisition unit 110. The sound output unit 116 may output the sound signal stored in the storage unit 114 as sound. The sound output unit 116 may output the acquired sound in near real time.

[0025] The sound output unit 116 outputs a sound that the control unit 117 determines to be an attention event sound from among the sounds including the second sound stored in the storage unit 114. The sound output unit 116 may output a sound that has been subjected to predetermined processing by the control unit 117. A method for outputting sounds from the sound output unit 116 will be described later. The sound output unit 116 may also output a pre-registered sound in addition to the sound acquired by the sound acquisition unit 110. The sound output unit 116 may continue to output a sound including a sound determined to be an attention event sound until the output of the sound is completed. The sound output unit 116 may also include a display or the like that visually displays the output sound. The sound output unit 116 may output information about the sound (e.g., at least one of the location of the sound source, the time the sound was generated, and the type of sound) to a display or the like to notify the user.

[0026] The control unit 117 includes at least one processor. The processor may be, for example, a general-purpose processor such as a CPU or a GPU, or a dedicated processor specialized for a specific process.

[0027] The control unit 117 controls the first sound detection unit 113 to determine whether or not the sound acquired by the sound acquisition unit 110 is the first sound. The control unit 117 may determine whether or not the sound is the first sound using one or more determination methods.

[0028] The control unit 117 controls the sound determination unit 115 to determine which of the multiple second sounds from the multiple sound sources acquired by the sound acquisition unit 110 and stored in the storage unit 114 is an attention event sound. When the second sound acquired by the sound acquisition unit 110 has one sound source, the second sound may be determined to be an attention event sound.

[0029] The second sound may be, for example, at least one of the sounds of people around the user, noise, the voice of an animal, and the sound of a doorbell. The extracorporeal sound acquisition unit 111 may acquire the second sound from a plurality of sound sources.

[0030] A series of flows of a processing method of the sound processing system 100 will be described with reference to Fig. 2. Fig. 2 is a flowchart for explaining a series of flows of a processing method of the sound processing system 100 in the first embodiment.

[0031] The sound acquisition unit 110 acquires a sound (step S1). The sound acquired by the sound acquisition unit 110 may be either the sound acquired by the extracorporeal sound acquisition unit 111 or the sound acquired by the intracorporeal sound acquisition unit 112, or may be both.

[0032] The sound acquired by the sound acquisition unit 110 is stored as a sound signal in the storage unit 114 (step S2).

[0033] The first sound detection unit 113 determines whether the sounds acquired by the sound acquisition unit 110 include the first sound (step S3).

[0034] If it is determined in step S3 that the sounds acquired by the sound acquisition unit 110 do not include the first sound, the sound output unit 116 outputs the sound signal stored in the storage unit 114 as sound (step S4). At this time, it is desirable that the sound output unit 116 outputs the sounds acquired by the sound acquisition unit 110 in close to real time.

[0035] The sound processing system 100 repeats steps S1 to S4 until the first sound detection unit 113 determines that the sound acquired by the sound acquisition unit 110 includes the first sound.

[0036] In step S3, if it is determined that the sound acquired by the sound acquisition unit 110 is the first sound, the sound determination unit 115 determines the attention event sound from the sound (i.e., the second sound) generated by the surrounding sound source before the user uttered the first sound (step S5).

[0037] The sound output unit 116 outputs the sound determined to be the attention event sound using a predetermined output method (step S6). The sound output unit 116 may continue outputting the sound until the generation of the attention event sound has ended. The method for outputting the attention event sound will be described later. Note that the sound output unit 116 may perform the processes of steps S1 to S3 while outputting the attention event sound in step S6. In this case, for example, the user can hear the attention event sound while listening to sounds generated from surrounding sound sources and acquired by the extracorporeal sound acquisition unit 111 in almost real time.

[0038] The control unit 117 determines whether the attention event sound has ended using a method described below (step S7). If the attention event sound has not ended, the sound output unit 116 continues to output the attention event sound (step S6).

[0039] When the attention event sound output in step S6 has ended, the sound output unit 116 stops outputting the attention event sound (step S8).

[0040] Next, a method for determining an attention event sound will be described using FIG. 3 . FIGS. 3A, 3B, and 3C are diagrams showing the relationship between the passage of time and the volume of three different sounds acquired from the same time using sound signals. The sound signals in FIGS. 3A and 3B are sound signals generated from sound sources around the user and acquired by the extracorporeal sound acquisition unit 111. The sounds in FIGS. 3A and 3B may include sound signals emitted from different sound sources. The sound signal in FIG. 3C is a sound signal acquired by the intracorporeal sound acquisition unit 112 or the extracorporeal sound acquisition unit 111, and includes a first sound. In FIGS. 3A and 3B, the timing when the volume becomes larger than a predetermined volume (e.g., 50 dBSPL) may be determined as the timing when the generation of the second sound begins. In FIG. 3C, the timing when the volume becomes larger than a predetermined volume (e.g., 40 dBSPL) may be determined as the timing when the production of the first sound begins. In this case, the timing at which the second sound in Fig. 3A starts is earlier than the timing at which the second sound in Fig. 3B starts. In other words, the timing at which the sounds in Fig. 3 start is such that the second sound in Fig. 3B starts after the second sound in Fig. 3A, and then the first sound in Fig. 3C starts. The volume values ​​used for these determinations may be preset to different values ​​for each individual.

[0041] The control unit 117 determines whether the second sound from the multiple sound sources acquired by the sound acquisition unit 110 is an attention event sound based on the time when the first sound was uttered. The control unit 117 may determine the second sound from the sound source that generated the sound closest to the time when the first sound was uttered, among the multiple sound sources, to be an attention event sound. The control unit 117 may determine the second sound from the sound source that started generating the sound closest to the time when the first sound was uttered, among the multiple sound sources, to be an attention event sound. In the case of FIG. 3 , the signal of the second sound that started generating the sound closest to the time when the first sound was uttered, as shown in FIG. 3B, can be determined to be the attention event sound signal. The control unit 117 may determine the sound generated from a predetermined time (sometimes referred to as the second time in this specification) before the user utters the first sound to the time when the user utters the first sound, to be an attention event sound. The second time may be, for example, one second. The control unit 117 causes the sound output unit 116 to output a second sound determined to be a caution event sound from among the second sounds stored in the storage unit 114 .

[0042] When the warning event sound has been generated for a predetermined time or longer (e.g., 2 seconds) and has fallen below a predetermined volume (e.g., 50 dBSPL), the control unit 117 may determine that the warning event sound has ended and stop output from the sound output unit 116.

[0043] If the user again utters the first sound while the output of the attention event sound is continuing, control unit 117 determines that a new attention event sound has been generated and causes sound output unit 116 to stop outputting the continuing attention event sound. Thereafter, control unit 117 determines whether the newly started attention event sound has been generated using a method that is the same as or similar to the method described above.

[0044] <Method for determining first sound> Described below are determination methods 1 to 7 for determining whether or not the sound acquired by the sound acquisition unit 110 includes the first sound. Note that the method for determining whether or not a sound is the first sound is not limited to determination methods 1 to 7. Furthermore, the first sound detection unit 113 may combine some of determination methods 1 to 7 to determine whether or not the detected sound is the first sound.

[0045] <Determination Method 1> In determination method 1, the control unit 117 determines whether a sound acquired by the internal sound acquisition unit 112 or the external sound acquisition unit 111 is a first sound based on the volume of the sound. For example, the control unit 117 may determine that the sound acquired by the internal sound acquisition unit 112 is a first sound if the volume of the sound acquired by the internal sound acquisition unit 112 is equal to or greater than a predetermined value (e.g., 30 dBSPL). The control unit 117 may determine that the sound acquired by the internal sound acquisition unit 112 is a first sound if the volume of the sound acquired by the internal sound acquisition unit 112 is equal to or greater than a first predetermined value and within a second predetermined value. The first predetermined value may be, for example, 30 dBSPL, and the second predetermined value may be, for example, 60 dBSPL. Furthermore, values ​​different from each other for each individual may be set in advance for these determinations.

[0046] <Determination Method 2> In determination method 2, the control unit 117 determines whether a sound is the first sound based on the volume relationship between the sound acquired by the internal sound acquisition unit 112 and the sound acquired by the external sound acquisition unit 111. The control unit 117 compares the sound acquired by the internal sound acquisition unit 112 with a sound acquired by the external sound acquisition unit 111 that is determined to be the same sound as the internal sound. The control unit 117 determines that the acquired sound is the first sound if the volume of the sound acquired by the internal sound acquisition unit 112 is greater than the volume of the sound acquired by the external sound acquisition unit 111. The control unit 117 may determine that the acquired sound is the first sound if the volume of the sound acquired by the internal sound acquisition unit 112 is at least twice as large as the volume of the sound acquired by the external sound acquisition unit 111. In addition, the fact that the sounds are the same means that the sounds are generated from the same sound source at the same time, and this may be determined based on at least one of the time when the internal sound acquisition unit 112 and the external sound acquisition unit 111 acquired the sounds, the voice quality, and the length of time the sounds were generated.

[0047] <Determination Method 3> In determination method 3, the control unit 117 determines that a sound acquired by the sound acquisition unit 110 is the first sound when the length of the sound acquired by the sound acquisition unit 110 is within a first time period. For example, the control unit 117 determines that a sound acquired by the sound acquisition unit 110 is the first sound when the length of the sound acquired by the sound acquisition unit 110 is 500 ms or less. For example, the control unit 117 determines that a sound acquired by the sound acquisition unit 110 is the first sound when the sound is an utterance consisting of two or fewer moras. In detection method 3, a sound may be determined to be the first sound based on the sound acquired by either or both of the extracorporeal sound acquisition unit 111 and the intracorporeal sound acquisition unit 112.

[0048] <Determination Method 4> In determination method 4, if the sound acquired by the sound acquisition unit 110 includes a sound higher than a predetermined pitch frequency, the control unit 117 determines that the sound is the first sound. For example, if the pitch frequency of the sound acquired by the sound acquisition unit 110 is 200 Hz or higher, the control unit 117 determines that the sound including the sound is the first sound. By detecting a sound higher than the predetermined pitch frequency, the control unit 117 can detect that the sound is a sound emitted during a short, strong inhalation. Furthermore, the predetermined pitch frequency may be at least one of a value that differs depending on gender, age, and user. Furthermore, the predetermined pitch frequency may be set to a pitch frequency that is higher by a predetermined value (e.g., 10 Hz) than the average pitch frequency of the user's inhalation sound measured in advance. In determination method 4, the first sound may be determined based on the sound acquired by either or both of the extracorporeal sound acquisition unit 111 and the internal sound acquisition unit 112.

[0049] <Determination Method 5> In determination method 5, when a sound acquired by the sound acquisition unit 110 includes a nasal sound, the control unit 117 determines that the sound is the first sound. The control unit 117 may detect the nasal sound using phoneme recognition. The control unit 117 may also acquire and detect the vibration of the nasal sound using a bone conduction microphone included in the sound acquisition unit 110. In detection method 5, the control unit 117 may determine that the sound is the first sound based on the sound acquired by either or both of the extracorporeal sound acquisition unit 111 and the intracorporeal sound acquisition unit 112.

[0050] <Determination Method 6> In determination method 6, when the pitch frequency of the end of a sound acquired by the sound acquisition unit 110 increases, the control unit 117 determines that the sound is the first sound. For example, when the pitch frequency of the end of a sound increases by 10 Hz / second, the control unit 117 can determine that the sound is the first sound. By confirming that the pitch frequency increases, the control unit 117 can detect that the user has uttered a question word. In determination method 6, the first sound may be determined based on the sound acquired by either or both of the extracorporeal sound acquisition unit 111 and the intracorporeal sound acquisition unit 112.

[0051] <Determination Method 7> In determination method 7, the control unit 117 performs phoneme recognition on the sound acquired by the sound acquisition unit 110 to determine that the sound is the first sound. Specifically, when the sound acquired by the sound acquisition unit 110 is a vowel sound that is louder than a predetermined threshold volume, and then immediately thereafter the volume decreases and becomes quieter than the predetermined threshold volume, the control unit 117 determines that the sound is the first sound. For example, when the control unit 117 detects a sound that has undergone a change in volume that decreases by 3 dB or more within a predetermined time, the control unit 117 determines that the sound is the first sound. In such a case, the control unit 117 can detect that the acquired sound includes a nasal consonant. In determination method 7, the first sound may be determined based on the sound acquired by either or both of the extracorporeal sound acquisition unit 111 and the intracorporeal sound acquisition unit 112.

[0052] The control unit 117 may determine whether the first sound is a first sound by combining determination methods 1 to 7. Specifically, the control unit 117 may determine whether the first sound is a first sound by using a determination tree that includes the above-mentioned determination methods 1 to 7. FIG. 4 is an example of the determination tree. The control unit 117 can determine whether the acquired sound is a first sound by making a determination according to the flow of the determination tree.

[0053] The control unit 117 may determine the first sound using a score table. An example of a method for determining the first sound using a score table will be described with reference to FIG. 5 . The score table is a table for converting an acquired sound into a score for determining whether it corresponds to the first sound. The higher the calculated score, the more likely the acquired sound is to be the first sound. The score table defines scores to be added when the requirements for determining that the sound is the first sound are met in each of the above-mentioned determination methods 1 to 7. The control unit 117 determines the acquired sound using scores calculated based on the determination results of determination methods 1 to 7 included in the score table. The control unit 117 determines that the acquired sound is the first sound if the total score is equal to or greater than a predetermined value (e.g., 70). When multiple sounds are acquired, the sound with the highest total score equal to or greater than a predetermined value may be determined to be the first sound.

[0054] <Output Method of Caution Event Sound> The following describes output methods 1 to 5 for outputting a sound that the control unit 117 has determined to be a caution event sound from the sound output unit 116. The sound output unit 116 can also use a combination of output methods 1 to 5.

[0055] <Output Method 1> In output method 1, when the sound acquired by the sound acquisition unit 110 includes sounds from multiple sound sources, the control unit 117 causes the sound output unit 116 to output a sound in which the attention event sound is emphasized more than sounds other than the attention event sound. By increasing the volume of the attention event sound and outputting it, the sound to which the user wants to pay attention can be reproduced so that it is easier to hear. In output method 1, the sound output unit 116 outputs the sound acquired by the sound acquisition unit 110 in near real time while emphasizing the attention event sound. The control unit 117 may emphasize the attention event sound by increasing the volume of the attention event sound more than the volume of sounds from the other sound sources. The control unit 117 may emphasize the attention event sound by increasing the volume of the attention event sound more than the original volume.

[0056] <Output Method 2> In output method 2, when the sound acquired by the sound acquisition unit 110 includes sounds from multiple sound sources, the control unit 117 causes the sound output unit 116 to output a sound in which sounds other than the attention event sound are suppressed compared to the attention event sound. By outputting sounds other than the attention event sound at a lower volume, the sound to which the user wants to pay attention can be reproduced so that it is easier to hear. In output method 2, the sound output unit 116 outputs the sound acquired by the sound acquisition unit 110 in near real time while suppressing sounds other than the attention event sound. The control unit 117 may make the volume of the sounds other than the attention event sound lower than the volume of the attention event sound. The control unit 117 may make the volume of the sounds other than the attention event sound lower than the original volume.

[0057] <Output Method 3> In output method 3, the control unit 117 outputs a second sound from two hours before the user utters the first sound as an attention event sound. The control unit 117 may cause the sound output unit 116 to output sound from the time the attention event sound starts. In output method 3, the user can rewind and listen to the sound. That is, the user listens back to an attention event sound that was output in the past. While an attention event sound acquired in the past is being output, the real-time sound acquired by the extracorporeal sound acquisition unit 111 may not be output. Furthermore, the control unit 117 may play the attention event sound that is rewinded and listened to at a slightly faster speed until it catches up with real time.

[0058] <Output Method 4> In output method 4, control unit 117 translates the attention event sound into another language and causes sound output unit 116 to output the translated attention event sound. For example, when the attention event sound is in a language other than the user's native language, control unit 117 translates the attention event sound into the user's native language and causes sound output unit 116 to output the translated attention event sound.

[0059] <Output Method 5> In output method 5, the control unit 117 stores the attention event sound as a sound signal in the storage unit 114. The control unit 117 may recognize the content of the attention event sound, and if the content matches at least one of a keyword and content registered in advance, store the sound as a sound signal in the storage unit. The control unit 117 may output the stored attention event sound and at least one text of the attention event sound to a display or the like. A configuration may be adopted in which the user can select text displayed on the display using an arbitrary user interface to listen back to the corresponding attention event sound stored in the storage unit 114.

[0060] Second Embodiment A sound processing system 200 according to a second embodiment of the present disclosure will be described with reference to FIG. 6 . The sound processing system 200 according to the second embodiment notifies the user when the user does not utter the first sound after a sound similar to a pre-registered sound is acquired by the sound acquisition unit. FIG. 6 is a diagram for describing a schematic configuration of the sound processing system 200. The sound processing system 200 includes an attention elicitation unit 218 in addition to the configuration of the first embodiment. Configurations other than the storage unit 214, the sound output unit 216, the control unit 217, and the attention elicitation unit 218 are the same as the corresponding configurations of the first embodiment, and therefore description thereof will be omitted. The attention elicitation unit 218 may be configured by software. The control unit 217 may control the attention elicitation unit 218. The attention elicitation unit 218 may be included in the control unit 217.

[0061] In addition to the configuration described in the first embodiment, the storage unit 214 stores pre-registered sounds that the user wants to hear. The sounds that the user wants to hear are sounds that the user does not want to miss, and may be registered in advance by the user. The sounds that the user wants to hear may be, for example, at least one of a sound indicating specific content, a specific keyword, and a sound from a specific sound source. The sounds indicating specific content may be, for example, at least one of a sound indicating the user's name and a sound indicating an event related to the user (such as an announcement for a train that the user plans to board). The sounds from a specific sound source may be, for example, at least one of a home phone call, a sound of a specific person, a sound from a speaker installed in a station, and a voice directed at the user.

[0062] The control unit 217 includes at least one processor. The processor may be, for example, a general-purpose processor such as a CPU or a GPU, or a dedicated processor specialized for a specific process.

[0063] The control unit 217 controls the attention eliciting unit 218 to determine whether the second sound acquired by the sound acquisition unit 210 includes a sound corresponding to a pre-registered sound. Whether the second sound corresponds to a pre-registered sound may be determined based on at least one of the similarity between the acquired sound and the registered sound and the similarity between the concept and meaning indicated by the sound. This determination may be made using sound recognition. The control unit 217 determines whether the user uttered the first sound within a predetermined time after the second sound including a sound corresponding to the pre-registered sound was generated. Whether the user uttered the first sound may be determined based on the sound acquired by the sound acquisition unit 210. If the user did not utter the first sound, the user may not have noticed that the sound they wanted to hear was generated, and therefore the control unit 217 causes the sound output unit 216 to notify the user. The notification may be made by outputting a sound from the sound output unit 216 or by vibrating a vibrator included in the sound output unit 216, but is not limited to these. The notification may be performed by outputting a second sound corresponding to a pre-registered sound as an attention event sound. The notification allows the user to realize that a sound they want to hear has been generated. The predetermined time may be set arbitrarily by the user or may be a predetermined time (e.g., 0.5 seconds). The predetermined time may be the time required from the previous generation of a pre-registered sound until the user utters the first sound. The predetermined time may be the average time required from the previous generation of a pre-registered sound until the user utters the first sound. If the user utters the first sound within a predetermined time after the generation of a second sound containing a sound corresponding to a pre-registered sound, an attention event sound identified using the same method as in the first embodiment is output from the sound output unit 216.

[0064] The attention eliciting unit 218 includes at least one processor. The processor may be, for example, a general-purpose processor such as a CPU or a GPU, or a dedicated processor specialized for a specific process.

[0065] A series of flows of a processing method of the sound processing system 200 will be described with reference to Fig. 7. Fig. 7 is a flowchart for explaining a series of flows of a processing method of the sound processing system 200 in the second embodiment.

[0066] The sound acquisition unit 210 acquires a sound (step S1). The sound acquired by the sound acquisition unit 210 may be either the sound acquired by the extracorporeal sound acquisition unit 211 or the sound acquired by the intracorporeal sound acquisition unit 212, or may be both.

[0067] The sound acquired by the sound acquisition unit 210 is stored as a sound signal in the storage unit 214 (step S2).

[0068] The control unit 217 determines whether or not the sound acquired by the sound acquisition unit 210 includes a sound corresponding to a pre-registered sound (step S3).

[0069] If it is determined in step S3 that the sound acquired by the sound acquisition unit 210 does not include a sound corresponding to a pre-registered sound, the sound output unit 216 outputs the sound signal stored in the storage unit 214 as sound (step S4). At this time, it is desirable that the sound output unit 216 outputs the sound acquired by the sound acquisition unit 210 in close to real time. The sound processing system 100 repeats steps S1 to S4 until it is determined in step 3 that the sound includes a sound corresponding to a pre-registered sound.

[0070] If it is determined in step S3 that the sound includes a sound corresponding to a pre-registered sound, the first sound detection unit 213 determines whether the sound acquired by the sound acquisition unit 210 includes the first sound (step S5).

[0071] If it is determined in step S5 that the sound acquired by the sound acquisition unit 210 includes the first sound, the sound processing system 200 terminates the processing. After this, the sound processing system 200 may perform the operations from step S5 onwards of the sound processing system 100 of the first embodiment.

[0072] If it is determined in step S5 that the sounds acquired by the sound acquisition unit 210 do not include the first sound, the control unit 417 causes the sound output unit 416 to notify the user. The notification may be a voice, alarm, vibration, or the like that notifies the user that a sound corresponding to a pre-registered sound has been emitted. The sound output unit 216 may output a sound corresponding to the sound registered as a notification as an event sound instead of a notification (step S6). The sound output unit 216 may output the event sound using any of the above-described output methods 1 to 5. The sound output unit 216 may output a sound by combining the above-described output methods 1 to 5.

[0073] According to the sound processing system 200 of the second embodiment, even if a desired sound has been generated, it is possible to notify the user that the desired sound has been generated, even if the user may not be aware of it. This reduces the possibility that the user will miss the desired sound.

[0074] Third Embodiment Sound processing systems 300 and 400 according to a third embodiment of the present disclosure will be described with reference to FIG. 8 . FIG. 8 is a diagram illustrating the schematic configuration of the sound processing systems 300 and 400. Here, the sound processing system used by a first user is referred to as the sound processing system 300, and the sound processing system used by a second user is referred to as the sound processing system 400. It is assumed that the first user and the second user are close to each other. A close distance is, for example, a distance at which one user can hear a sound emitted from a surrounding sound source. In other words, it is a distance at which the sound acquisition unit 310 can acquire a sound acquired by the sound acquisition unit 410. The sound processing systems 300 and 400 may have the same configuration. In the sound processing system illustrated in FIG. 8 , only the configuration necessary for explaining the third embodiment is described. The sound processing system 300 used by the first user includes a transmission unit 319. The configuration other than the control unit 317 and the transmission unit 319 is the same as the corresponding configuration in the first or second embodiment, and therefore a description thereof will be omitted. The sound processing system 400 used by the second user includes a receiving unit 420. The configuration other than the control unit 417, the attention eliciting unit 418, and the receiving unit 420 is the same as the corresponding configuration in the first or second embodiment, and therefore a description thereof will be omitted.

[0075] The control unit 317 includes at least one processor. The processor may be, for example, a general-purpose processor such as a CPU or a GPU, or a dedicated processor specialized for a specific process.

[0076] When the control unit 317 determines that the sound acquisition unit 310 of the sound processing system 300 has acquired the first sound, it transmits information that the first sound has been detected to the receiving unit 420 of the sound processing system 400 via the transmitting unit 319. The control unit 317 may also cause the transmitting unit 319 to transmit the second sound determined to be the attention event sound, i.e., the attention event sound, as a sound signal to the receiving unit 420. The transmitting unit 319 may also transmit information related to the attention event sound together with the sound signal of the attention event sound. The information related to the attention event sound includes, for example, at least one of the time the sound was generated, the type of sound, the direction from which the sound is coming, and the position of the sound.

[0077] The receiving unit 420 receives the information transmitted from the transmitting unit 319 .

[0078] The control unit 417 includes at least one processor. The processor may be, for example, a general-purpose processor such as a CPU or a GPU, or a dedicated processor specialized for a specific process.

[0079] The control unit 417 controls the attention eliciting unit 418 to determine whether the second user uttered the first sound within a predetermined time after the receiving unit 420 received information indicating that the first sound was detected, transmitted from the transmitting unit 319. The control unit 417 may determine whether the second user uttered the first sound within a predetermined time after the time when the attention event sound, included in the information transmitted by the transmitting unit 319, occurred. If the second user uttered the first sound, it is considered that the second user was also able to perceive the attention event sound perceived by the first user. However, if the second user did not utter the first sound, it is considered that the second user did not notice this. Therefore, the control unit 417 may control the sound output unit 416 to notify the second user. If the second user did not utter the first sound, the control unit 417 may cause the sound output unit 416 to output the attention event sound received by the receiving unit 419 instead of or as a notification. The control unit 417 may cause the sound output unit 416 to notify the user of information related to the caution event sound.

[0080] When the second user who has received the notification utters the first sound, an attention event sound may be extracted from the second sound and output from the sound output unit 416 in a predetermined output method, in the same or similar manner as the sound processing system 100 of the first embodiment.

[0081] The attention eliciting unit 418 includes at least one processor. The processor may be, for example, a general-purpose processor such as a CPU or a GPU, or a dedicated processor specialized for a specific process.

[0082] 9 and 10, a series of flows of the sound processing method of the sound processing system 300 and the sound processing system 400 will be described. Fig. 9 is a diagram for explaining the series of flows of the sound processing method of the sound processing system 300 in the third embodiment, and Fig. 10 is a diagram for explaining the series of flows of the sound processing method of the sound processing system 400.

[0083] First, the processing flow of the sound processing system 300 will be described.

[0084] The sound acquisition unit 310 acquires a sound (step S1). The sound acquired by the sound acquisition unit 310 may be either the sound acquired by the extracorporeal sound acquisition unit 311 or the sound acquired by the intracorporeal sound acquisition unit 312, or may be both.

[0085] The sound acquired by the sound acquisition unit 310 is stored as a sound signal in the storage unit 314 (step S2).

[0086] The first sound detection unit 313 determines whether the first sound is included in the sounds acquired by the sound acquisition unit 310 (step S3).

[0087] If it is determined in step S3 that the sounds acquired by the sound acquisition unit 310 do not include the first sound, steps S1 to S3 are repeated until it is determined that the sounds include the first sound. At this time, the sounds acquired by the sound acquisition unit 310 may be output in near real time from the sound output unit 316 (not shown), in the same or similar manner as in the first and second embodiments.

[0088] If it is determined in step S3 that the sounds acquired by the sound acquisition unit 310 include the first sound, the sound determination unit 315 determines whether the second sound is an attention event sound (step S4).

[0089] The transmitting unit 319 transmits information that the first sound has been detected to the receiving unit 420 of the sound processing system 400 (step S5). At this time, the transmitting unit 319 may cause the receiving unit 420 of the sound processing system 400 to transmit, as a sound signal, information on at least one of the second sound determined to be the attention event sound, i.e., the attention event sound, and the attention event sound.

[0090] Next, the processing flow of the sound processing system 400 will be described.

[0091] The sound acquisition unit 410 acquires a sound (step S1). The sound acquired by the sound acquisition unit 410 may be either the sound acquired by the extracorporeal sound acquisition unit 411 or the sound acquired by the intracorporeal sound acquisition unit 412, or may be both.

[0092] The sound acquired by the sound acquisition unit 410 is stored as a sound signal in the storage unit 414 (step S2).

[0093] The first sound detection unit 413 determines whether the sounds acquired by the sound acquisition unit 410 include the first sound (step S3).

[0094] In step S3, if the first sound is included in the sounds acquired by the sound acquisition unit 410, an attention event sound is extracted and output by a predetermined output method in the same or similar manner as steps S5 to S6 of the sound processing system 100 of the first embodiment. After that, when the attention event sound has ended, the output is stopped (steps S6 and S7).

[0095] In step S3, if the first sound is not included in the sounds acquired by the sound acquisition unit 410, the control unit 417 determines whether the receiving unit 420 has received from the transmitting unit 319 a notification that the sound acquisition unit 310 has acquired the first sound (step S4).

[0096] In step S4, if the receiving unit 420 has not received from the transmitting unit 319 the notification that the sound acquiring unit 310 has acquired the first sound, steps S1 to S4 are repeated.

[0097] In step S4, when the receiving unit 420 receives from the transmitting unit 319 that the sound acquisition unit 310 has acquired the first sound, it determines whether the second user has uttered the first sound within a predetermined time from the reception (step S5).

[0098] In step S5, when the second user utters the first sound, an attention event sound is extracted and output by a predetermined output method in the same or similar manner as steps S5 to S6 of the sound processing system 100 of the first embodiment. After that, when the attention event sound has ended, its output is stopped (steps S6 and S7).

[0099] If the second user does not utter the first sound in step S5, the control unit 417 controls the sound output unit 416 to notify the second user (step S8). The control unit 417 may notify the second user by causing the sound output unit 416 to output the attention event sound received by the receiving unit 419. If the second user who has received the notification utters the first sound, the attention event sound may be extracted from the second sound and output from the sound output unit 416 in a predetermined output method, in the same way as or similar to the sound processing system 100 of the first embodiment.

[0100] According to the sound processing systems 300 and 400 of the third embodiment, even if a second user misses a second sound to which a first user has selectively focused their attention, the second user can be notified that the second sound has been generated.

[0101] Other Embodiments In a fourth embodiment, the sound processing system 500 includes a sensor that detects the direction in which the user's face is facing, and two or more extracorporeal sound acquisition units 511. The control unit 517 detects the direction in which the user's face is facing using information obtained from the sensor. The control unit 517 also determines the direction of arrival of each of the multiple second sounds acquired by the extracorporeal sound acquisition unit 511. When a second sound is generated, the user may look in the direction of the sound source. Therefore, the control unit 517 may control the sound determination unit 515 to determine that the second sound arriving from the direction in which the user's face is facing is an attention event sound.

[0102] According to the sound processing system 500 of the fourth embodiment, the sound determination unit 515 can accurately determine an attention event sound from the direction to which the user directs selective attention.

[0103] The sound processing system 600 in the fifth embodiment includes two or more external sound acquisition units 611 and two or more internal sound acquisition units 612. The multiple internal sound acquisition units 612 may be arranged in different positions. For example, one of the internal sound acquisition units 612 may be attached to the right ear, and the other may be attached to the left ear. The control unit 617 controls the first detection unit 613 to compare the volume of the sounds acquired by the respective internal sound acquisition units 612. The control unit 617 then controls the sound determination unit 615 to determine that a second sound arriving from the side where the internal sound acquisition unit 612 that acquired a louder sound is located, among the multiple internal sound acquisition units 612, is an attention event sound.

[0104] In the sound processing system 600 of the fifth embodiment, the direction in which the user is directing his or her attention can be determined without using a sensor that detects the user's movements, and attention event sounds can be determined with high accuracy.

[0105] <Effects of this embodiment> The sound processing system of the above-described embodiment includes a sound acquisition unit that acquires sounds including a first sound, which is a sound uttered by a user and has a duration of a first time or less, and a second sound, which is a sound other than the sound uttered by the user, and a sound output unit that outputs the second sound when the first sound is acquired. By outputting the second sound after the first sound is acquired, it becomes possible to easily select a sound.

[0106] The present technology can also be configured as follows.

[0107] (1) A sound processing system includes a sound acquisition unit that acquires a first sound, which is a sound uttered by a user and has a duration equal to or shorter than a first time, and a second sound, which is a sound other than the sound uttered by the user; and a sound output unit that outputs the second sound when the sound acquisition unit acquires the first sound.

[0108] (2) In the sound processing system described in (1) above, the second sound is a sound from a sound source that generated a sound before the user uttered the first sound.

[0109] (3) In the sound processing system described in (1) or (2) above, the sound acquisition unit acquires the second sound from the plurality of sound sources, and the sound output unit outputs the second sound from one of the plurality of sound sources that was generating a sound most recently before the user uttered the first sound.

[0110] (4) In the sound processing system described in (1) to (3) above, the first time period is two moras.

[0111] (5) In the sound processing system described in (1) to (4) above, the sound acquisition unit includes an internal body sound acquisition unit that acquires sound arriving from inside the user's body, and has a control unit that determines whether the sound acquired by the internal body sound acquisition unit is a first sound.

[0112] (6) In the sound processing system described in (5) above, the control unit determines whether the sound acquired by the internal sound acquisition unit is a first sound based on the volume.

[0113] (7) In the sound processing system described in (6) above, the sound acquisition unit includes an extracorporeal sound acquisition unit that acquires sound coming from outside the user's body, and the control unit determines whether the sound acquired by the intracorporeal sound acquisition unit is a first sound based on the relationship in volume between the sounds acquired by the intracorporeal sound acquisition unit and the extracorporeal sound acquisition unit.

[0114] (8) In the sound processing system described in (1) to (7) above, the sound output unit notifies the user when the user does not utter the first sound after the sound acquisition unit acquires a sound similar to a pre-registered sound.

[0115] (9) A sound processing device includes: a sound acquisition unit that acquires sounds including a first sound, which is a sound uttered by a user and has a duration equal to or shorter than a first time, and a second sound, which is a sound other than the sound uttered by the user; and a sound output unit that outputs the second sound when the sound acquisition unit acquires the first sound.

[0116] (10) A sound processing method includes: a sound acquisition unit acquiring sounds including a first sound, which is a sound uttered by a user and has a duration equal to or shorter than a first time, and a second sound, which is a sound other than the sound uttered by the user; and a sound output unit outputting the second sound when the sound acquisition unit acquires the first sound.

[0117] (11) The sound processing program causes the sound acquisition unit to acquire sounds including a first sound, which is a sound uttered by a user and has a duration equal to or shorter than a first time, and a second sound, which is a sound other than the sound uttered by the user; and when the sound acquisition unit acquires the first sound, causes the sound output unit to output the second sound.

[0118] While the present disclosure has been described based on the drawings and embodiments, it should be noted that those skilled in the art can easily make various modifications and alterations based on the present disclosure. Therefore, it should be noted that these modifications and alterations are within the scope of the present disclosure. For example, the functions included in each component can be rearranged so as not to cause logical inconsistencies. Multiple components can be combined into one or separated. The above-described embodiments of the present disclosure are not limited to faithful implementation of each of the described embodiments, but can be implemented by combining features or omitting some of them as appropriate. In other words, those skilled in the art can make various modifications and alterations to the contents of the present disclosure based on the present disclosure. Therefore, these modifications and alterations are within the scope of the present disclosure. For example, in each embodiment, each functional unit, each means, each step, etc. can be added to other embodiments so as not to cause logical inconsistencies, or can be replaced with each functional unit, each means, each step, etc. of other embodiments. Furthermore, in each embodiment, multiple functional units, each means, each step, etc. can be combined into one or separated. Furthermore, each of the above-described embodiments of the present disclosure is not limited to being implemented faithfully according to each of the described embodiments, but can also be implemented by combining each feature or omitting some of them as appropriate.

[0119] Although the embodiments of the present disclosure have been described mainly in terms of a system, the embodiments of the present disclosure may also be realized as a method including steps executed by each component of the system. The embodiments of the present disclosure may also be realized as a method executed by a processor included in the system, a program, or a storage medium on which a program is recorded. It should be understood that these are also encompassed within the scope of the present disclosure.

[0120] 100, 200, 300, 400, 500, 600 Sound processing system 110, 210, 310, 410 Sound acquisition unit 111, 211, 311, 411, 511, 611 Extracorporeal sound acquisition unit 112, 212, 312, 412, 612 Internal sound acquisition unit 113, 213, 313, 413, 613 First sound detection unit 114, 214, 314 Storage unit 115, 215, 315, 515, 615 Sound determination unit 116, 216, 416 Sound output unit 117, 217, 317, 417, 517, 617 Control unit 218, 418 Attention induction unit 319 Transmission unit 420 Reception unit

Claims

1. A sound processing system comprising: a sound acquisition unit that acquires sounds including a first sound, which is a sound uttered by a user and has a duration of a first time or less, and a second sound, which is a sound other than the sound uttered by the user; and a sound output unit that outputs the second sound when the sound acquisition unit acquires the first sound.

2. The sound processing system according to claim 1, wherein the second sound is a sound from a sound source that generated a sound before the user uttered the first sound.

3. The sound processing system according to claim 1 or 2, wherein the sound acquisition unit acquires a second sound from a plurality of the sound sources, and the sound output unit outputs the second sound from one of the plurality of sound sources that generated a sound most recently before the user uttered the first sound.

4. The sound processing system according to any one of claims 1 to 3, wherein the first time period is two moras.

5. A sound processing system according to any one of claims 1 to 4, wherein the sound acquisition unit includes an internal body sound acquisition unit that acquires sound arriving from within the user's body, and has a control unit that determines whether the sound acquired by the internal body sound acquisition unit is a first sound.

6. The sound processing system according to claim 5, wherein the control unit determines whether the sound acquired by the internal body sound acquisition unit is the first sound based on the volume.

7. The sound processing system according to claim 6, wherein the sound acquisition unit includes an extracorporeal sound acquisition unit that acquires sound coming from outside the user's body, and the control unit determines whether the sound acquired by the intracorporeal sound acquisition unit is a first sound based on the relationship in volume between the sounds acquired by the intracorporeal sound acquisition unit and the extracorporeal sound acquisition unit.

8. The sound processing system according to any one of claims 1 to 7, wherein the sound output unit notifies the user when the user does not utter the first sound after the sound acquisition unit acquires a sound similar to a pre-registered sound.

9. A sound processing device comprising: a sound acquisition unit that acquires sounds including a first sound, which is a sound uttered by a user and has a duration of a first time or less, and a second sound, which is a sound other than the sound uttered by the user; and a sound output unit that outputs the second sound when the sound acquisition unit acquires the first sound.

10. A sound processing method comprising: a sound acquisition unit acquiring a sound including a first sound, which is a sound uttered by a user and has a duration equal to or less than a first time, and a second sound, which is a sound other than the sound uttered by the user; and a sound output unit outputting the second sound when the sound acquisition unit acquires the first sound.

11. A sound processing program that causes a control unit to execute the following: causing a sound acquisition unit to acquire a sound including a first sound, which is a sound uttered by a user and has a duration of a first time or less, and a second sound, which is a sound other than the sound uttered by the user; and, when the sound acquisition unit acquires the first sound, causing a sound output unit to output the second sound.

Citation Information

Patent Citations

  • Hearing aid, hearing-aid processing method and integrated circuit for hearing-aid

    JP2010011447A

  • Estimation apparatus, estimation method, and estimation program

    JP2017228164A

  • Throat microphone system and method

    US20200128317A1