Output system, output device, and output method
The output system addresses the issue of missed ambient sound by using speech interval-based stimuli, ensuring user engagement and awareness of surroundings.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-25
- Publication Date
- 2026-04-06
AI Technical Summary
Conventional earphones with a hear-through mode may cause users to miss ambient sound.
An output system that includes an acquisition unit to capture speech and an output unit to provide visual, auditory, or tactile stimuli based on the interval between beats in the speech, enhancing user engagement.
The system effectively provides stimuli to users based on speech, ensuring they do not miss ambient sound and enhancing their interaction with the environment.
Smart Images

Figure 2026058785000001_ABST
Abstract
Description
Technical Field
[0001] This Disclosure relates to an output system, an output device, and an output method.
Background Art
[0002] Conventionally, earphones equipped with a hear-through mode for providing acquired ambient sound to a user have been disclosed. However, the user may miss the ambient sound.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] This Disclosure aims to provide an output system that stimulates a user based on speech included in sound and supports the user in hearing the speech.
Means for Solving the Problems
[0005] To solve the above problems, an output system according to an embodiment includes an acquisition unit that acquires sound including speech, and an output unit that outputs a visual, auditory, or tactile stimulus. The output unit outputs the stimulus based on the length of the interval between beats included in the speech when the speech is being made.
[0006] Also, an output device according to an embodiment includes an acquisition unit that acquires sound including speech, and an output unit that outputs a visual, auditory, or tactile stimulus. The output unit outputs the stimulus based on the length of the interval between beats included in the speech when the speech is being made. This
[0007] Furthermore, an output method according to one embodiment includes an acquisition unit acquiring sound including speech, and an output unit outputting visual, auditory, or tactile stimuli, wherein the output unit outputs the stimuli based on the length of the interval between beats included in the speech while the speech is being performed. [Effects of the Invention]
[0008] According to one embodiment, it is possible to provide an output system, output device, or output method that can provide a stimulus to a user based on speech. [Brief explanation of the drawing]
[0009] [Figure 1] This is a diagram illustrating the schematic configuration of output system 1. [Figure 2] This diagram illustrates the parameters of a stimulus when the stimulus is vibration. [Figure 3] This diagram shows that chunks are output at output intervals that are multiplied by the set mora length. [Figure 4] This diagram shows the relationship between the sound played by the first output unit 120 and the chunk output by the second output unit 240. [Figure 5] This is a flowchart illustrating the processing flow of the output device 100. [Figure 6] This figure shows that in the output system 1 of Example 2, the second output unit 240 outputs chunks based on the speech envelope of the sound signal. [Figure 7] This figure shows that in the output system 1 of Example 3, the second output unit 240 outputs chunks based on the speech envelope of the sound signal. [Modes for carrying out the invention]
[0010] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In each figure, the same reference numerals indicate components having the same or equivalent functions. Note that the configurations, numerical values, processing flows, functions, elements, etc. described in the following embodiments are merely examples, and modifications and changes thereof are possible freely, and the Disclosure technical scope of the present invention is not intended to be limited to the following description.
[0011] <Example 1> FIG. 1 is a diagram for explaining the schematic configuration of the output system 1 in Example 1.
[0012] The output system 1 includes an output device 100 and a control device 200. The output device 100 and the control device 200 may be configured as separate devices or may be configured as one device.
[0013] The output device 100 is a device that outputs sound to a user. The output device 100 is, for example, an earphone. The output device 100 is not limited to an earphone and may be a headset, a headphone, a speaker, a smartphone, or the like.
[0014] The schematic configuration of the output device 100 will be described. The output device 100 includes an acquisition unit 110, a first output unit 120, and a first communication unit 130.
[0015] The acquisition unit 110 is, for example, a microphone. The acquisition unit 110 acquires ambient sound and converts it into a sound signal. The ambient sound includes speech. Speech is the sound generated by a language being uttered as sound. Speech includes, for example, speech uttered by the user, speech uttered by a person around the user, and speech emitted from a sound source (such as a speaker) around the user.
[0016] The acquisition unit 110 may be provided outside the output device 100 and may be configured to be communicable with one or both of the output device 100 and the control device 200 by wire or wirelessly. The acquisition unit 110 may be configured to be included in the configuration of the control device 200.
[0017] The first output unit 120 is, for example, a speaker. The first output unit 120 outputs sound to the user. When the output device 100 is a so-called canal-type earphone, the acquired sound obtained by the acquisition unit 110 can be output from the first output unit 120 in real time as it is, thereby enabling a so-called hear-through mode to function. When the output device 100 is a so-called open-ear type earphone and the user can directly hear the surrounding sound, the first output unit 120 may not be included in the output device 100.
[0018] The first communication unit 130 is, for example, a communication module, and is connected to the control device 200 to perform communication. The communication module is a communication module corresponding to an arbitrary communication standard. The communication standard is, for example, a wired communication standard or a short-range wireless communication standard including Bluetooth (registered trademark), infrared rays, NFC, etc.
[0019] The output device 100 can transmit the sound signal of the sound acquired by the acquisition unit 110 to the control device 200 via the first communication unit 130.
[0020] Next, the schematic configuration of the control device 200 will be described.
[0021] The control device 200 may be, for example, a smartphone, a smartwatch, smart glasses, etc., and is a terminal capable of controlling the output device 100.
[0022] The control device 200 includes a second communication unit 210, an input unit 220, a second output unit 240, a storage unit 230, and a control unit 250. One or more components included in the control device 200 are the output device 10 It may be included in 0. One or more components included in the control device 200 may be located outside the control device 200 and configured to be connectable to one or both of the control device 200 and the output device 100 by wire or wireless connection. For example, the storage unit 230 may be located on a remote server configured to communicate with the control device 200 via a network system or the like. For example, the second output unit 240 may be configured as a device configured to communicate with either or both of the output device 100 and the control device 200 via a network system or the like.
[0023] The second communication unit 210 is, for example, a communication module, which is connected to the first communication unit 130 and used for communication between the output device 100 and the control device 200. The communication module may support the same communication standard as the first communication unit 130. By connecting the second communication unit 210 and the first communication unit 130, communication can be performed between the output device 100 and the control device 200. The second communication unit 210 can receive sound signals transmitted from the output device 100 via the first communication unit 130.
[0024] The input unit 220 includes an input interface capable of receiving input from the user. The input unit 220 may include, for example, one or more buttons, a touch panel, etc. The input unit 220 may be used by the user to select a process to be executed by the output device 100.
[0025] The second output unit 240 outputs a stimulus to the user. The second output unit 240 may output a stimulus when the acquisition unit 110 is acquiring ambient sound. The second output unit 240 may be configured to include one or more of the following: the display unit 241, the light-emitting unit 242, the vibration unit 243, and the sound-producing unit 244.
[0026] The display unit 241 is, for example, a display located on the surface of the control device 200. The display unit 241 displays an image or video and outputs a visual stimulus to the user.
[0027] If the input unit 220 is a touch panel display or the like, the control device 200 Display section 241 Alternatively, instead of providing a separate unit, the input unit 220 of the control device 200 may be used as the display unit 241.
[0028] The light-emitting unit 242 is, for example, an LED (Light Emitting) placed on the surface of the control device 200. It may be a diode, etc. The light-emitting unit 242 emits light and outputs a visual stimulus to the user.
[0029] The vibrating unit 243 includes, for example, a vibrating element such as a piezoelectric element. The vibrating unit 243 vibrates and outputs tactile stimulation to the user.
[0030] The sound-producing unit 244 is, for example, a speaker. The sound-producing unit 244 generates sound and outputs auditory stimuli to the user. Instead of providing the sound-producing unit 244 in the control device 200, the first output unit 120 of the output device 100 may be used as the sound-producing unit 244.
[0031] The storage unit 230 is a storage medium including ROM (Read Only Memory) or RAM (Random Access Memory), etc. The storage unit 230 may store a program executed by the output system 1. The program may include a sound processing program executed by the output system 1.
[0032] The control unit 250 is configured to include at least one processor, at least one dedicated circuit, or a combination thereof. The processor may be a general-purpose processor such as a CPU or GPU, or a dedicated processor specialized for a specific process. The dedicated circuit may be, for example, an FPGA or ASIC. The control unit 250 executes processes related to the operation of the output system 1 while controlling each component of the output system 1.
[0033] The components included in the control unit 250 may consist of hardware, software, or both. Each component included in the control unit 250 may be controlled by the control unit 250. The control unit 250 can perform various processes on the output system 1 in response to input from the input unit 220 by the user.
[0034] The control unit 250 includes a stimulus parameter setting unit 251, a sound processing unit 252, and an output control unit 253. One or more components included in the control unit 250 may be provided outside the control unit 250.
[0035] The stimulus parameter setting unit 251 calculates and stores the parameters used by the second output unit 240 when outputting a stimulus. The parameters are set by the user by inputting them into the input unit 220 and may be set to user-specific values.
[0036] This section describes the parameters when the stimulus is a tactile stimulus, i.e., vibration. Figure 2 is a diagram illustrating the parameters of a stimulus when the stimulus is vibration. When the stimulus is vibration, the parameters include vibration frequency, vibration amplitude, and vibration duration. The vibration frequency is the number of times the vibration repeats per unit time. The vibration frequency may be, for example, 150-250 Hz. The vibration amplitude indicates the intensity of the vibration. The vibration intensity may be an intensity that can be perceived by a human. The vibration duration is the length of time the vibration output continues. The vibration duration may be, for example, 100 ms or more.
[0037] If the stimulus is a visual stimulus, the parameters include duration and color. Duration is the length of time that the display unit 241 or light-emitting unit 242 maintains a predetermined display or light emission. The duration may be, for example, 100 ms or more.
[0038] The color is the color of the visual stimulus (display or light emission) output by the display unit 241 or the light-emitting unit 242. The color output by the display unit 241 or the light-emitting unit 242 may be a color based on the emotional information contained in the utterance acquired by the acquisition unit 110. The emotional information is information that indicates the emotions of the speaker making the utterance, and includes information that indicates negative emotions such as sadness and anger, and positive emotions such as joy, happiness, and contentment. The stimulus parameter setting unit 251 may estimate the emotions of the person who made the utterance using a known method. For example, if the emotional information of the utterance acquired by the acquisition unit 110 indicates a positive emotion, the color may be a warm color (e.g., red, orange, yellow). If the emotional information of the utterance acquired by the acquisition unit 110 indicates a negative emotion, the color may be a cool color (e.g., green, blue, purple).
[0039] If the stimulus is an auditory stimulus, the parameters include frequency, volume, and duration. The frequency of the sound is a parameter that indicates the pitch of the sound. The frequency may be set based on the fundamental frequency obtained from the speech interval of the utterance acquired by the acquisition unit 110. The stimulus parameter setting unit 251 may obtain the fundamental frequency from the speech interval of the utterance using a known method such as the autocorrelation method. The stimulus parameter setting unit 251 may set the frequency to, for example, twice the acquired fundamental frequency. If the utterance acquired by the acquisition unit 110 is male (fundamental frequency = 80~200Hz), the sound frequency may be set to 100Hz~400Hz, and if it is female (fundamental frequency = 150~400Hz), the frequency may be set to 200~800Hz. Volume is the loudness of the sound. Duration is the length of time that the sound output unit 244 sustains the sound output. The duration may be, for example, 100ms or more.
[0040] The stimulus parameter setting unit 251 may prompt the user to perform an input operation to set the value of each parameter. The stimulus parameter setting unit 251 may allow the user to set the value of each parameter within a predetermined range from a predetermined upper limit to a predetermined lower limit. The stimulus parameter setting unit 251 may store the value of the parameter selected by the user.
[0041] The second output unit 240 outputs a stimulus based on the parameters held by the stimulus parameter setting unit 251. In this specification, the stimuli output by the second output unit 240 during a single duration may be collectively referred to as a chunk.
[0042] The sound processing unit 252 processes the sound signal acquired from the control unit 200.
[0043] The sound processing unit 252 detects speech intervals from the sound signal acquired from the control device 200. The sound processing unit 252 may detect speech intervals by performing speech recognition processing on the sound signal. A speech interval is a period in which the speech state continues. A speech interval may be a period in which speech continues at a predetermined volume or higher. If the period of speech at a volume lower than the predetermined volume, or if no speech is made, is within a predetermined time, it may be considered that speech is continuing. The predetermined time may be, for example, 400ms. The start point of a speech interval is also referred to as the "start of speech". The end point of a speech interval is also referred to as the "end of speech". The sound processing unit 252 may detect non-speech intervals. A non-speech interval is a period in which there is no speech state. A non-speech interval may be a period in which the volume of speech is lower than a predetermined volume. If the period of speech at a volume lower than the predetermined volume, or if no speech is made, is longer than a predetermined time, it may be considered a non-speech interval. The specified time may be, for example, 400ms. The utterance interval may be the interval between two non-utterance intervals.
[0044] The speech interval may be the interval of speech output from a single sound source.
[0045] The sound processing unit 252 detects the number of morae contained within a predetermined time period (sometimes referred to as the first time period) within the speech segment based on the sound signal of the detected speech segment. A mora is a beat and beat During by distanceYes. The length of the interval between beats may be the length between the start of one beat and the start of the next beat. The length of the interval between beats may be extracted based on the length of the moras included in the utterance. The first time is a time shorter than the utterance interval and may be the time it takes to utter multiple moras. The first time may be set in advance and may be the time during which at least two or more moras are generated. The first time may be, for example, one second. For the first time, the sound processing unit 252 may, for example, detect the number of moras from the text information of the utterance converted from the sound signal using speech recognition.
[0046] The sound processing unit 252 may detect the number of syllables instead of the number of moras. The length of the interval between beats may include syllables.
[0047] The sound processing unit 252 calculates the average mora length based on the number of morae detected and the duration of the speech segment targeted for detection (first time). The average mora length is the average duration of the morae included in the speech segment. The average mora length may be the number of morae included in the speech segment divided by the first time. The sound processing unit 252 may retain the calculated average mora length as the set mora length. In this specification, the retained average mora length may be referred to as the set mora length. The sound processing unit 252 may also calculate the average mora length based on the time it takes for the number of morae detected from the sound signal in one speech segment to reach a predetermined number (e.g., 2), without using a predetermined time set in advance.
[0048] The sound processing unit 252 calculates the average mora length, which allows for the calculation of the average length of the interval between beats in an utterance.
[0049] The sound processing unit 252 calculates the output interval at which the second output unit 240 outputs a chunk, based on the set mora length. The second output unit 240 outputs the chunk multiple times at each output interval calculated by the sound processing unit 252.
[0050] The sound processing unit 252 may calculate the output interval as an interval obtained by multiplying the set mora length. That is, the output interval may be a length obtained by multiplying the length of the interval between beats. The multiplication factor may be set by the user by inputting it into the input unit 220. Figure 3 shows the output of chunks at an output interval obtained by multiplying the set mora length. In Figure 3, the set mora length is the length of the interval between the dotted lines. Figure 3(a) shows the output of chunks at an output interval of 3 times the set mora length. Figure 3(b) shows the output of chunks at an output interval of 4 times the set mora length. Figure 3(c) shows the output of chunks at an output interval of 5 times the set mora length. Figure 3(d) shows the output of chunks at an output interval of 6 times the set mora length.
[0051] The length of the output interval may be longer than the vibration duration. The length of time during the output interval when no chunks are output may be within a predetermined time range. The output interval may be set, for example, based on equation (1) shown below. 300ms < (Output interval - Oscillation duration) ≦ 800ms...Equation (1) (Output interval - Oscillation duration) is the length of time within the output interval when no chunks are being output.
[0052] Table 1 shows the output interval when the duration is 100 ms. The first row of the table is the set mora length. The first column of the table is the multiplication factor set by the user. The sound processing unit 252 may calculate the output interval as shown in Table 1, for example. For example, if the set mora length is 1 / 6 s and the multiplication factor is 1:3, the sound processing unit 252 can calculate the output interval as 500 ms (1 / 6 s × 3). For example, if the set mora length is 1 / 6 s Therefore, if the multiplication ratio is 1:6, it can be calculated as 1 / 6s × 6 = 1000ms. However, based on equation (1), it is acceptable to calculate the output interval as 900ms, which is close to 1000ms.
[0053] [Table 1]
[0054] The output control unit 253 causes the stimulus (chunk) of the parameter held by the stimulus parameter setting unit 251 to output to the second output unit 240 based on the output interval calculated by the sound processing unit 252. The output control unit 253 may output the chunk to the second output unit 240 when the acquisition unit 110 acquires the utterance contained in the sound, that is, when an utterance is taking place. The output control unit 253 also determines the average mora length, that is, the length of the interval between beats contained in the utterance. Average Based on this, the second output unit 240 is made to output a stimulus.
[0055] The output control unit 253 causes the second output unit 240 to output chunks based on the speech and non-speech intervals detected by the sound processing unit 252. Specifically, the output control unit 253 causes the second output unit 240 to output chunks during the speech interval at a calculated output interval. When the first output unit 120 outputs the sound acquired by the acquisition unit 110, a slight delay may occur. If such a delay occurs, the output control unit 253 may control the timing of chunk output so that chunks are output at an appropriate output interval during the speech interval. If the non-speech interval of the sound acquired from the output device 100 is longer than a predetermined length (e.g., 400 ms), the second output unit 240 may weaken the chunk stimulus output or stop outputting the chunks. smaller If it is a non-speech interval, the chunk may be continuously output to the second output unit 240 at the calculated output interval.
[0056] Figure 4 shows the relationship between the sound played by the first output unit 120 and the output of chunks by the second output unit 240. Figure 4(a) shows the sound signal of the sound output by the first output unit 120. Figure 4(b) shows whether the sound signal in Figure 4(a) is in a speech interval or a non-speech interval. Figure 4(c) shows the output of chunks by the second output unit 240.
[0057] In the first utterance section in Figure 4(b), the output control unit 253 may start outputting chunks at an output interval calculated from the sound signals of the utterance section within a predetermined time from the start of the utterance. Since the first non-utterance section in Figure 4(b) is shorter than 400ms, it may be considered that the utterance section is continuing, and the output of the chunks that were output in the first utterance section may be continued. In the second non-utterance section in Figure 4(c), since it lasts longer than 400ms, after the non-utterance section has lasted for 400ms, the output control unit 253 may consider that the utterance section has been interrupted and reduce the size of the chunks output by the second output unit 240.
[0058] The output control unit 253 may vary the output interval based on the average mora length calculated first by the sound processing unit 252 and the set mora length held in the sound processing unit 252. For example, in the third utterance section following the second non-utterance section, the output control unit 253 may vary the output interval using the average mora length calculated from the first and second utterance sections.
[0059] The output control unit 253 may cause the first output unit 120 to output a stimulus if the stimulus to be output is an auditory stimulus. If the first output unit 120 is a stereo earphone consisting of two earphones, one for the right ear and one for the left ear, it may cause the sound including speech to be output to the speaker of one earphone and the auditory stimulus to be output to the speaker of the other earphone.
[0060] Next, we will explain the output method, which is the processing flow of the output system 1, using Figure 5. Figure 5 is a flowchart for explaining the processing flow of the output device 100.
[0061] The sound processing unit 252 receives the sound signal acquired by the acquisition unit 110 (Step 1).
[0062] The sound processing unit 252 detects the speech segment of the acquired sound signal (step 2).
[0063] The sound processing unit 252 extracts the number of morae contained in the utterance during the first hour of the speech interval (step 3).
[0064] The sound processing unit 252 is, Sound processing unit 252 Based on the number of moras extracted and the first hour, the average mora length is calculated (Step 4).
[0065] The sound processing unit 252 determines whether the set mora length is retained in the sound processing unit 252 (step 5). If the set mora length is not retained, a new utterance interval has started, and if the set mora length is retained, existing This refers to a situation where the utterance interval continues without interruption.
[0066] In step 5, if the set mora length is not retained, the sound processing unit 252 sets the average mora length calculated in step 4 as the set mora length (step 6).
[0067] In step 5, if the set mora length is retained, the sound processing unit 252 determines whether the value calculated by the following equation (2) is greater than a predetermined value a (step 7). |Average mora length - set mora length| / set mora length···Equation (2) The predetermined value a may be, for example, 0.2.
[0068] In step 7, if the value calculated by equation (2) is less than or equal to a predetermined value a, the sound processing unit 252 sets the average mora length as the new set mora length (step 6).
[0069] In step 7, the value calculated by formula (2) is the predetermined value. a If it is greater, the sound processing unit 252 will hold the set mora length if the average mora length is greater than the set mora length. ( 1+a ) ,or ( 1-a ) The length is doubled and set as the new set mora length (step 8). The sound processing unit 252 adjusts the new set mora length so that it is close to the calculated average mora length, depending on the relationship between the set mora length and the average mora length. ( 1+a ) ,or( 1-a ) You may choose to multiply the set mora length by either of the two options to obtain the new set mora length. By performing steps 6 to 8, the set mora length can be updated even if the average mora length during the speech interval changes. This allows stimuli to be output to the user at appropriate output intervals.
[0070] The sound processing unit 252 holds the calculated set mora length (step 9).
[0071] The stimulus parameter setting unit 251 generates chunks based on the preset stimulus parameters (step 10).
[0072] Sound processing unit 252 This includes the set mora length held in the sound processing unit 252, and the user's pre-set Based on the multiplier, the output interval is calculated (step 11).
[0073] The output control unit 253 determines whether the sound acquired by the acquisition unit 110 is in a speech interval or a non-speech interval (step 12).
[0074] In step 12, if the sound acquired by the acquisition unit 110 is within a speech interval, the output control unit 253 causes the second output unit 240 to output chunks at the calculated output interval. (Step 13) The output control unit 253 can output stimuli in accordance with the sound the user is hearing.
[0075] In step 12, if the sound acquired by the acquisition unit 110 is in a non-speech section, Sound processing unit 252 This determines whether the non-speech interval is longer than a predetermined length (step 14).
[0076] In step 14, the non-speech interval is of a predetermined length. shorter In this case, the chunks are output at the calculated output interval (step 13).
[0077] Step 14In this case, the non-speech interval is of a predetermined length. The above In this case, the sound processing unit 252 resets the retained set mora length (step 15).
[0078] The second output unit 240 is, Reduce the chunk amplitude (step 16), Output chunks with reduced amplitude at the calculated output interval (step 13 ).
[0079] <Example 2> Figure 6 shows that in the output system 1 of Embodiment 2, chunks are output based on the speech envelope of the sound signal. Figure 6(a) shows the sound signal output by the first output unit 120. Figure 6(b) shows 6 Figure 6(a) shows whether the sound signal in Figure 6(a) is in the speech or non-speech section. Figure 6(c) shows the speech envelope calculated by the sound processing unit 252 from the sound signal in Figure 6(a). Figure 6(d) shows the second output unit 240 6 This figure shows that the second output unit outputs chunks with an amplitude or magnitude of 240 based on the value of the speech envelope in (c).
[0080] In the output system 1 according to Example 2, the sound processing unit 252 in Example 1 calculates a speech envelope that indicates the strength of the sound signal during the speech interval of the acquired sound signal. The speech envelope is obtained by extracting the envelope component of the power (squared value) of the acquired sound signal. The envelope may be extracted using well-known methods such as Hilbert transform or short-time Fourier transform.
[0081] The output control unit 253 detects the value of the speech envelope corresponding to the sound output from the first output unit 120 at the time the second output unit 240 outputs the chunk. Based on the detected value of the speech envelope, the output control unit 253 may determine the vibration amplitude or magnitude of the chunk output by the second output unit 240. If the value of the speech envelope is large, the output control unit 253 may increase the vibration amplitude or magnitude of the stimulus output by the second output unit 240. If the value of the speech envelope fluctuates only slightly, the output control unit 253 may decrease the vibration amplitude or magnitude of the stimulus output by the second output unit 240. By varying the vibration amplitude or magnitude of the stimulus output by the second output unit 240 based on the value of the speech envelope, it is expected that a strong sense of synchronization between sound and stimulus will be felt, and attention to the sound will be promoted.
[0082] <Example 3> Figure 7 shows how the output system 1 of Example 3 outputs chunks based on the speech envelope of the sound signal. 7 (a) is a diagram showing the sound signal output by the first output unit 120. 7 (b) is Figure 7 (a) This figure shows whether the sound signal of the sound in (a) is in the speech or non-speech interval. 7 (c) shows the sound processing unit 252 in Figure 7 (a) Speech envelope calculated from the sound signal This is a diagram showing lines. 7 (d) shows the second output unit 240 in Figure 7 This figure shows that chunks are output at output intervals based on the peak interval of the speech envelope in (c).
[0083] In the output system 1 according to Example 3, the sound processing unit 252 in Example 1 calculates the speech envelope of the acquired sound signal during the speech interval. The speech envelope is obtained by extracting the envelope component of the power (squared value) of the acquired sound signal. The envelope may be detected using well-known methods such as the Hilbert transform or the short-time Fourier transform.
[0084] Instead of the sound processing unit 252 in Example 1 calculating the average mora length, the sound processing unit 252 in Example 3 calculates the length between peaks of the calculated speech envelope. , The average peak interval may be calculated by averaging the lengths between multiple peaks within a speech interval. The length between peaks is the length of the interval between beats. The length of the interval between beats may be extracted based on the accent of the utterance. By calculating the average peak interval, the sound processing unit 252 can calculate the average length of the interval between beats included in the utterance.
[0085] The sound processing unit 252 may calculate the output interval based on the calculated average peak interval instead of the average mora length in Example 1.
[0086] By calculating the output interval based on the intervals between peaks in the speech envelope, it is possible to output stimuli based on the accent of the speech.
[0087] <Example 4> In the output system 1 according to Embodiment 4, the acquisition unit 110 may be included in the control device 200. The acquisition unit 110 may acquire sound received by the control device 200 via a communication line. In this case, the acquisition unit 110 may be configured as software as a function of the control unit 250. For example, if the output device 100 is a smartphone, the acquisition unit 110 may acquire sound arriving from another device via a communication line. The other device may be the device of the person making the call, or it may be a server that provides content such as video or audio.
[0088] The sound acquired by the acquisition unit 110 may be transmitted from the control device 200 to the output device 100 via the first communication unit 130 and the second communication unit 210. The output device 100 may output the sound transmitted from the control device 200 from the first output unit 120.
[0089] <Example 5> In the output system 1 according to Embodiment 5, the acquisition unit 110 may be configured to acquire the sound output by the first output unit 120. The control device 200 may output a stimulus based on the sound output by the first output unit 120.
[0090] <Effects of this embodiment> Book Disclosure The output system comprises an acquisition unit that acquires sounds including speech, and an output unit that outputs visual, auditory, or tactile stimuli. The output unit outputs the stimuli based on the length of the interval between beats included in the speech while the speech is being spoken, thereby preventing the sound from being missed.
[0091] Furthermore, this technology can also be configured as follows.
[0092] (1) The output system comprises an acquisition unit that acquires sounds including speech, and an output unit that outputs visual, auditory, or tactile stimuli, wherein the output unit outputs the stimuli based on the length of the interval between beats included in the speech while the speech is being performed.
[0093] (2) In the output system described in (1) above, the output unit outputs the stimulus when the acquisition unit is acquiring sound.
[0094] (3) In the output system described in (1) or (2) above, Between beats The interval is, These are the moras included in the aforementioned utterance. .
[0095] (4) In the output system described in (1) or (3) above, the output unit outputs the stimulus at a first output interval based on the average length of a plurality of morae included in the utterance.
[0096] (5) (1) above ~ ( 4 In the output system described above, First output The interval is as described above. Mora length Average It is the length obtained by multiplying by a factor of 1.
[0097] (6) (1) above ~ ( 5 In the output system described in ( ), the output unit outputs the stimulus in accordance with the sound the user is listening to.
[0098] (7) (1) above ~ ( 6 In the output system described above, Between beats The interval is, This is the peak interval of the speech envelope included in the aforementioned utterance. .
[0099] (8) In the output system described in (1) to (7) above, the output unit outputs the stimulus at a second output interval based on the average of the lengths of the multiple peak intervals included in the utterance.
[0100] (9) (1) above ~ ( 8 In the output system described above, The output unit controls the volume of the sound to a predetermined level. smaller The stimulus is output based on the length of the interval between beats contained in the utterance during the interval.
[0101] (10) (1) above ~ ( 9 In the output system described in ( ), the output unit outputs the stimulus based on the length of the interval between beats included in the speech output from a single sound source, which is included in the sound.
[0102] (11) The output device comprises an acquisition unit that acquires sound including speech, and an output unit that outputs visual, auditory, or tactile stimuli, wherein the output unit outputs the stimuli based on the length of the interval between beats included in the speech while the speech is being spoken.
[0103] (12) The output method involves the acquisition unit acquiring sound including speech, The output unit outputs visual, auditory, or tactile stimuli, The output unit outputs the stimulus based on the length of the interval between beats included in the utterance when the utterance is being spoken.
[0104] That's all. Disclosure This has been explained using embodiments, Disclosure The technical scope is not limited to the scope described in the above embodiments, and various modifications and changes are possible within the scope of its gist. For example, all or part of the device can be configured as a state determination system by functionally or physically distributing and integrating them in any unit. Furthermore, new embodiments resulting from any combination of multiple embodiments are also possible. Disclosure This is included in the embodiment. The effects of the new embodiment resulting from the combination also possess the effects of the original embodiment. [Explanation of symbols]
[0105] Output System 1 Output device 100 Acquisition part 110 First output unit 120 1st Communications Department 130 Control device 200 Second Communications Department 210 Input section 220 Storage section 230 Second output unit 240 Display section 241 Light-emitting part 242 Vibrating part 243 Pronunciation section 244 Control unit 250 Stimulation parameter setting unit 251 Sound processing unit 252 Output control unit 253
Claims
1. An acquisition unit that acquires sounds including speech, It comprises an output unit that outputs visual, auditory, or tactile stimuli, The output unit outputs the stimulus based on the length of the interval between beats included in the utterance when the utterance is being spoken. Output system.
2. The output system according to claim 1, wherein the output unit outputs the stimulus when the acquisition unit is acquiring sound.
3. The interval is based on the average length of the intervals between multiple beats included in the utterance. The output system according to claim 1.
4. The aforementioned interval is a length obtained by multiplying the length of the interval between beats. The output system according to claim 1.
5. The output unit outputs the stimulus in accordance with the sound the user is listening to. The output system according to claim 1.
6. The interval is extracted based on either the accent of the utterance or the length of the mora contained in the utterance. The output system according to claim 1.
7. The output unit outputs the stimulus based on the length of the interval between beats included in the speech within two intervals of the sound that are below a predetermined volume. The output system according to claim 1.
8. The output unit outputs the stimulus based on the length of the interval between beats in the speech output from a single sound source, which is included in the sound. The output system according to claim 1.
9. An acquisition unit that acquires sounds including speech, It comprises an output unit that outputs visual, auditory, or tactile stimuli, The output unit is an output device that outputs the stimulus based on the length of the interval between beats included in the utterance when the utterance is being spoken.
10. The acquisition unit acquires sounds including speech, An output method comprising an output unit outputting visual, auditory, or tactile stimuli, The output unit outputs the stimulus based on the length of the interval between beats included in the utterance when the utterance is being spoken. Output method.
Citation Information
Patent Citations
User interface for ANR headphones with active hear-through
JP2015537466A