Acoustic processing device, acoustic processing system, acoustic processing method, and program

The acoustic processing device and system enhance hearing aids by enabling users to confirm and re-listen to missed speech content through voice recognition and text display, addressing the limitations of existing technologies.

JP2025094324APending Publication Date: 2025-06-25CASIO COMPUTER CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023209768
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-06-25

AI Technical Summary

Technical Problem

Existing hearing aids, such as those described in Patent Document 1, primarily focus on reducing noise and amplifying speech, but do not effectively allow users to confirm the speech content of a speaker when they miss important information.

Method used

An acoustic processing device and system that includes a wearable ear device and a mobile terminal, allowing users to confirm missed speech content through voice recognition and text display, with features like voice re-output and parameter adjustment.

Benefits of technology

Enables users to accurately confirm and re-listen to missed speech content in both voice and text formats, enhancing the usability of hearing aids by providing clear confirmation of speech content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025094324000001_ABST
    Figure 2025094324000001_ABST
Patent Text Reader

Abstract

To check the speech content of a speaker properly.SOLUTION: The acoustic processing device 100 comprises a controller that stores voice obtained from outside sources in real-time as voice data, automatically determines a voice block contained in the stored voice data based on user instructions, and outputs the data according to the determined voice block.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an acoustic processing device, an acoustic processing system, an acoustic processing method, and a program.

Background Art

[0002] Hearing aids that are attached to the user's ear to make it easier to hear ambient sounds are known. In order to make the sound collected by the microphone in the hearing aid into a sound that is easier for the user to hear, processing such as adjusting the volume for each frequency is performed. Further, Patent Document 1 discloses a hearing aid or the like that can reduce the influence of ambient noise in order to make the sound even easier to hear.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The hearing aid disclosed in Patent Document 1 outputs clear voice with less noise by performing speech recognition on all the speech contents of the speaker, converting them into text data, and performing speech synthesis on this text data.

[0005] An object of the present invention is to provide an acoustic processing device, an acoustic processing system, an acoustic processing method, and a program that can satisfactorily confirm the speech content of a speaker.

Means for Solving the Problems

[0006] To achieve the above object, one aspect of the acoustic processing device according to the present invention is to store in real time the voice acquired from the outside as voice data, Automatically determine the voice blocks included in the stored voice data based on an instruction from a user, Output data corresponding to the determined voice blocks, and includes a control unit.

Advantages of the Invention

[0007] According to the present invention, the speech content of a speaker can be confirmed well.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Embodiments for Carrying Out the Invention

[0009] A description will be given with reference to the drawings. In the drawings, the same or corresponding parts are denoted by the same reference numerals.

[0010] As shown in FIG. 1, the acoustic processing system 1000 includes an acoustic processing device 100 as an ear device worn by a user on the ear, and a mobile terminal 200 used by the user while carried. These operate by being communicatively connected according to a short-range wireless communication standard such as Bluetooth (registered trademark).

[0011] The acoustic processing device 100 has a function of adjusting the voice being spoken around to a voice that is easy for the user to hear and outputting it, and a function of being able to replay the utterance content that could not be heard. Also, the mobile terminal 200 has a function of, for example, displaying on the display unit 230 the text obtained by voice recognition of the utterance content that the user could not hear.

[0012] More specifically, in the normal mode, like a general hearing aid, the acoustic processing device 100 acquires the voice of the conversation partner with the microphone 131, adjusts it to a voice that is easy for the user to hear, and outputs it from the speaker 141. Then, when the user cannot hear the output voice and touches the touch sensor 161, the acoustic processing device 100 shifts to the confirmation mode so that the user can confirm the voice that could not be heard.

[0013] In the confirmation mode, when the acoustic processing device 100 cannot perform data communication with the mobile terminal 200, the acoustic processing device 100 outputs the voice (the utterance content that could not be heard) from the speaker 141 again. Also, in the confirmation mode, when the acoustic processing device 100 can perform data communication with the mobile terminal 200, such as when the user is carrying the mobile terminal 200, the acoustic processing device 100 transmits the data (voice data) of the voice (the utterance content that could not be heard) to the mobile terminal 200, and the mobile terminal 200 displays on the display unit 230 the text obtained by voice recognition of the received voice data. Therefore, in any case, the user can confirm the utterance content that could not be heard in voice or text.

[0014] For example, assume that the user carrying the mobile terminal 200 fails to hear the voice of the interlocutor who says "Hello", and touches the touch sensor 161 of the acoustic processing device 100. Then, the mobile terminal 200 performs voice recognition on the voice data and displays "Hello" on the voice recognition text screen 301 as shown in FIG. 1. If the user can confirm the utterance content of "Hello" that could not be heard on the voice recognition text screen 301, the user touches the OK button 302. Thereafter, the acoustic processing device 100 also returns to the normal mode of adjusting the voice of the interlocutor to a voice that is easy for the user to hear and outputting it from the speaker 141, similar to a general hearing aid.

[0015] In the example of FIG. 1, the result correctly recognized as "Hello" is displayed on the voice recognition text screen 301. However, in some cases, it is conceivable that an unclear text string is displayed. In that case, when the user touches the voice re-recognition button 303, the mobile terminal 200 performs voice recognition again and redisplays the voice recognition result on the voice recognition text screen 301. When performing voice recognition again, the mobile terminal 200 may change the voice recognition algorithm or perform some processing (such as consonant emphasis or noise removal) on the voice data before performing voice recognition.

[0016] Also, by turning on the voice output mode switch 304, the user can confirm the utterance content that could not be heard not only as a text string but also as voice from the speaker 141 of the acoustic processing device 100. The playback speed of the voice at that time can be adjusted by the playback speed adjustment bar 305. Also, the user can confirm the utterance content that could not be heard any number of times as voice by touching the voice re-output button 306.

[0017] Also, regarding the sound output by the audio processing device 100, it can be set with the setting button 307. When the user touches the setting button 307, a parameter setting screen (not shown) is displayed on the display unit 230 of the mobile terminal 200, and the setting values (acoustic characteristics) of various parameters (values such as the amplitude increase for each frequency used when adjusting the sound output by the audio processing device 100, volume, and default playback speed during voice reconfirmation) can be adjusted (reset).

[0018] For example, it is possible to set the amplitude increase for each specific frequency band (for example, frequency bands centered around 250 Hz, 500 Hz, 1 kHz, 2 kHz, 4 kHz) in the audible band, set the overall volume, and set the default playback speed when confirming the voice with the voice re - output button 306, etc.

[0019] The above is a description of the outline of the functions of the audio processing system 1000. Next, the functional configurations of the audio processing device 100 and the mobile terminal 200 included in the audio processing system 1000 will be described.

[0020] The audio processing device 100 is a small wearable device. Specifically, as shown in FIG. 1, it is an earphone - type ear device worn on the user's ear. As shown in FIG. 2, the audio processing device 100 includes a control unit 110, a storage unit 120, a sound input unit 130, a sound output unit 140, a communication unit 150, and an operation input unit 160 as its functional components.

[0021] The control unit 110 is composed of at least one processor such as a CPU (Central Processing Unit) or a DSP (Digital Signal Processor). By executing the program stored in the storage unit 120, the control unit 110 performs various processes for operating the acoustic processing device 100. For example, the control unit 110 adjusts the voice data acquired by the voice input unit 130 to a voice that is easy for the user to hear by removing noise or performing processes such as amplification for each frequency band, and then outputs (plays back) it from the voice output unit 140. In addition, the control unit 110 includes a clock unit inside that can acquire time information. Further, the control unit 110 may include the storage unit 120 inside.

[0022] The storage unit 120 stores the program executed by the control unit 110 and necessary data. The storage unit 120 may include a RAM (Random Access Memory), a ROM (Read Only Memory), a flash memory, etc., but is not limited thereto. Note that each parameter set on the parameter setting screen described above is received by the control unit 110 from the mobile terminal 200 via the communication unit 150 and stored as a set value in the storage unit 120. The control unit 110 performs adjustment (processing such as amplification) of the sound output from the speaker 141 based on this set value.

[0023] In addition, the storage unit 120 includes a ring buffer 121 as shown in FIG. 3, and voice data is stored in this ring buffer 121 in real time. The size of the ring buffer is arbitrary, but in this embodiment, the ring buffer 121 is a 1MB memory to which addresses from the start address (0x00000) to the end address (0xFFFFF) are allocated as shown in FIG. 3, and it is assumed that voice data for about 3 minutes (voice information storage period) can be stored. Then, the ring buffer 121 is treated as a buffer connected in a ring shape such that the address next to the end address (0xFFFFF) becomes the start address (0x00000).

[0024] Further, the memory unit 120 includes a head pointer and a tail pointer as pointers for indicating the addresses where valid data is stored in the ring buffer 121. When the ring buffer 121 is initialized, the head address is set for both the head pointer and the tail pointer. At the time of data writing, the control unit 110 writes data from the address of the tail pointer and increments the value of the tail pointer one by one. Also, at the time of data reading, the control unit 110 reads data from the address of the head pointer, and when it is no longer necessary to read that data again (use it), it increments the value of the head pointer. That is, the area from the address of the head pointer to the address of tail pointer - 1 (the area hatched with diagonal lines in FIG. 3) becomes the area where valid data is stored.

[0025] The sound input unit 130 includes a microphone 131, an A / D (Analog / Digital) converter, etc., and acquires ambient sound. Then, the control unit 110 records the audio data obtained by A / D converting the sound acquired by the sound input unit 130 with the microphone 131 at a predetermined sampling frequency (for example, 44.1 kHz) in the ring buffer 121 in the memory unit 120.

[0026] The sound output unit 140 includes a D / A (Digital / Analog) converter, a speaker 141, etc., and outputs the audio data adjusted by the control unit 110 from the speaker 141. More specifically, the control unit 110 amplifies the digital data of the sound acquired by the sound input unit 130 (data obtained by acquiring and A / D converting the sound around the user) based on the set values set for each user (various parameters related to hearing such as the amplification degree of each frequency band), and outputs it as sound from the speaker 141 of the sound output unit 140.

[0027] The communication unit 150 is a communication interface for transmitting and receiving data between the acoustic processing device 100 and the mobile terminal 200 via Bluetooth (registered trademark). By communicating with the mobile terminal 200 through the communication unit 150, the acoustic processing device 100 can, for example, transmit voice data to the mobile terminal 200 or receive information regarding setting values such as the amplitude increase of each frequency band from the mobile terminal 200. Note that the communication standard supported by the communication unit 150 is not limited to Bluetooth (registered trademark), and a communication interface corresponding to wireless LAN (Local Area Network) or the like may also be used.

[0028] The operation input unit 160 includes a touch sensor 161 and receives operation inputs from the user. Note that the sensor (switch) included in the operation input unit 160 is not limited to the touch sensor 161, and any sensor that can receive operation inputs from the user (for example, a push button switch or the like) can be used.

[0029] The mobile terminal 200 is, for example, a smartphone and includes, as functional components, a control unit 210, a storage unit 220, a display unit 230, an operation input unit 240, and a communication unit 250 as shown in FIG. 4.

[0030] The control unit 210 is composed of a processor such as a CPU, for example. The control unit 210 executes terminal voice processing and the like, which will be described later, according to the programs stored in the storage unit 220. Note that the control unit 210 also includes a clock unit inside that can acquire time information, similar to the control unit 110 of the acoustic processing device 100. The control unit 210 and the control unit 110 synchronize the time counted by the clock unit periodically via the communication unit 250 and the communication unit 150. Also, the control unit 210 may include the storage unit 220 inside.

[0031] The storage unit 220 stores programs executed by the control unit 210 and necessary data. The storage unit 220 may include RAM, ROM, flash memory, etc., but is not limited thereto.

[0032] The display unit 230 includes a display device such as a liquid crystal display or an organic EL (Electro-Luminescence) display.

[0033] The operation input unit 240 is a user interface such as a push button switch or a touch panel integrated with the display unit 230, and receives operation inputs from the user. The control unit 210 can acquire what kind of operation input the user has made based on detection results such as a tap operation on the touch panel of the operation input unit 240 or the pressed state of a switch.

[0034] The communication unit 250 is a communication interface for the mobile terminal 200 to perform data communication with the audio processing device 100, an external device (e.g., another smartphone, tablet, PC (Personal Computer), etc.), or to acquire information from the Internet. The communication unit 250 may include a wireless communication interface for communicating, for example, via Bluetooth (registered trademark) or wireless LAN.

[0035] When distinguishing which device's control units 110 and 210, the control unit 110 is called the audio processing control unit or simply the control unit, and the control unit 210 is called the terminal control unit. Also, when distinguishing which device's communication units 150 and 250, the communication unit 150 is called the audio processing communication unit or simply the communication unit, and the communication unit 250 is called the terminal communication unit. The same applies to the storage units 120 and 220 and the operation input units 160 and 240.

[0036] Next, the speech content confirmation process, which is a process for the acoustic processing device 100 to realize the functions as described above, will be described with reference to FIG. 5. Note that the acoustic processing device 100 and the mobile terminal 200 are paired in advance. When the mobile terminal 200 exists near the acoustic processing device 100 (for example, when the user wears the acoustic processing device 100 on the ear and carries the mobile terminal 200), it is assumed that the acoustic processing device 100 and the mobile terminal 200 are automatically communicatively connected. Then, when the power of the acoustic processing device 100 is turned on, after the control unit 110 performs necessary initialization processing etc. (when the mobile terminal 200 exists in the vicinity, the communication connection with the mobile terminal 200 is also included in this initialization processing), the control unit 110 starts executing the speech content confirmation process.

[0037] First, the control unit 110 initializes the recording flag and the instruction flag stored in the storage unit 120 to OFF, and also initializes the head pointer and the tail pointer of the ring buffer 121 (step S101). Then, the control unit 110 acquires voice data with the microphone 131 of the sound input unit 130 (step S102).

[0038] Next, the control unit 110 removes noise from the acquired voice data, or adjusts (processes) the voice data based on each of the above parameters (first setting) (step S103). This adjustment of the voice data is the same as the adjustment of a general hearing aid, and is an adjustment such as increasing the amplification degree in the high frequency band according to the user's auditory characteristics (which can be grasped by, for example, an audiogram). Then, the adjusted voice data is output from the speaker 141 of the sound output unit 140 (step S104).

[0039] Then, the control unit 110 determines whether the recording flag stored in the storage unit 120 is ON (step S105). If the recording flag is ON (step S105; Yes), the process proceeds to step S109.

[0040] On the other hand, if the recording flag is not ON (step S105; No), the control unit 110 determines whether the start point of speech has been detected in the voice data (step S106).

[0041] The starting point of speech is the point in time when the state changes from a state where no voice is heard to a state where some voice is heard, and it can be automatically detected by the control unit 110 analyzing the audio data acquired up to that point. For example, a predetermined volume threshold is set in advance for the volume value of the audio data, the audio data is analyzed, and after a state where the volume is less than the volume threshold continues for a predetermined silent time threshold (e.g., 0.5 seconds), if a state where the volume is greater than or equal to the volume threshold continues for a predetermined speech time threshold (e.g., 0.1 seconds), the control unit 110 detects the moment when the volume becomes greater than or equal to the volume threshold as the starting point of speech.

[0042] If the control unit 110 does not detect the starting point of speech (step S106; No), it returns to step S102. On the other hand, if the starting point of speech is detected (step S106; Yes), the control unit 110 turns on the recording flag stored in the storage unit 120 (step S107). Then, the control unit 110 records the time (date and time) of the starting point and the address in the ring buffer 121 (the value of the end pointer at that time) in the storage unit 120 as index data 122 as shown in FIG. 6 (step S108). Then, it proceeds to step S109.

[0043] In step S109, the control unit 110 records the audio data in the ring buffer 121 of the storage unit 120. The value of the end pointer of the ring buffer 121 is incremented by the amount of audio data recorded. Then, the control unit 110 determines whether there is an instruction from the user (e.g., whether the touch sensor 161 has detected a touch) (step S110). If there is no instruction from the user (step S110; No), it proceeds to step S113.

[0044] On the one hand, if there is an instruction from the user (step S110; Yes), the control unit 110 turns on the instruction flag stored in the storage unit 120 (step S111). Then, the control unit 110 stores in the storage unit 120, taking the time point when the user's instruction was given as the instruction point, that time (date and time), and the address in the ring buffer 121 (the value of the end pointer at that time) as index data 122 as shown in FIG. 6 (step S112). Then, it proceeds to step S113.

[0045] In step S113, the control unit 110 determines whether or not the end point of speech has been detected in the audio data (step S113).

[0046] The end point of speech is the time point when the state changes from a state where some voices can be heard to a state where no voices can be heard, and the control unit 110 can detect it by analyzing the audio data acquired until then. For example, after a state where the volume is equal to or greater than the volume threshold continues for a predetermined speech time threshold (e.g., 0.1 second), and then a state where the volume is less than the volume threshold continues for a predetermined silent time threshold (e.g., 0.5 second), the control unit 110 automatically detects the moment when the volume becomes less than the volume threshold as the end point of speech.

[0047] If the control unit 110 does not detect the end point of speech (step S113; No), it returns to step S102. On the other hand, if the end point of speech is detected (step S113; Yes), the control unit 110 stores in the storage unit 120 the time (date and time) of the end point and the address in the ring buffer 121 (the value of the end pointer at that time) as index data 122 as shown in FIG. 6 (step S114). Then, the control unit 110 turns off the recording flag stored in the storage unit 120 (step S115).

[0048] Then, the control unit 110 determines whether the instruction flag stored in the storage unit 120 is ON (step S116). If the instruction flag is not ON (step S116; No), the control unit 110 updates the head pointer of the ring buffer 121 (step S119). Specifically, since there will be no opportunity to use the voice block recorded this time (the fact that there is no user instruction means that the user has heard all the content spoken by the conversation partner, so there is no need to confirm the spoken content in the future), the head pointer of the ring buffer 121 is incremented by the amount of the voice block recorded this time. Then, it returns to step S102.

[0049] On the other hand, if the instruction flag is ON (step S116; Yes), the control unit 110 executes a voice block output process described later (step S117). Then, the control unit 110 turns off the instruction flag stored in the storage unit 120 (step S118) and returns to step S102.

[0050] Next, the voice block output process executed in step S117 will be described with reference to FIG. 7.

[0051] First, the control unit 110 refers to the ring buffer 121 and the index data 122, and acquires the voice data from the immediately previous start point to the end point as a voice block (step S201). That is, in step S201, the control unit 110 automatically determines, based on the detected start point, the voice block before (immediately before) the timing instructed by the user from the stored voice data, and acquires the determined voice block. Next, the control unit 110 determines whether it is possible to perform data communication with the mobile terminal 200 (step S202). Note that the immediately previous voice block is a voice block based on the start point closest to the timing instructed by the user among a plurality of start points existing before the timing instructed by the user.

[0052] If the mobile terminal 200 is not capable of data communication (step S202; No), the control unit 110 performs a predetermined process on the voice block to generate output voice data (step S203). Note that, as the predetermined process (first process) on the voice block performed here, consonant enhancement, noise removal, etc. (based on a second setting different from the above-described first setting) can be considered, but doing nothing (using the voice block as it is as the output voice data) is also included in the predetermined process. Then, the control unit 110 outputs the generated output voice data from the speaker 141 of the sound output unit 140 (step S204).

[0053] In the second setting, changing (shifting / processing) the frequency band when outputting the voice block to a frequency band that is easier for the user to hear than the first setting, changing the playback speed of outputting the voice block to be slower than the first setting, emphasizing the pronunciation of the attack part of the voice block more than the first setting, etc. can be considered. Note that, in step S204 (the same applies to steps S210, S226, and S246 described later), the voice data input from the microphone 131 while the output voice data is being output from the speaker 141 is controlled by the control unit 110 so as not to be output from the speaker 141 at the same time as the output of the output voice data.

[0054] Then, the control unit 110 determines whether there is an instruction for re-output from the user (step S205). For example, if a touch on the touch sensor 161 is detected, the control unit 110 determines that the user has instructed re-output.

[0055] If there is an instruction for re-output from the user (step S205; Yes), the process returns to step S203 to output the voice block again. At that time, the control unit 110 may output exactly the same output voice data as the previous time, or may perform a predetermined process different from the previous time (second process of emphasizing consonants, reducing noise, or slowing down the playback speed) and then output it.

[0056] If there is no instruction for re-output from the user (step S205; No), the control unit 110 updates the head pointer of the ring buffer 121 (step S206). Specifically, since the voice block output this time is no longer necessary for future output, the head pointer of the ring buffer 121 is incremented by the amount of the voice block output this time. Then, the control unit 110 ends the voice block output process.

[0057] On the other hand, in step S202, if it is possible to perform data communication with the mobile terminal 200 (step S202; Yes), the control unit 110 transmits the voice block to the mobile terminal 200 via the communication unit 150 (step S207).

[0058] Then, the control unit 110 determines whether an instruction regarding voice output (e.g., voice output, voice re-output, etc.) has been received from the mobile terminal 200 (step S208). If no instruction regarding voice output has been received (step S208; No), the process proceeds to step S211.

[0059] If an instruction regarding voice output has been received from the mobile terminal 200 (step S208; Yes), the control unit 110 performs a predetermined process (third process) on the voice block according to the received instruction to generate output voice data (step S209). Then, the control unit 110 outputs the generated output voice data from the speaker 141 of the voice output unit 140 (step S210). The third process may be the same as the above-described first process or second process.

[0060] Then, the control unit 110 determines whether it has received data from the mobile terminal 200 indicating that the user has confirmed the speech content of the speech block (step S211). If the control unit 110 has not received the data (step S211; No), it returns to step S208. If it has received the data (step S211; Yes), it proceeds to step S206. In step S206, the head pointer of the ring buffer 121 is incremented by the number of speech blocks transmitted to the mobile terminal 200 in step S207. Then, the control unit 110 ends the speech block output process.

[0061] The processing on the side of the acoustic processing device 100 has been described above. Next, the terminal voice processing, which is the processing on the side of the mobile terminal 200 that performs text display through data communication with the acoustic processing device 100, will be described with reference to FIG. 8. When the distance between the mobile terminal 200 and the acoustic processing device 100 becomes short (for example, when a user wearing the acoustic processing device 100 on the ear carries the mobile terminal 200), the control unit 210 automatically establishes a communication connection with the acoustic processing device 100 and starts executing the terminal voice processing. However, the terminal voice processing is not limited to being automatically executed in this way, and it may be executed based on a user's instruction. The distance between the mobile terminal 200 and the acoustic processing device 100 is determined based on the signal strength of Bluetooth (registered trademark).

[0062] First, the control unit 210 determines whether a speech block has started being transmitted from the ear device, which is the acoustic processing device 100, via the communication unit 250 (step S301). If transmission has not started (step S301; No), it returns to step S301.

[0063] If the voice block is started to be transmitted (step S301; Yes), the control unit 210 acquires the voice block via the communication unit 250 (step S302). Then, the control unit 210 performs voice recognition processing on the acquired voice block to generate display data (step S303). Note that the voice recognition processing performed on the voice block here is also included in the predetermined processing for the voice block. Then, the control unit 210 outputs the text (display data), which is the result of the voice recognition processing, (displays the text on the display unit 230) (step S304).

[0064] Then, the control unit 210 determines whether the voice output mode is ON (step S305). Regarding the voice output mode, it can be switched by the voice output mode changeover switch 304 in FIG. 1. In this step S305, the control unit 210 determines whether the voice output mode changeover switch 304 is ON or OFF.

[0065] If the voice output mode is OFF (step S305; No), the process proceeds to step S307. If the voice output mode is ON (step S305; Yes), the control unit 210 transmits an instruction regarding voice output to the ear device, which is the acoustic processing device 100 (step S306). Specifically, the control unit 210 transmits an instruction to reproduce the voice block to the ear device together with information on each currently set parameter (volume, reproduction speed, amplitude increase degree of each frequency band, etc.). As a result, the uttered content of the voice block is output by the ear device, which is the acoustic processing device 100, by the processing of steps S208 to S210 of the above-described voice block output processing (FIG. 7).

[0066] Then, the control unit 210 determines whether there is a user operation (step S307). If there is no user operation (step S307; No), the process returns to step S307 and waits until there is some user operation.

[0067] If there is a user operation (step S307; Yes), the control unit 210 determines whether the user operation is a touch on the OK button 302 (step S308). If the user operation is a touch on the OK button 302 (step S308; Yes), the control unit 210 transmits data indicating that the user has confirmed it via the communication unit 250 (step S309), and returns to step S301.

[0068] If the user operation is other than a touch on the OK button 302 (step S308; No), the control unit 210 performs processing according to the user operation (processing based on operations of various UIs (User Interfaces) shown in FIG. 1 etc., such as voice recognition, voice output mode switching, playback speed adjustment, voice re-output, parameter setting, etc.) (step S310), and returns to step S307.

[0069] By the above-described speech content confirmation process (FIG. 5) being executed by the control unit 110 of the acoustic processing apparatus 100 and the terminal voice processing (FIG. 8) being executed by the control unit 210 of the mobile terminal 200, the acoustic processing system 1000 can normally provide functions such as a hearing aid and can provide a function of allowing the user to confirm the speech content that the user could not hear.

[0070] Also, the control unit 110 of the acoustic processing apparatus 100 acquires a voice block, which is a unit of speech that is somewhat coherent, from the voice data acquired from the outside (of the conversation partner), and based on the user's instruction, can confirm the part that the user could not hear in units of voice blocks. Therefore, even if the user touches the touch sensor 161 at an arbitrary timing, the speech content of the conversation partner will not be interrupted during confirmation. For this reason, the user can grasp the speech content of the conversation partner in units of voice blocks and then confirm the part that could not be heard.

[0071] In the above-described voice block output process (FIG. 7) and terminal voice process (FIG. 8), when the voice output mode switching switch 304 is ON, the control unit 210 of the mobile terminal 200 transmits an instruction regarding voice output to the acoustic processing device 100 (step S306), and the control unit 110 of the acoustic processing device 100 performs a predetermined process on the voice block according to the instruction (step S209). However, this is only an example of the process when the acoustic processing device 100 and the mobile terminal 200 cooperate to output voice in the confirmation mode.

[0072] For example, instead of transmitting an instruction regarding voice output in step S306, the control unit 210 of the mobile terminal 200 may generate output voice data by performing a predetermined process on the voice block according to the instruction and transmit the generated output voice data to the acoustic processing device 100. In this case, the control unit 110 of the acoustic processing device 100 receives the output voice data from the mobile terminal 200 in step S209 of the voice block output process (FIG. 7), and outputs the received output voice data from the voice output unit 140 in step S210. By doing so, since the predetermined process on the voice block is performed on the mobile terminal 200 side, the power consumption of the ear device, which is the acoustic processing device 100, can be further reduced.

[0073] Also, in the above-described embodiment, the control unit 110 of the acoustic processing device 100 detects the start point and end point of the speech, but these detections do not necessarily have to be performed by the acoustic processing device 100. For example, the control unit 110 may transmit all the acquired voice data to the mobile terminal 200, and the control unit 210 of the mobile terminal 200 may detect the start point and end point from the received voice data. Further, the control unit 210 may store the ring buffer 121 and index data 122 for storing the acquired voice data in the storage unit 220 of the mobile terminal 200. Also, not limited to the above-described content, each process necessary in the acoustic processing system 1000 may be executed by either the control unit 110 or the control unit 210.

[0074] As described above, the acoustic processing system 1000 includes the acoustic processing device 100 and the mobile terminal 200. However, if an acoustic processing device has both the functions of the acoustic processing device 100 and the mobile terminal 200, the acoustic processing device alone can provide the same functions as the acoustic processing system 1000. Such an embodiment will be described.

[0075] As shown in FIG. 9, the acoustic processing system 1001 includes an acoustic processing device 101. In FIG. 9, the acoustic processing device 101 is a smartphone including a microphone 131, a speaker 141, and a display unit 170, but it may be a smartwatch or the like including these components.

[0076] Similar to the acoustic processing device 100 and the mobile terminal 200, the acoustic processing device 101 has functions such as adjusting and outputting the voice being spoken around into a voice that is easy for the user to hear, replaying the speech content that could not be heard, and displaying on the display unit 170 the text obtained by voice recognition of the speech content that the user could not hear.

[0077] More specifically, in the normal mode, the acoustic processing device 101 acquires the voice of the conversation partner with the microphone 131, adjusts it into a voice that is easy for the user to hear, and outputs it from the speaker 141. If the user cannot hear the output voice, when the replay button 308 on the touch panel is touched, the acoustic processing device 101 determines that there is an instruction from the user, and then shifts to the confirmation mode.

[0078] In the confirmation mode, the acoustic processing device 101 displays on the display unit 170 the text obtained by voice recognition of the data (voice data) of the voice (the speech content that could not be heard). Also, when the voice output mode is ON, the voice data is output again from the speaker 141 at a specified playback speed. Therefore, the user can confirm the speech content that could not be heard in the form of voice or text.

[0079] In addition, various UIs such as the OK button 302 and the voice recognition button 303 shown in FIG. 9 have the same functions as those in FIG. 1, and the setting button 307 can also be used to adjust the setting values of various parameters in the same way.

[0080] As shown in FIG. 10, the audio processing device 101 includes a control unit 110, a storage unit 120, an audio input unit 130, an audio output unit 140, a display unit 170, and an operation input unit 180 as its functional components. Among these, since the control unit 110, the storage unit 120, the audio input unit 130, and the audio output unit 140 are the same as those of the above-described audio processing device 100, the description thereof is omitted.

[0081] The display unit 170 includes a display device such as a liquid crystal display or an organic EL (Electro-Luminescence) display.

[0082] The operation input unit 180 is a user interface such as a push button switch or a touch panel integrated with the display unit 170, and receives operation inputs from the user. The control unit 110 can obtain what operation input the user has made based on the detection results such as tap operations on the touch panel of the operation input unit 180 or the pressed state of the switch.

[0083] Since the speech content confirmation process executed by the audio processing device 101 is almost the same as the above-described speech content confirmation process (FIG. 5), the description is omitted except for steps S110 and S117. However, the determination of "Is there a user instruction?" in step S110 of the speech content confirmation process is made by determining whether the re-listening button 308 has been touched. In addition, since the voice block output process executed in step S117 of the speech content confirmation process is different from the above-described embodiment, the voice block output process will be described with reference to FIG. 11.

[0084] First, the control unit 110 refers to the ring buffer 121 and the index data 122, and acquires the audio data from the immediately previous start point to the end point as an audio block (step S221). This process is the same as the process of step S201 in the audio block output process (Figure 7) in the above-described embodiment.

[0085] Then, the control unit 110 performs speech recognition on the acquired audio block (step S222), and text-displays the recognition result on the display unit 170 (step S223).

[0086] Then, the control unit 110 determines whether the voice output mode is ON (step S224). Regarding the voice output mode, it can be switched by the voice output mode changeover switch 304 shown in FIG. 9. In this step, the control unit 110 determines whether the voice output mode changeover switch 304 is ON or OFF. If the voice output mode is OFF (step S224; No), the process proceeds to step S227.

[0087] If the voice output mode is ON (step S224; Yes), the control unit 110 performs a predetermined process on the audio block to generate output audio data (step S225). Then, the control unit 110 outputs the generated output audio data from the speaker 141 of the audio output unit 140 (step S226). The processes of step S225 and step S226 are the same as the processes of step S203 and step S204 in the audio block output process (Figure 7) in the above-described embodiment.

[0088] Then, the control unit 110 determines whether there is a user operation (step S227). If there is no user operation (step S227; No), the process returns to step S227, and the control unit 110 waits until there is some user operation.

[0089] If there is a user operation (step S227; Yes), the control unit 110 determines whether the user operation is a touch on the OK button 302 (step S228). If the user operation is a touch on the OK button 302 (step S228; Yes), the control unit 110 updates the head pointer of the ring buffer 121 (step S230). Specifically, the head pointer of the ring buffer 121 is incremented by the amount of the voice block recognized this time. Then, the control unit 110 ends the voice block output process.

[0090] On the other hand, if the user operation is other than a touch on the OK button 302 (step S228; No), the control unit 110 performs processing according to the user operation (voice re-recognition, voice output mode switching, playback speed adjustment, voice re-output, parameter setting, etc., processing according to operations on various UIs (User Interfaces) shown in FIG. 9 etc.) (step S229), and returns to step S227.

[0091] By executing the above-described voice block output process (FIG. 11) and the above-described utterance content confirmation process (FIG. 5) by the control unit 110 of the acoustic processing device 101, the control unit 110 automatically determines the voice blocks included in the voice data, and outputs data according to the determined voice blocks. Therefore, the acoustic processing system 1001 can provide a function that allows the user to confirm the utterance content that the user could not hear even without an ear device.

[0092] Also in this embodiment, since the user can confirm the utterance content that the user could not hear in units of voice blocks, even if the user touches the replay button 308 at an arbitrary timing, the utterance content of the conversation partner will not be interrupted in the middle. For this reason, the user can grasp the utterance content of the conversation partner in units of voice block units and then confirm the parts that could not be heard.

[0093] In the above-described embodiment, when the user touches the touch sensor 161 (or the repeat button 308), at step S110 of the speech content confirmation process (Fig. 5), the control unit 110 determines that "there is a user instruction" (judging that the content spoken by the conversation partner at that time could not be heard), and then, when an end point is detected (when the conversation partner's speech is interrupted), the mode is transitioned from the normal mode to the confirmation mode. However, this mode transition is not limited to being caused by the touch sensor 161 (or the repeat button 308). For example, if the control unit 110 recognizes that the user has uttered a specific word (pardun word) such as "excuse me" or "wait a moment" to confirm again what the conversation partner has said, the mode may be transitioned to the confirmation mode.

[0094] At step S110 of the speech content confirmation process (Fig. 5), when the control unit 110 recognizes the voice data recorded so far and determines that the user has uttered a specific word (pardun word) such as "excuse me" or "wait a moment" to confirm again what the conversation partner has said, this specific word is regarded as a user instruction and it is determined that "there is a user instruction".

[0095] In this modification, the user does not need to touch the touch sensor 161 (or the repeat button 308), and by using the pardun word, the user can quickly and easily confirm what the conversation partner has said. Furthermore, when the user utters the pardun word, the conversation partner often interrupts their speech and waits (because most people interrupt their speech and wait when they are said "excuse me", "wait a moment", etc.), so the user can calmly confirm what the conversation partner has said immediately before that.

[0096] In addition, in order to improve the recognition accuracy of the pardon word, the control unit 110 may have the user pronounce the pardon word in advance and learn the characteristics of the user's voice. Through this learning, it becomes possible to accurately recognize that the user has spoken the pardon word. Moreover, even if a conversation partner with a voice quality different from that of the user accidentally speaks the pardon word, the control unit 110 can recognize that it is not the user speaking the pardon word, thus preventing an accidental transition to the confirmation mode.

[0097] Also, in the above-described embodiment, in the voice block output process (FIGS. 7 and 11) and the terminal voice process (FIG. 8), the data (the time of the user instruction and the address of the ring buffer 121 written at the time of the user instruction) whose type recorded in the index data 122 is "instruction point" was not used (conversely, in the above-described embodiment, it was not necessary to record the instruction point data in the index data 122). However, by using this instruction point data, it is possible to more easily confirm the voice that the user could not hear.

[0098] When the control unit 110 acquires a voice block in step S201 of the voice block output process (FIG. 7), it refers to the index data 122 and also acquires information on at which point of the acquired voice block there was a user instruction such as a touch of the touch sensor 161, the time of the user instruction, and the address.

[0099] Then, when the control unit 110 generates the output voice data in step S203, it slows down the playback speed or increases the volume when playing back the content spoken during a predetermined period (for example, 1 second before and after) around the time of the user instruction, making it easier to hear the voice of the voice block during the predetermined period.

[0100] Also, when the control unit 110 transmits the voice block to the mobile terminal 200 in step S207, it also transmits the information on the time and address of the user instruction.

[0101] Then, when the control unit 210 of the mobile terminal 200 acquires the voice block in step S302 of the terminal voice processing (FIG. 8), it also acquires the information on the time and address indicated by the user. Then, when the control unit 210 displays the recognition result in text in step S304, it highlights (for example, makes it bold, underlines it, or displays it in a large font) the content spoken during a predetermined period (for example, 1 second before and after) around the time indicated by the user, thereby making it easier to view the voice recognition result of the voice block during the predetermined period.

[0102] Through these processes, among the speech content of the conversation partner, the part that the user could not hear (the content spoken during a predetermined period before and after the time indicated by the user) is output in an easy-to-check manner (such as slowing down the playback speed, increasing the volume, making it bold, underlining it, or displaying it in a large font), so that the user can easily check the part that could not be heard.

[0103] Also, when the control unit 110 acquires the voice block in step S221 of the voice block output process (FIG. 11), it refers to the index data 122 to acquire information on at which point of the acquired voice block there was a user instruction such as a touch on the replay button 308, and the time and address information indicated by the user.

[0104] Then, when the control unit 110 displays the recognition result in text in step S223, it highlights (for example, makes it bold, underlines it, or displays it in a large font) the part spoken at the time indicated by the user.

[0105] Then, when the control unit 110 generates the output voice data in step S225, it slows down the playback speed or increases the volume when playing back a predetermined period (for example, 1 second before and after) around the time indicated by the user. Through these processes, the user can easily check the part that could not be heard in the speech content of the conversation partner.

[0106] Also, in the above-described embodiment, even if it is determined in step S110 of the utterance content confirmation process (Fig. 5) that there is a user instruction, the mode transition from the normal mode to the confirmation mode is not performed until the end point is detected in step S113. However, the mode transition to the confirmation mode may be immediately performed when it is determined that there is a user instruction.

[0107] In step S113 of the utterance content confirmation process (Fig. 5), the control unit 110 determines "end point detection or instruction flag = ON". As a result, the time point when there is a user instruction is also regarded as the end point, and the confirmation mode can be immediately entered. By doing so, when the user cannot hear the utterance content of the conversation partner, the user can immediately confirm the content.

[0108] Also, in the above-described embodiment, the head pointer of the ring buffer 121 is updated at the end (steps S206 and S230) of the voice block output process (Figs. 7 and 11), but this update does not necessarily have to be performed here.

[0109] As shown in Fig. 12, the acoustic processing system 1002 includes an acoustic processing device 102. The acoustic processing device 102 has the same functional configuration as the acoustic processing device 101. However, as shown in Fig. 12, on the voice recognition text screen 301, voice block display frames 321, 322, and 323 are displayed, and the voice recognition results are displayed surrounded by each voice block display frame 321, 322, and 323 for each voice block.

[0110] On the voice recognition text screen 301, the voice recognition results of a plurality of voice blocks are displayed. One of them (initially the voice recognition result of the latest voice block) is displayed surrounded by a solid-line voice block display frame (voice block display frame 323 in the example shown in Fig. 12) as the selected voice block (selected voice block). The unselected voice blocks are displayed surrounded by a dotted-line voice block display frame (voice block display frames 321 and 322 in the example shown in Fig. 12).

[0111] When the voice recognition button 303, voice re-output button 306, etc. are touched, the voice block to be processed is the selected voice block. The selected voice block can have its selection changed one by one to the previous voice block using the previous button 326 and to the next block using the next button 325. For example, when the previous button 326 is touched in the state shown in FIG. 12, the voice block display frame 323 becomes a dotted line, the voice block display frame 322 becomes a solid line, and the voice block displayed as "It's a nice day" becomes the selected voice block and becomes the target of voice recognition and voice re-output.

[0112] Also, when the block deletion button 324 is touched, regardless of the selected voice block, the oldest voice block stored in the ring buffer 121 is deleted, and the head pointer of the ring buffer 121 is incremented by the amount of that voice block.

[0113] Regarding the speech content confirmation process in the acoustic processing system 1002, since it is the same as the speech content confirmation process in the above-described acoustic processing system 1001, the description is omitted. Regarding the voice block output process in the acoustic processing system 1002, since it is almost the same as the above-described voice block output process (FIG. 11), mainly the different parts will be described with reference to FIG. 13.

[0114] First, the control unit 110 refers to the ring buffer 121 and the index data 122, and acquires the voice data from the immediately previous start point to the end point as the latest voice block (step S241).

[0115] Then, the control unit 110 performs voice recognition on the acquired voice block (step S242), and displays the recognized text together with the voice block display frame on the display unit 170 (step S243).

[0116] The processing from the next step S244 to step S248 is the same as the processing from step S224 to step S228 in the voice block output process (FIG. 11), so the description is omitted.

[0117] Then, in step S248, if the user operation is a touch on the OK button 302 (step S248; Yes), the control unit 110 ends the voice block output process. If the user operation is not a touch on the OK button 302 (step S248; No), the control unit 110 determines whether the user operation is a touch on the forward button 326 or the backward button 325 (step S249).

[0118] If the user operation is a touch on the forward button 326 or the backward button 325 (step S249; Yes), the control unit 110 changes the selected voice block according to the touched button (step S250), and returns to step S247. For example, in the example shown in FIG. 12, the selected voice block is "Where are you going out?", but if the forward button 326 is touched here, the selected voice block is changed to "Nice weather", and the voice block display frame 322 is changed to a solid line and the voice block display frame 323 is changed to a dotted line, respectively.

[0119] If the user operation is neither a touch on the forward button 326 nor a touch on the backward button 325 (step S249; No), the control unit 110 determines whether the user operation is a touch on the block deletion button 324 (step S251).

[0120] If the user operation is a touch on the block deletion button 324 (step S251; Yes), the control unit 110 deletes the display of the oldest voice block from the display unit 170 together with the voice block display frame, and updates the head pointer of the ring buffer 121 (step S252). The update of the head pointer specifically increments the head pointer of the ring buffer 121 by the amount of the oldest voice block. Then it returns to step S247.

[0121] If the user operation is not a touch on the block deletion button 324 (step S251; No), the control unit 110 performs processing according to the user operation (processing according to operations on various UIs (User Interfaces) shown in FIG. 12, etc., such as voice recognition, voice output mode switching, playback speed adjustment, voice re-output, parameter setting, etc.) (step S253), and returns to step S227.

[0122] By the above processing, until receiving an instruction to delete a voice block from the user, the past voice blocks are continuously saved, and the user can select a desired voice block by touching the previous button 326 or the next button 325 and output the voice again. In some cases, there may be speech content that cannot be understood (correctly heard) because the flow of the conversation from the past is not known. However, even in such cases, the voice blocks can be confirmed retroactively, making it easier for the user to confirm the speech content that could not be heard.

[0123] Note that in this example, unless the user touches the block deletion button 324, the head pointer of the ring buffer 121 is not updated, so there is a possibility that the ring buffer 121 will cause a capacity overflow. Therefore, if a situation occurs where the tail pointer of the ring buffer 121 is about to overtake the head pointer, the control unit 110 may automatically perform processing to increment the head pointer by the amount of the oldest voice block. By doing so, even if the user does not manually delete the voice block, the head pointer of the ring buffer 121 can be automatically updated, preventing the ring buffer 121 from causing a capacity overflow.

[0124] In the above-described embodiment, the result of voice recognition of the voice block including the timing of the instruction was displayed only when there was a user instruction such as the user touching the touch sensor 161. However, the result of voice recognition may be displayed for all voice blocks regardless of the presence or absence of a user instruction.

[0125] When the determination in step S116 of the utterance content confirmation process (Figure 5) is No, the control unit 110 of the audio processing device 100 transmits the audio block to the mobile terminal 200 before proceeding to step S119. At that time, information indicating that there was no user instruction is also transmitted. Then, in the mobile audio processing (Figure 8) of the mobile terminal 200, when the information indicating that there was no user instruction is included in the information transmitted from the ear device, after finishing the process of step S304, the control unit 210 is configured to return to step S301.

[0126] By performing the above-described processing, regardless of the presence or absence of a user instruction, for all audio blocks, the result of voice recognition is text-displayed on the display unit 230 of the mobile terminal 200. As a result, the user can not only confirm the content that could not be heard, but also look back and confirm the content of the conversation so far in text display, making it easier to understand the conversation with the conversation partner.

[0127] In the above-described embodiment, the audio from the start point to the end point of the utterance is treated as one audio block, but this is merely an example of an audio block. For example, the control unit 110 may sequentially perform voice recognition on the uttered audio and treat the audio data within a period that can be recognized as one sentence as one audio block. By doing so, the user can confirm the words that could not be heard in units of one sentence.

[0128] Also, although the audio processing device 100 has been described as being wirelessly connected to the mobile terminal 200, the audio processing device 100 and the mobile terminal 200 may be wired-connected.

[0129] Also, in the above-described acoustic processing system 1001, although the acoustic processing device 101 provides the same functions as a hearing aid in the normal mode, it may be configured not to output voice in the normal mode. That is, in the normal mode, the acoustic processing device 101 acquires the voice of the conversation partner with the microphone 131, but does not control to output all the voices of the conversation partner from the speaker 141. When the user cannot hear the voice of the conversation partner, etc., in response to the detection of an operation by the user (such as a touch on the replay button 308 on the touch panel), the acoustic processing device 101 determines the immediately preceding voice block (the voice of the conversation partner) stored in the ring buffer 121, and controls to output the voice corresponding to the determined voice block from the speaker 141 or to display the voice recognition text on the display unit 230. To do this, the process of step S104 of the utterance content confirmation process (Figure 5) may be omitted.

[0130] Note that the acoustic processing devices 100 and 101 can also be realized by a computer equipped with the microphone 131 and the speaker 141. Also, the mobile terminal 200 can be realized not only by a smartphone, but also by a computer such as a smartwatch, a tablet, or a PC that can communicate with the acoustic processing device 100 (ear device).

[0131] Specifically, it has been described that the programs executed by the control units 110 of the acoustic processing devices 100, 101, and 102 are stored in advance in the storage unit 120, and the programs executed by the control unit 210 of the mobile terminal 200 are stored in advance in the storage unit 220. However, the program may be stored and distributed in a computer-readable recording medium such as a flexible disk, a CD-ROM (Compact Disc Read Only Memory), a DVD (Digital Versatile Disc), an MO (Magneto-Optical disc), a memory card, or a USB memory, and a computer capable of executing the above-described respective processes may be configured by reading and installing the program into the computer.

[0132] Furthermore, the program can be superimposed on a carrier wave and applied via a communication medium such as the Internet. For example, the program may be posted and distributed on a bulletin board (BBS: Bulletin Board System) on a communication network. Then, by starting this program and executing it in the same manner as other application programs under the control of an OS (Operating System), each of the above-described processes may be configured to be executable.

[0133] In addition, the control unit 110 and the control unit 210 may be configured not only by any single processor such as a single processor, a multi-processor, or a multi-core processor, but also by a combination of any of these processors and a processing circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field-Programmable Gate Array).

[0134] As described above, the preferred embodiments of the present invention have been described. However, the present invention is not limited to such specific embodiments, and the present invention includes the invention described in the claims and its equivalent scope. The present invention has the advantage that even if there is a word (phrase) that the user could not hear during a conversation with the other party, the user can confirm the content of the other party's speech by himself / herself without asking the other party to speak again.

Explanation of Reference Numerals

[0135] 100, 101, 102... audio processing device, 110, 210... control unit, 120, 220... memory unit, 121... ring buffer, 122... index data, 130... audio input unit, 131... microphone, 140... audio output unit, 141... speaker, 150, 250... communication unit, 160, 180, 240... operation input unit, 161... touch sensor, 170, 230... display unit, 200... mobile terminal, 301... voice recognition text screen, 302... OK button, 303... voice re-recognition button, 304... voice output mode switch, 305... playback speed adjustment bar, 306... voice re-output button, 307... settings button, 308... replay button, 321, 322, 323... voice block display frame, 324... block deletion button, 325... forward button, 326... backward button, 1000, 1001, 1002... audio processing system

Claims

1. Real-time store the voice obtained from the outside as voice data, automatically determine the voice blocks included in the stored voice data based on an instruction from the user, output data corresponding to the determined voice blocks, and include a control unit, an acoustic processing device.

2. The control unit detects the start point of speech from the voice data, and the voice block is determined based on the detected start point, The acoustic processing device according to Claim 1.

3. further include a display unit, and the control unit performs voice recognition on the voice block to generate display data, and displays the generated display data on the display unit, The acoustic processing device according to Claim 1.

4. The control unit when the specific words spoken by the user are included in the voice data, determines the voice blocks included in the voice data using the specific words as the instruction from the user, The acoustic processing device according to Claim 1.

5. The voice data is stored in a ring buffer, The acoustic processing device according to Claim 1.

6. An acoustic processing system including an ear device and a mobile terminal, wherein the ear device real-time stores the voice obtained from the outside as voice data, automatically determines the voice blocks included in the stored voice data based on an instruction from the user, and transmits the determined voice blocks to the mobile terminal, and includes a control unit, wherein the mobile terminal receives the voice blocks transmitted from the ear device, performs voice recognition processing on the received voice blocks, and outputs the text obtained by the voice recognition processing, and includes a terminal control unit, an acoustic processing system.

7. A control unit real-time stores the voice obtained from the outside as voice data, automatically determines the voice blocks included in the stored voice data based on an instruction from the user, and outputs data corresponding to the determined voice blocks, an acoustic processing method.

8. A program that causes a control unit to real-time store the voice obtained from the outside as voice data, automatically determine the voice blocks included in the stored voice data based on an instruction from the user, and output data corresponding to the determined voice blocks. ​

Citation Information

Patent Citations

  • Hearing aid and program

    JP2019213001A