Sound processing method, sound processing device, and sound processing program
The audio processing device addresses inconsistent notification in existing technologies by using a voice recognition system to provide customizable and efficient voice detection notifications, improving user experience through tailored sound, vibration, and visual alerts.
Patent Information
- Application Number
- JP2025006751
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-01-21
- Filing Date
- 2025-01-17
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2043-01-10
AI Technical Summary
Existing audio processing technologies do not adequately account for user preferences regarding notification of detected voices, leading to inconsistent user experiences.
An audio processing device and method that includes a voice recognition system to detect voices satisfying preset conditions and notify users accordingly, using various notification methods based on priority settings, such as sound, vibration, and visual cues, allowing users to customize their notification preferences.
Enhances user convenience by providing tailored notifications based on voice detection, ensuring users are promptly informed of relevant sounds while minimizing distractions.
Smart Images

Figure 0007763367000001 
Figure 0007763367000002 
Figure 0007763367000003
Abstract
Description
Cross-reference to related applications
[0001] This application claims priority to Japanese Patent Application No. 2022-008227, filed on January 21, 2022, the entire disclosure of which is incorporated herein by reference. [Technical Field]
[0002] The present disclosure relates to an audio processing device, an audio processing method, and an audio processing system. [Background technology]
[0003] Conventionally, there is known a technology that enables a user to hear surrounding sounds while wearing an audio output device such as headphones or earphones. In such a technology, a portable music playback device is known that is equipped with a notification means that notifies the user through headphones when an external sound matches a predetermined phrase (Patent Document 1). [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2001-256771 Summary of the Invention
[0005] An audio processing device according to an embodiment of the present disclosure includes: The control unit acquires the results of a voice recognition process that recognizes voice from voice data, and when voice that satisfies predetermined setting conditions for the voice is detected based on the results of the voice recognition process, notifies the user that voice that satisfies the setting conditions has been detected, in accordance with notification conditions corresponding to the detected voice that are set in the setting conditions.
[0006] An audio processing method according to an embodiment of the present disclosure includes: Obtaining a result of a speech recognition process for recognizing speech from the speech data; When a voice that satisfies a preset setting condition regarding the voice is detected based on the results of the voice recognition processing, the method includes notifying the user that a voice that satisfies the set condition has been detected, in accordance with a notification condition corresponding to the detected voice that is set in the set condition.
[0007] An audio processing system according to an embodiment of the present disclosure includes: A sound collector that collects surrounding sounds, and a voice processing device that acquires the results of a voice recognition process that recognizes voice from the voice data collected by the sound collector, and when voice that satisfies predetermined setting conditions for the voice is detected based on the results of the voice recognition process, notifies the user that voice that satisfies the setting conditions has been detected, in accordance with notification conditions corresponding to the detected voice that are set in the setting conditions. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a diagram illustrating a schematic configuration of a voice processing system according to an embodiment of the present disclosure. [Figure 2] FIG. 2 is a block diagram of the voice processing system shown in FIG. [Figure 3] FIG. 10 is a diagram illustrating an example of a search list. [Figure 4] FIG. 10 is a diagram illustrating an example of a notification sound list. [Figure 5] FIG. 10 is a diagram for explaining notification means and notification timing according to priority. [Figure 6] FIG. 10 is a diagram illustrating an example of a main screen. [Figure 7] FIG. 10 is a diagram illustrating an example of a setting screen. [Figure 8] FIG. 10 is a diagram illustrating an example of a notification screen. [Figure 9] 3 is a diagram illustrating an example of processing performed by a section detection unit shown in FIG. 2. FIG. [Figure 10] FIG. 3 is a block diagram of an utterance storage unit shown in FIG. 2. [Figure 11] 3 is a flowchart showing the operation of an event detection process executed by the voice processing device shown in FIG. 2. [Figure 12] 3 is a flowchart showing an operation of a playback data output process executed by the audio processing device shown in FIG. 2; [Figure 13] 3 is a flowchart showing an operation of a playback data output process executed by the audio processing device shown in FIG. 2; [Figure 14] FIG. 10 is a diagram illustrating a schematic configuration of a voice processing system according to another embodiment of the present disclosure. [Figure 15] FIG. 10 is a diagram illustrating a schematic configuration of a voice processing system according to yet another embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0009] The conventional techniques have room for improvement. For example, depending on the content of the detected voice, a user may preferentially want to be notified of the detection of voice in some cases, or may preferentially not want to be notified in other cases. According to an embodiment of the present disclosure, an improved voice processing device, a voice processing method, and a voice processing system can be provided.
[0010] In the present disclosure, "audio" includes any sound, including, for example, human voices, machine-generated sounds, animal cries, and environmental sounds.
[0011] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings.
[0012] 1, the sound processing system 1 includes a sound collector 10 and a sound processing device 20. The sound collector 10 and the sound processing device 20 can communicate with each other via a communication line. The communication line includes at least one of a wired line and a wireless line.
[0013] In this embodiment, the sound collector 10 is an earphone. However, the sound collector 10 is not limited to an earphone. The sound collector 10 may be a headphone or the like. The sound collector 10 is worn by the user. The sound collector 10 is capable of outputting music or the like. The sound collector 10 may include an earphone unit worn on the user's left ear and an earphone unit worn on the user's right ear.
[0014] The sound collector 10 collects sounds around the sound collector 10. When worn by a user, the sound collector 10 collects sounds around the user. The sound collector 10 outputs the collected sounds around the user based on the control of the sound processing device 20. With this configuration, the user can hear sounds around them while wearing the sound collector 10.
[0015] In this embodiment, the voice processing device 20 is a terminal device. The terminal device that serves as the voice processing device 20 is, for example, a mobile phone, a smartphone, a tablet, or a personal computer (PC), etc. However, the voice processing device 20 is not limited to a terminal device.
[0016] The sound processing device 20 is operated by a user. The user can operate the sound processing device 20 to set the sound collector 10 and the like.
[0017] The sound processing device 20 controls the sound collector 10 to collect sounds around the user. When the sound processing device 20 detects sounds that satisfy preset conditions from the collected sounds around the user, it notifies the user that sounds that satisfy the set conditions have been detected. Details of this process will be described later.
[0018] Fig. 2 is a block diagram of the voice processing system 1 shown in Fig. 1. In Fig. 2, the main flow of data and the like is indicated by solid lines.
[0019] The sound collector 10 includes a microphone 11, a speaker 12, a communication unit 13, a storage unit 14, and a control unit 15.
[0020] The microphone 11 is capable of collecting sounds around the sound collector 10. The microphone 11 includes a left microphone and a right microphone. The left microphone may be included in an earphone unit included in the sound collector 10 that is worn on the left ear of the user. The right microphone may be included in an earphone unit included in the sound collector 10 that is worn on the right side of the user. For example, the microphone 11 is a stereo microphone or the like.
[0021] The speaker 12 is capable of outputting sound. The speaker 12 includes a left speaker and a right speaker. The left speaker may be included in an earphone unit included in the sound collector 10 that is worn on the left ear of the user. The right speaker may be included in an earphone unit included in the sound collector 10 that is worn on the right side of the user. For example, the speaker 12 may be a stereo speaker.
[0022] The communication unit 13 includes at least one communication module capable of communicating with the audio processing device 20 via a communication line. The communication module is a communication module compatible with the standard of the communication line. The standard of the communication line is, for example, a wired communication standard or a short-range wireless communication standard including Bluetooth (registered trademark), infrared, and NFC (Near Field Communication).
[0023] The storage unit 14 is configured to include at least one semiconductor memory, at least one magnetic memory, at least one optical memory, or a combination of at least two of these. The semiconductor memory is, for example, a RAM (Random Access Memory) or a ROM (Read Only Memory). The RAM is, for example, an SRAM (Static Random Access Memory) or a DRAM (Dynamic Random Access Memory). The ROM is, for example, an EEPROM (Electrically Erasable Programmable Read Only Memory). The storage unit 14 may function as a main storage device, an auxiliary storage device, or a cache memory. The storage unit 14 stores data used in the operation of the sound collector 10 and data obtained by the operation of the sound collector 10. For example, the storage unit 14 stores system programs, application programs, embedded software, etc.
[0024] The control unit 15 is configured to include at least one processor, at least one dedicated circuit, or a combination of these. The processor is a general-purpose processor such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit), or a dedicated processor specialized for specific processing. The dedicated circuit is, for example, an FPGA (Field-Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). The control unit 15 executes processing related to the operation of the sound collector 10 while controlling each part of the sound collector 10.
[0025] In this embodiment, the control unit 15 includes an audio acquisition unit 16, an audio playback unit 17, and a storage unit 18. The storage unit 18 is configured to include components that are the same as or similar to those of the storage unit 14. At least a part of the storage unit 18 may be part of the storage unit 14. The operation of the storage unit 18 is executed by a processor or the like of the control unit 15.
[0026] The voice acquisition unit 16 acquires digital voice data from the analog voice data collected by the microphone 11. In this embodiment, the voice acquisition unit 16 acquires voice sampling data as digital voice data by sampling the analog voice data at a preset sampling rate.
[0027] The voice acquisition unit 16 outputs the voice sampling data to the voice reproduction unit 17. The voice acquisition unit 16 also transmits the voice sampling data to the voice processing device 20 via the communication unit 13.
[0028] When the microphone 11 includes a left microphone and a right microphone, the audio acquisition unit 16 may acquire left audio sampling data from analog data of audio collected by the left microphone. The audio acquisition unit 16 may also acquire right audio sampling data from analog data of audio collected by the right microphone. The audio acquisition unit 16 may transmit the left audio sampling data and the right audio sampling data to the audio processing device 20 via the communication unit 13. Hereinafter, when there is no particular distinction between the left audio sampling data and the right audio sampling data, they are also simply referred to as "audio sampling data."
[0029] The audio reproducing unit 17 acquires audio sampling data from the audio acquiring unit 16. The audio reproducing unit 17 receives a replay flag from the audio processing device 20 via the communication unit 13.
[0030] The replay flag is set to True or False. When the replay flag is False, the sound processing system 1 operates in through mode. Through mode is a mode in which sound data collected by the sound collector 10 is output from the sound collector 10 without going through the sound processing device 20. When the replay flag is True, the sound processing system 1 operates in playback mode. Playback mode is a mode in which the sound collector 10 outputs playback data acquired from the sound processing device 20. The conditions under which the replay flag is set to True or False will be described later.
[0031] When the replay flag is False, that is, when the audio processing system 1 is in the through mode, the audio reproducing unit 17 causes the speaker 12 to output the audio sampling data acquired from the audio acquiring unit 16 .
[0032] When the replay flag is True, that is, when the audio processing system 1 is in the playback mode, the audio playback unit 17 causes the speaker 12 to output the playback data stored in the storage unit 18 .
[0033] The audio playback unit 17 receives the notification sound file from the audio processing device 20 via the communication unit 13. When the audio processing device 20 detects audio that satisfies the set conditions, the notification sound file is transmitted from the audio processing device 20 to the audio playback unit 17. When the audio playback unit 17 receives the notification sound file, it outputs the notification sound to the speaker 12. With this configuration, the user can know that audio that satisfies the set conditions has been detected.
[0034] The storage unit 18 stores playback data. The playback data is data transmitted from the audio processing device 20 to the sound collector 10. When the control unit 15 receives playback data from the audio processing device 20 via the communication unit 13, it stores the received playback data in the storage unit 18. The control unit 15 can receive a playback stop instruction and a replay stop instruction, which will be described later, from the audio processing device 20 via the communication unit 13. When the control unit 15 receives a playback stop instruction or a replay stop instruction, it erases the playback data stored in the storage unit 18.
[0035] The control unit 15 may receive the left playback data and the right playback data from the audio processing device 20 and store them in the storage unit 18. In this case, the audio playback unit 17 may output the left playback data stored in the storage unit 18 to the left speaker of the speakers 12, and may output the right playback data stored in the storage unit 18 to the right speaker of the speakers 12.
[0036] The voice processing device 20 includes a communication unit 21 , an input unit 22 , a display unit 23 , a vibration unit 24 , a storage unit 26 , and a control unit 27 .
[0037] The communication unit 21 is configured to include at least one communication module capable of communicating with the sound collector 10 via a communication line. The communication module is a communication module that complies with the standard of the communication line. The standard of the communication line is, for example, a wired communication standard, or a short-range wireless communication standard including Bluetooth (registered trademark), infrared, NFC, etc.
[0038] The communication unit 21 may further include at least one communication module that can be connected to any network including a mobile communication network and the Internet, etc. The communication module is, for example, a communication module that complies with mobile communication standards such as LTE (Long Term Evolution), 4G (4th Generation), or 5G (5th Generation).
[0039] The input unit 22 can receive input from a user. The input unit 22 includes at least one input interface that can receive input from a user. The input interface is, for example, a physical key, a capacitance key, a pointing device, a touch screen that is integrated with a display, a microphone, or the like.
[0040] The display unit 23 is capable of displaying data. The display unit 23 is, for example, a display. The display is, for example, an LCD (Liquid Crystal Display) or an organic EL (Electro Luminescence) display.
[0041] The vibration unit 24 is capable of vibrating the sound processing device 20. The vibration unit 24 includes a vibration element. The vibration element is, for example, a piezoelectric element.
[0042] The light emitting unit 25 is capable of emitting light and is, for example, an LED (Light Emitting Diode).
[0043] The storage unit 26 is configured to include at least one semiconductor memory, at least one magnetic memory, at least one optical memory, or a combination of at least two of these. The semiconductor memory is, for example, a RAM or a ROM. The RAM is, for example, an SRAM or a DRAM. The ROM is, for example, an EEPROM. The storage unit 26 may function as a main storage device, an auxiliary storage device, or a cache memory. The storage unit 26 stores data used in the operation of the audio processing device 20 and data obtained by the operation of the audio processing device 20. For example, the storage unit 26 stores system programs, application programs, embedded software, etc.
[0044] The storage unit 26 stores, for example, a search list as shown in Fig. 3, which will be described later, and a notification sound list and notification sound files as shown in Fig. 4, which will be described later. The storage unit 26 stores, for example, a notification list, which will be described later.
[0045] The control unit 27 is configured to include at least one processor, at least one dedicated circuit, or a combination of these. The processor is a general-purpose processor such as a CPU or a GPU, or a dedicated processor specialized for a specific process. The dedicated circuit is, for example, an FPGA or an ASIC. The control unit 27 executes processes related to the operation of the audio processing device 20 while controlling each unit of the audio processing device 20.
[0046] The control unit 27 executes a voice recognition process to recognize voice data. However, the control unit 27 may also acquire the result of the voice recognition process executed by an external device. When the control unit 27 detects voice that satisfies a set condition based on the result of the voice recognition process, the control unit 27 notifies the user that voice that satisfies the set condition has been detected. The set condition is a condition that is set in advance regarding voice. The control unit 27 notifies the user that voice that satisfies the set condition has been detected in accordance with a notification condition corresponding to the detected voice that is set in the set condition.
[0047] The notification condition is a condition for determining the priority for notifying the user that sound has been detected. The higher the priority, the earlier the timing of notifying the user may be. The higher the priority, the more easily the user may be notified by a notification means. For example, as described above, the user is wearing earphones that are the sound collector 10. Therefore, when sound such as a notification sound is output from the sound collector 10, the user can immediately notice the sound. In other words, a notification means using sound such as a notification sound has a higher priority than a notification means such as presenting visual information. As will be described later, the user may be notified that sound has been detected by playing the detected sound. In this case, the higher the priority, the earlier the timing of playing the detected sound may be. Furthermore, when the priority is low, the detected sound may be played at any timing.
[0048] The notification condition includes a first condition and a second condition. If the notification condition satisfies the second condition, the priority of notifying the user is lower than if the notification condition satisfies the first condition. The first condition includes a third condition and a fourth condition. If the notification condition satisfies the fourth condition, the priority of notifying the user is lower than if the notification condition satisfies the third condition.
[0049] In this embodiment, the notification conditions are set by priority. The priority indicates the order of priority for notifying the user of audio that satisfies the set conditions. The priority may be set in multiple stages. The higher the priority, the higher the priority for notifying the user. In this embodiment, the priority is set in three stages including "high", "medium", and "low", as shown in Figures 3 and 4. The "high" priority is the highest priority of the three priority stages. The "medium" priority is the middle priority of the three priority stages. The "low" priority is the lowest priority of the three priority stages.
[0050] In this embodiment, the notification condition satisfies the first condition when the priority is "high" or "medium". The notification condition satisfies the second condition when the priority is "low". The notification condition satisfies the third condition when the priority is "high". The notification condition satisfies the fourth condition when the priority is "medium".
[0051] When the notification condition corresponding to the detected voice satisfies the first condition, that is, when the priority corresponding to the detected voice is "medium" or "high," the control unit 27 may notify the user that voice satisfying the set condition has been detected by playing a notification sound. With this configuration, the user can immediately notice that voice has been detected.
[0052] When the notification condition corresponding to the detected sound satisfies the first condition, i.e., when the priority corresponding to the detected sound is "medium" or "high," the control unit 27 may play the notification sound and then play the sound that satisfies the set condition. With this configuration, when the priority is "medium" or "high," the detected sound is automatically played. When the priority is "medium" or "high," there is a high possibility that the user will want to immediately check the content of the detected sound. By automatically playing the sound that satisfies the set condition when the notification condition satisfies the first condition, the user can immediately check the detected sound. This improves user convenience.
[0053] When the notification condition corresponding to the detected sound satisfies the second condition, that is, when the priority corresponding to the detected sound is "low," the control unit 27 may notify the user that sound satisfying the set condition has been detected by presenting visual information to the user. When the priority is "low," a notification appropriate for the low priority can be given by presenting visual information to the user instead of playing a notification sound.
[0054] In this embodiment, the set condition is that the speech contains a preset search word. A priority is set for each search word. The search word may contain, for example, at least one of letters and numbers. The search word may be any information that can be processed as text data. In this embodiment, when the control unit 27 detects an utterance containing the search word as a voice that satisfies the set condition, it notifies the user that the utterance has been detected.
[0055] FIG. 3 shows a search list. The search list associates search words with priorities set for the search words. A "high" priority is set for the search word "flight 153." A "medium" priority is set for the search word "hello." A "low" priority is set for the search word "good morning." For example, the control unit 27 generates the search list based on a user's input to a setting screen 50 as shown in FIG. 7, which will be described later.
[0056] FIG. 4 shows a notification sound list. The notification sound list associates priorities with notification sounds set for the priorities. The notification sound is used to notify the user that an utterance has been detected. In this embodiment, the notification sound is also used when an utterance corresponding to a priority of "low" is detected. However, when an utterance corresponding to a priority of "low" is detected, the notification sound does not have to be used. In FIG. 4, priorities are associated with notification sound files. The notification sound file is a file for storing the notification sound on a computer. The notification sound file "ring.wav" is associated with a priority of "high." The notification sound file "alert.wav" is associated with a priority of "medium." The notification sound file "notify.wav" is associated with a priority of "low."
[0057] Notification according to priority according to this embodiment will be described with reference to Fig. 5. In Fig. 5, notification means is means for notifying a user that an utterance including a search word has been detected. Notification timing is the timing at which the user is notified that an utterance including a search word has been detected. Playback timing is the timing at which the detected utterance is played back.
[0058] As shown in FIG. 5 , when the control unit 27 detects an utterance corresponding to a priority of “high,” the control unit 27 uses a notification sound and vibration by the vibration unit 24 as notification means. The control unit 27 sets the notification timing to the timing immediately after detecting a search word. The control unit 27 sets the playback timing to the timing immediately after notifying that the utterance has been detected. That is, when the notification condition satisfies the third condition, the control unit 27 controls so that the notification sound is played and the playback of the utterance begins immediately after detecting the search word. For example, assume that the priority of “high” is set for the search word “flight 153.” In this case, the control unit 27 sets the notification timing to the timing immediately after detecting the search word “flight 153.” That is, the control unit 27 vibrates the vibration unit 24 and plays the notification sound immediately after detecting the search word “flight 153.” Furthermore, the control unit 27 sets the playback timing to the timing immediately after playing the notification sound, and controls so that the utterance “Flight 153 is scheduled to depart 20 minutes late” is played. With this configuration, immediately after detecting the search word "flight 153," a notification sound is played and playback of the utterance "Flight 153 is scheduled to depart 20 minutes late" begins. In addition, utterances containing the search word are automatically played.
[0059] As shown in FIG. 5 , when the control unit 27 detects an utterance corresponding to a priority of “medium,” the control unit 27 uses a notification sound and vibration by the vibration unit 24 as notification means. The control unit 27 sets the notification timing to the timing immediately after the utterance including the search word ends. The control unit 27 sets the playback timing to the timing immediately after notifying that the utterance has been detected. That is, when the notification condition satisfies the fourth condition, the control unit 27 controls so that the notification sound is played and playback of the utterance begins immediately after the utterance including the search word ends. For example, assume that the priority of “medium” is set for the search word “flight 153.” In this case, the control unit 27 sets the notification timing to the timing immediately after the utterance “flight 153 is scheduled to depart 20 minutes late” ends, and vibrates the vibration unit 24 and plays the notification sound. Furthermore, the control unit 27 controls so that the playback timing is immediately after the utterance “flight 153 is scheduled to depart 20 minutes late” is played. With this configuration, immediately after the utterance "Flight 153 is scheduled to depart 20 minutes late" is finished, a notification sound is played and playback of the utterance "Flight 153 is scheduled to depart 20 minutes late" begins. In addition, utterances that include the search word are automatically played.
[0060] As shown in FIG. 5, when the control unit 27 detects an utterance corresponding to a priority of "low," the control unit 27 uses screen display by the display unit 23 and light emission by the light-emitting unit 25 as notification means. Screen display and light emission are examples of notification means for presenting visual information to the user. The control unit 27 causes the display unit 23 to display a notification list as the screen display. The notification list is a list of audio information that satisfies the detected setting conditions. In this embodiment, the notification list is a list of event information. An event is an utterance that includes a search word. Details of the notification list will be described later. The control unit 27 sets the playback timing to immediately after the user instructs playback of the utterance. In other words, when the notification condition satisfies the second condition, the control unit 27 plays back the utterance that includes the search word based on the user's input. With this configuration, the utterance that includes the search word is played back manually.
[0061] <Input / output processing> The control unit 27 receives input from the user through the input unit 22. The control unit 27 selects a screen to be displayed on the display unit 23 based on the input received by the input unit 22. For example, the control unit 27 displays a screen such as that shown in FIG. 6, FIG. 7, or FIG. 8 based on the input received by the input unit 22. In the configurations shown in FIGS. 6 to 8, the input unit 22 is a touch screen that is provided integrally with the display of the display unit 23.
[0062] The main screen 40 shown in FIG. 6 includes an area 41, an area 42, an area 43, and an area 44.
[0063] The area 41 displays the state of the sound collector 10. In Fig. 6, the area 41 displays the information "Replaying..." which indicates that the sound collector 10 is replaying.
[0064] When the audio processing system 1 is in the through mode, the words "Start replay" are displayed in the area 42. When the audio processing system 1 is in the playback mode, the words "Stop replay" are displayed in the area 42. The control unit 27 can accept input to the area 42 via the input unit 22.
[0065] When the words "Start replay" are displayed in the area 42, that is, when the voice processing system 1 is in the through mode, the control unit 27 can accept a request to start a replay by receiving an input to the area 42 via the input unit 22. When the control unit 27 accepts a request to start a replay, it sets a replay flag to True and outputs a replay instruction to the utterance accumulation unit 32, which will be described later.
[0066] When the words "Stop Replay" are displayed in the area 42, that is, when the sound processing system 1 is in playback mode, the control unit 27 can accept a command to stop replay by accepting an input to the area 42 via the input unit 22. When the control unit 27 accepts a command to stop replay, it sets the replay flag to False and transmits a command to stop replay to the sound collector 10 via the communication unit 21.
[0067] The words "Notification List" are displayed in area 43. Control unit 27 can accept input to area 43 via input unit 22. When control unit 27 accepts input to area 43 via input unit 22, it causes display unit 23 to display a notification screen 60 as shown in FIG.
[0068] The word "Settings" is displayed in area 44. Control unit 27 can accept input to area 44 via input unit 22. When control unit 27 accepts input to area 44 via input unit 22, it causes display unit 23 to display a setting screen 50 as shown in FIG.
[0069] 7 is a screen for the user to make various settings. The setting screen 50 includes an area 51, an area 52, an area 53, an area 54, an area 55, and an area 56.
[0070] The words "Add search word" are displayed in area 51. Control unit 27 can accept input to area 51 via input unit 22. Control unit 27 accepts input of a search word and an input of a priority corresponding to the search word from area 51.
[0071] The set search word is displayed in area 52. In FIG. 7, the search words "flight 153," "hello," and "good morning" are displayed in area 52. The control unit 27 can accept input to area 52 by the input unit 22. When the control unit 27 accepts input to area 52 by the input unit 22, it causes the display unit 23 to display a search list such as that shown in FIG. 3.
[0072] The words "Recording buffer setting" are displayed in area 53. Area 53 is used to set the length of recording time for recording sound collected by the sound collector 10. In this embodiment, sound sampling data for the recording time is accumulated in a ring buffer 34 as shown in FIG. 10 , which will be described later. The control unit 27 can accept input to area 53 via the input unit 22. The control unit 27 accepts input of recording times such as 5 seconds, 10 seconds, and 15 seconds. The control unit 27 stores the accepted recording time information in the memory unit 26.
[0073] The area 54 displays the words "speed setting." The area 54 is used to set the playback speed of the sound output from the sound collector 10. The control unit 27 can accept input to the area 54 via the input unit 22. The control unit 27 accepts input of sound speeds such as 1x speed, 1.1x speed, and 1.2x speed. The control unit 27 stores the accepted sound speed information in the memory unit 26.
[0074] The words "Audio Threshold Setting" are displayed in area 55. Area 55 is used to set an audio threshold at which audio collected by the sound collector 10 is cut off as noise. In this embodiment, audio below the audio threshold is cut off as noise. The control unit 27 can accept input to area 55 via the input unit 22. The control unit 27 accepts input of audio thresholds ranging from -50 [dBA] to -5 [dBA], for example. The control unit 27 stores the information on the accepted audio threshold in the memory unit 26.
[0075] The words "SETTING END" are displayed in area 56. Control unit 27 can accept input to area 56 via input unit 22. When control unit 27 accepts input to area 56 via input unit 22, it causes main screen 40 as shown in FIG. 6 to be displayed on display unit 23.
[0076] 8 is a screen for notifying the user of various information. The notification screen 60 includes an area 61, an area 62, an area 63, and an area 64.
[0077] A notification list is displayed in area 61. As described above, the notification list is a list of event information. As described above, an event is an utterance including a search word. The control unit 27 displays event information of events with a "low" priority among the events included in the notification list in area 61. However, the control unit 27 may display all event information included in the notification list in area 61 regardless of priority. The control unit 27 can accept input for each event in the notification list displayed in area 61 via the input unit 22. The control unit 27 accepts selection of an event in the notification list by accepting input for each event in the notification list from area 61 via the input unit 22.
[0078] The words "display details" are displayed in area 62. The control unit 27 can accept input to area 62 via the input unit 22. The control unit 27 can accept a selection of an event included in the notification list from area 61, and can also accept input to area 62 via the input unit 22. In this case, the control unit 27 causes the display unit 23 to display details of the event information selected from area 61 in area 61. For example, the control unit 27 causes the display unit 23 to display the voice recognition results for the left and right, which will be described later, as details of the event information.
[0079] The words "Start Playback / Stop Playback" are displayed in area 63. The control unit 27 can accept a playback start or playback stop by accepting an input to area 63 via the input unit 22. When an utterance is not being played back, the control unit 27 accepts a selection of an event included in the notification list from area 61, and further accepts an input to area 63 via the input unit 22 to accept a playback start of the event. When an utterance is not being played back, the control unit 27 accepts a playback start of the event by accepting an input to area 63 via the input unit 22. When an utterance is being played back, the control unit 27 controls the playback of the event selected from area 61, i.e., the utterance. In this embodiment, the control unit 27 references the notification list in the storage unit 26 and acquires an event ID (described below) of the selected event from area 61. The control unit 27 outputs the event ID and a playback start instruction to the utterance storage unit 36 (described below) and controls the playback of the utterance. Furthermore, when an utterance is being played back, the control unit 27 accepts an input to area 63 via the input unit 22 to accept a playback stop of the event. When a command to stop playback of an event is received, the control unit 27 controls the playback of the speech to stop. In this embodiment, a playback stop command is sent to the sound collector 10 by the communication unit 21, and control is performed to stop the playback of the speech.
[0080] The word "Back" is displayed in area 64. The control unit 27 can accept input to area 64 via the input unit 22. When the control unit 27 accepts input to area 64 via the input unit 22, it causes the display unit 23 to display a main screen 40 as shown in FIG.
[0081] <Audio processing> 2, control unit 27 includes a section detection unit 28, a voice recognition unit 29, an event detection unit 30, an utterance notification unit 31, an utterance accumulation unit 32, a voice modulation unit 35, and an utterance holding unit 36. Utterance holding unit 36 is configured to include components that are the same as or similar to those of memory unit 26. At least a part of utterance holding unit 36 may be part of memory unit 26. The operation of utterance holding unit 36 is executed by a processor of control unit 27, etc.
[0082] The section detection unit 28 receives audio sampling data from the sound collector 10 via the communication unit 21. The section detection unit 28 detects speech sections from the audio sampling data. A speech section is a section in which a speech state continues. The section detection unit 28 can also detect non-speech sections by detecting speech sections from the audio sampling data. A non-speech section is a section in which a non-speech state continues. The start point of a speech section is also referred to as the "start time of speech." The start point of a speech section is the end point of a non-speech section. The end point of a speech section is also referred to as the "end time of speech." The end point of a speech section is the start point of a non-speech section.
[0083] An example of the processing of the section detection unit 28 will be described with reference to Fig. 9. However, the processing of the section detection unit 28 is not limited to the processing described with reference to Fig. 9. The section detection unit 28 may detect the speech section from the audio sampling data by any method. As another example, the section detection unit 28 may detect the speech section from the audio sampling data by a machine learning model generated using any machine learning algorithm.
[0084] In Fig. 9, the horizontal axis represents time. The audio sampling data shown in Fig. 9 is acquired by the audio acquisition unit 16 of the sound collector 10. The section detection unit 28 acquires audio activity detection data from the audio sampling data. The audio activity detection data is data obtained by averaging the power of the audio sampling data over a preset time width. The time width of the audio activity detection data may be set based on the specifications of the audio processing device 20, etc. In Fig. 9, one piece of audio activity detection data is shown as one square. The time width of this one square, i.e., the time width of one piece of audio activity detection data, is, for example, 200 [ms].
[0085] The section detection unit 28 acquires information about the voice threshold from the storage unit 26 and classifies the voice activity detection data into voice data and non-voice data. In FIG. 9, voice data is data that is darkly shaded among the voice activity detection data shown as squares. Non-voice data is data that is white among the voice activity detection data shown as squares. If the value of the voice activity detection data is null, the section detection unit 28 classifies the voice activity detection data as non-voice data. If the value of the voice activity detection data is not null and is less than the voice threshold, the section detection unit 28 classifies the voice activity detection data as non-voice data. If the value of the voice activity detection data is not null and is equal to or greater than the voice threshold, the section detection unit 28 classifies the voice activity detection data as voice data.
[0086] The section detection unit 28 detects a section in which voice data continues without a gap of a set time as a speech section. The set time may be set based on the language processed by the voice processing device 20. When the processed language is Japanese, the set time is, for example, 500 [ms]. In FIG. 9 , when the section detection unit 28 detects voice data after non-voice data has continued for more than the set time, the section detection unit 28 identifies the time when the voice data was detected as the speech start time. For example, the section detection unit 28 identifies time t1 as the speech start time. After identifying the speech start time, if the section detection unit 28 determines that the non-voice data has continued for more than the set time, the section detection unit 28 identifies the time when the determination was made as the speech end time. For example, the section detection unit 28 identifies time t2 as the speech end time. The section detection unit 28 detects the section from the speech start time to the speech end time as the speech section.
[0087] The section detection unit 28 may receive left audio sampling data and right audio sampling data from the sound collector 10. In this case, when non-audio data continues for a set time in both the left and right audio sampling data and then audio data is detected in either the left or right audio sampling data, the section detection unit 28 may identify the time point at which the audio data is detected as the speech start time. Furthermore, when the section detection unit 28 determines that non-audio data has continued for a set time in both the left and right audio sampling data, it may identify the time point at which this determination is made as the speech end time.
[0088] The section detection unit 28 identifies the start time of an utterance from the voice sampling data and generates an utterance ID. The utterance ID is identification information that can uniquely identify each utterance. The section detection unit 28 outputs the information on the start time of the utterance and the utterance ID to the voice recognition unit 29 and the utterance storage unit 32, respectively.
[0089] When the section detection unit 28 identifies the end point of the utterance from the voice sampling data, the section detection unit 28 outputs information on the end point of the utterance to the voice recognition unit 29 and the utterance storage unit 32, respectively.
[0090] The section detection unit 28 sequentially outputs the voice sampling data received from the sound collector 10 to the voice recognition unit 29 and the speech accumulation unit 32, respectively.
[0091] The speech recognition unit 29 acquires information on the start time of an utterance and an utterance ID from the section detection unit 28. When the speech recognition unit 29 acquires the information on the start time of an utterance and the like, it executes a speech recognition process to recognize speech on the speech sampling data successively acquired from the section detection unit 28. In this embodiment, the speech recognition unit 29 recognizes speech by converting the speech data included in the speech sampling data into text data through the speech recognition process.
[0092] The speech recognition unit 29 outputs the information on the start time of the utterance and the utterance ID acquired from the section detection unit 28 to the event detection unit 30. After outputting the information on the start time of the utterance and the like to the event detection unit 30, the speech recognition unit 29 sequentially outputs text data as the speech recognition result to the event detection unit 30.
[0093] The speech recognition unit 29 acquires information about the end of the speech from the section detection unit 28. When the speech recognition unit 29 acquires the information about the end of the speech, it terminates the speech recognition process. The speech recognition unit 29 outputs the information about the end of the speech acquired from the section detection unit 28 to the event detection unit 30. Thereafter, the speech recognition unit 29 can acquire information about the start of a new speech and an utterance ID from the section detection unit 28. When the speech recognition unit 29 acquires the information about the start of a new speech, etc., it performs the speech recognition process again on the speech sampling data successively acquired from the section detection unit 28.
[0094] The voice recognition unit 29 may acquire left voice sampling data and right voice sampling data from the section detection unit 28. In this case, the voice recognition unit 29 may convert each of the left voice sampling data and the right voice sampling data into text data. Hereinafter, the text data acquired from the left voice sampling data will also be referred to as "left text data" or "left voice recognition result." The text data acquired from the right voice sampling data will also be referred to as "right text data" or "right voice recognition result."
[0095] The event detection unit 30 acquires information about the start of an utterance and an utterance ID from the voice recognition unit 29. After acquiring the information about the start of an utterance, the event detection unit 30 sequentially acquires text data from the voice recognition unit 29. The event detection unit 30 refers to a search list such as that shown in Fig. 3 and determines whether the text data sequentially acquired from the voice recognition unit 29 includes any of the search words in the search list.
[0096] When the event detection unit 30 determines that the text data contains a search word, it detects the utterance containing the search word as an event. When the event detection unit 30 detects an event, it acquires the utterance ID acquired from the voice recognition unit 29 as an event ID. Furthermore, the event detection unit 30 refers to a search list such as that shown in FIG. 3 and acquires a priority corresponding to the search word contained in the text data. When the event detection unit 30 acquires the priority, it executes a notification process according to the priority.
[0097] When the priority is "high", if the event detection unit 30 determines that the text data contains a search word, it outputs an event ID and an output instruction to the utterance storage unit 32 and outputs a priority of "high" to the utterance notification unit 31. The output instruction is an instruction to the utterance storage unit 32 to output the audio sampling data corresponding to the event ID as playback data to the audio modulation unit 35. When the event detection unit 30 outputs the output instruction, it sets the replay flag to True. In this way, when the priority is "high", an output instruction etc. is output to the utterance storage unit 32 etc. immediately after the search word included in the text data is detected. With this configuration, when the priority is "high", as shown in FIG. 5, a notification sound is played and playback of the utterance begins immediately after the search word is detected.
[0098] When the priority is "medium", the event detection unit 30, upon acquiring information on the end of the utterance from the voice recognition unit 29, outputs an event ID and an output instruction to the utterance storage unit 32, and outputs a priority of "medium" to the utterance notification unit 31. When the event detection unit 30 outputs the output instruction, it sets the replay flag to True. In this way, when the priority is "medium", an output instruction etc. is output to the utterance storage unit 32 etc. at the end of the utterance. With this configuration, when the priority is "medium", as shown in FIG. 5, a notification sound is played and playback of the utterance begins immediately after the end of the utterance including the search word.
[0099] When the priority is "low," the event detection unit 30, upon acquiring information on the end point of the utterance from the voice recognition unit 29, outputs an event ID and a storage instruction to the utterance storage unit 32 and outputs a priority of "low" to the utterance notification unit 31. The storage instruction is an instruction to the utterance storage unit 32 to output the voice sampling data corresponding to the event ID to the utterance storage unit 36. The voice sampling data stored in the utterance storage unit 36 is played back when the user issues a playback instruction, as described above with reference to FIG. 8. With this configuration, when the priority is "low," as shown in FIG. 5, the utterance including the search word is played back immediately after the user issues a playback instruction.
[0100] The event detection unit 30 updates the notification list stored in the storage unit 26 based on the event ID, priority, detection date and time when the event was detected, and search words included in the text data. The notification list in the storage unit 26 includes, for example, associations between the event ID, priority, detection date and time when the event was detected, search words, and text data. As an example of the update process, the event detection unit 30 associates the event ID, priority, detection date and time, search words, and text data. The event detection unit 30 updates the notification list by including this association in the notification list.
[0101] The event detection unit 30 determines whether the search word is included in the text data until it acquires information about the end of the utterance from the voice recognition unit 29. If the event detection unit 30 determines that the search word is not included in the text data sequentially acquired from the voice recognition unit 29 at the time when it acquires the information about the end of the utterance, it acquires the utterance ID acquired from the voice recognition unit 29 as a clear event ID. The event detection unit 30 outputs the clear event ID to the utterance accumulation unit 32.
[0102] The event detection unit 30 can acquire information about the start time of a new utterance and an utterance ID from the voice recognition unit 29. When the event detection unit 30 acquires the information about the start time of a new utterance, etc., it determines whether or not any of the search words in the search list is included in the text data successively newly acquired from the voice recognition unit 29.
[0103] The event detection unit 30 may acquire left text data and right text data from the voice recognition unit 29. In this case, if the event detection unit 30 determines that the search word is included in either the left text data or the right text data, it may detect an utterance including the search word as an event. If the event detection unit 30 determines that the search word is not included in either the left text data or the right text data, it may acquire an utterance ID corresponding to the text data as a clear event ID.
[0104] The utterance notification unit 31 acquires the priority from the event detection unit 30. The utterance notification unit 31 acquires a notification sound file corresponding to the priority from the storage unit 26. The utterance notification unit 31 transmits the acquired notification sound file to the sound collector 10 via the communication unit 21.
[0105] If the priority is "high", the utterance notification unit 31 refers to the notification sound list as shown in Fig. 4 and acquires the notification sound file "ring.wav" associated with the priority of "high" from the storage unit 26. The utterance notification unit 31 transmits the acquired notification sound file to the sound collector 10 via the communication unit 21.
[0106] If the priority is "medium", the utterance notification unit 31 refers to the notification sound list as shown in Fig. 4 and acquires the notification sound file "alert.wav" associated with the priority of "medium" from the storage unit 26. The utterance notification unit 31 transmits the acquired notification sound file to the sound collector 10 via the communication unit 21.
[0107] If the priority is "low", the utterance notification unit 31 refers to the notification sound list as shown in Fig. 4 and acquires the notification sound file "notify.wav" associated with the priority of "low" from the storage unit 26. The utterance notification unit 31 transmits the acquired notification sound file to the sound collector 10 via the communication unit 21.
[0108] 10, the utterance accumulation unit 32 has a data buffer 33 and a ring buffer 34. The data buffer 33 and the ring buffer 34 are configured to include the same or similar components as those of the storage unit 26. At least a part of the data buffer 33 and the ring buffer 34 may be part of the storage unit 26. The operation of the utterance accumulation unit 32 is executed by a processor or the like of the control unit 27.
[0109] The utterance accumulation unit 32 acquires information on the start of an utterance and an utterance ID from the section detection unit 28. When the utterance accumulation unit 32 acquires the information on the start of an utterance, etc., the utterance accumulation unit 32 associates the voice sampling data successively acquired from the section detection unit 28 with the utterance ID and accumulates it in the data buffer 33. When the utterance accumulation unit 32 acquires information on a new start of an utterance and a new utterance ID from the section detection unit 28, the utterance accumulation unit 32 associates the voice sampling data successively acquired from the section detection unit 28 with the new utterance ID and accumulates it in the data buffer 33. In FIG. 10 , the data buffer 33 accumulates a plurality of voice sampling data corresponding to utterance ID1, a plurality of voice sampling data corresponding to utterance ID2, and a plurality of voice sampling data corresponding to utterance ID3.
[0110] The speech accumulation unit 32 receives voice sampling data from the sound collector 10 via the communication unit 21. The speech accumulation unit 32 accumulates the voice sampling data received from the sound collector 10 in the ring buffer 34. The speech accumulation unit 32 references the recording time information stored in the memory unit 26, and accumulates voice sampling data for the recording time in the ring buffer 34. The speech accumulation unit 32 accumulates the voice sampling data sequentially in chronological order in the ring buffer 34.
[0111] The utterance accumulation unit 32 can acquire a clear event ID from the event detection unit 30. When the utterance accumulation unit 32 acquires a clear event ID, the utterance accumulation unit 32 deletes, from the voice sampling data accumulated in the data buffer 33, the voice sampling data associated with the utterance ID that matches the clear event ID.
[0112] The utterance accumulation unit 32 can acquire an event ID and an output instruction from the event detection unit 30. When the utterance accumulation unit 32 acquires the output instruction, it identifies an utterance ID that matches the event ID acquired together with the output instruction from among the voice sampling data accumulated in the data buffer 33. The utterance accumulation unit 32 outputs the voice sampling data corresponding to the identified utterance ID as playback data to the voice modulation unit 35. The utterance accumulation unit 32 outputs the voice sampling data to the voice modulation unit 35 so that the voice sampling data is played back from the first voice sampling data. The first voice sampling data is the voice sampling data with the oldest time among multiple voice sampling data arranged in time series.
[0113] The utterance accumulation unit 32 can acquire an event ID and a storage instruction from the event detection unit 30. Upon acquiring the storage instruction, the utterance accumulation unit 32 identifies an utterance ID that matches the event ID acquired together with the storage instruction from among the voice sampling data accumulated in the data buffer 33. The utterance accumulation unit 32 outputs the voice sampling data associated with the identified utterance ID to the utterance storage unit 36 together with the event ID.
[0114] The utterance accumulation unit 32 may receive a replay instruction. When the utterance accumulation unit 32 receives a replay instruction, the utterance accumulation unit 32 outputs the audio sampling data accumulated in the ring buffer 34 to the audio modulation unit 35 as playback data so that the audio sampling data is played back from the leading audio sampling data.
[0115] 2, the audio modulation unit 35 acquires playback data from the speech accumulation unit 32. When the replay flag is True, the audio modulation unit 35 references the audio speed information stored in the storage unit 26 and modulates the playback data so that the playback data is played back as audio at that audio speed. The audio modulation unit 35 transmits the modulated playback data to the sound collector 10 via the communication unit 21.
[0116] The utterance holding unit 36 acquires the event ID and the voice sampling data from the utterance accumulation unit 32. The utterance holding unit 36 stores the acquired voice sampling data in association with the acquired event ID.
[0117] The utterance storage unit 36 can acquire the event ID and the playback start instruction. Upon acquiring the playback start instruction, the utterance storage unit 36 identifies the audio sampling data associated with the event ID. The utterance storage unit 36 transmits the identified audio sampling data as playback data to the sound collector 10 via the communication unit 21.
[0118] Fig. 11 is a flowchart showing the operation of the event detection process executed by the sound processing device 20 shown in Fig. 2. This operation corresponds to an example of the sound processing method according to the present embodiment. For example, when transmission of sound sampling data from the sound collector 10 to the sound processing device 20 starts, the sound processing device 20 starts the process of step S1.
[0119] The section detection unit 28 receives the sound sampling data from the sound collector 10 via the communication unit 21 (step S1).
[0120] In the process of step S2, the section detection unit 28 sequentially outputs the voice sampling data acquired in the process of step S1 to the voice recognition unit 29 and the utterance accumulation unit 32, respectively.
[0121] In the process of step S2, the section detection unit 28 identifies the start time of the utterance from the voice sampling data acquired in the process of step S1. After identifying the start time of the utterance, the section detection unit 28 generates an utterance ID. The section detection unit 28 outputs the information on the start time of the utterance and the utterance ID to the voice recognition unit 29 and the utterance accumulation unit 32, respectively.
[0122] In the process of step S2, the section detection unit 28 identifies the end point of the utterance from the voice sampling data acquired in the process of step S1. After identifying the end point of the utterance, the section detection unit 28 outputs information on the end point of the utterance to the voice recognition unit 29 and the utterance accumulation unit 32.
[0123] In the processing of step S3, when the speech recognition unit 29 acquires information such as the start of an utterance from the section detection unit 28, it sequentially converts the speech sampling data sequentially acquired from the section detection unit 28 into text data. When the speech recognition unit 29 outputs information such as the start of an utterance to the event detection unit 30, it sequentially outputs text data as a speech recognition result to the event detection unit 30. When the speech recognition unit 29 acquires information such as the end of an utterance from the section detection unit 28, it terminates the speech recognition process. However, when the speech recognition unit 29 acquires information such as a new start of an utterance from the section detection unit 28, it sequentially converts the speech sampling data sequentially acquired from the section detection unit 28 into text data.
[0124] In the process of step S4, the event detection unit 30 refers to the search list as shown in FIG. 3 and determines whether or not the text data sequentially acquired from the voice recognition unit 29 includes any of the search words in the search list.
[0125] If the event detection unit 30 determines that the search word is not included in the successively acquired text data at the time when the information on the end of the utterance is acquired from the voice recognition unit 29 (step S4: NO), the event detection unit 30 proceeds to the processing of step S5. If the event detection unit 30 determines that the search word is included in the successively acquired text data from the voice recognition unit 29 before acquiring the information on the end of the utterance (step S4: YES), the event detection unit 30 proceeds to the processing of step S6.
[0126] In the process of step S5, the event detection unit 30 acquires the utterance ID acquired from the voice recognition unit 29 as a clear event ID. The event detection unit 30 outputs the clear event ID to the utterance accumulation unit 32.
[0127] In the process of step S6, the event detection unit 30 detects an utterance including the search word as an event.
[0128] In the process of step S7, the event detection unit 30 acquires, as an event ID, the utterance ID acquired from the voice recognition unit 29. Furthermore, the event detection unit 30 refers to the search list as shown in Fig. 3 and acquires a priority corresponding to the search word included in the text data.
[0129] In the process of step S8, the event detection unit 30 executes a notification process according to the priority acquired in the process of step S7.
[0130] In the process of step S9, the event detection unit 30 updates the notification list stored in the storage unit 26 based on the event ID, the priority, the detection date and time when the event was detected, and the search word included in the text data.
[0131] 12 and 13 are flowcharts showing the operation of the playback data output process executed by the sound processing device 20 shown in Fig. 2. This operation corresponds to an example of the sound processing method according to the present embodiment. For example, when transmission of sound sampling data from the sound collector 10 to the sound processing device 20 starts, the sound processing device 20 starts the process of step S11 as shown in Fig. 12.
[0132] In the process of step S11, the sound processing device 20 operates in the through mode. In the sound collector 10, the sound reproducing unit 17 outputs the sound sampling data acquired from the sound acquiring unit 16 to the speaker 12. In the process of step S11, the replay flag is set to False.
[0133] In the process of step S12, the control unit 27 determines whether or not a replay start has been accepted by accepting an input to an area 42 as shown in Fig. 6 from the input unit 22. If the control unit 27 determines that a replay start has been accepted (step S12: YES), the control unit 27 proceeds to the process of step S13. If the control unit 27 does not determine that a replay start has been accepted (step S12: NO), the control unit 27 proceeds to the process of step S18.
[0134] In the process of step S13, the control unit 27 sets the replay flag to True and outputs a replay instruction to the utterance accumulation unit 32.
[0135] In the process of step S14, the utterance accumulation unit 32 acquires a replay instruction. Upon acquiring the replay instruction, the utterance accumulation unit 32 starts outputting the reproduced data from the ring buffer to the audio modulation unit .
[0136] In the process of step S15, the control unit 27 determines whether or not all of the reproduced data has been output from the ring buffer 34 to the audio modulation unit 35. If the control unit 27 determines that all of the reproduced data has been output (step S15: YES), the control unit 27 proceeds to the process of step S17. If the control unit 27 does not determine that all of the reproduced data has been output (step S15: NO), the control unit 27 proceeds to the process of step S16.
[0137] In the process of step S16, the control unit 27 determines whether or not a replay stop has been received by receiving an input to an area 42 as shown in Fig. 6 from the input unit 22. If the control unit 27 determines that a replay stop has been received (step S16: YES), the control unit 27 proceeds to the process of step S17. If the control unit 27 does not determine that a replay stop has been received (step S16: NO), the control unit 27 returns to the process of step S15.
[0138] In the process of step S17, the control unit 27 sets the replay flag to False. After executing the process of step S17, the control unit 27 returns to the process of step S11.
[0139] In the process of step S18, the control unit 27 determines whether or not a request to start playing an event has been received by receiving an input to an area 63 as shown in Fig. 8 via the input unit 22. If the control unit 27 determines that a request to start playing an event has been received (step S18: YES), the control unit 27 proceeds to the process of step S19. If the control unit 27 does not determine that a request to start playing an event has been received (step S18: NO), the control unit 27 proceeds to the process of step S24 as shown in Fig. 13.
[0140] In the process of step S19, the control unit 27 sets the replay flag to True. The control unit 27 also refers to the notification list in the storage unit 26 and acquires the event ID of the selected event from the area 61 as shown in Fig. 8. The control unit 27 outputs the event ID and a playback start instruction to the utterance storage unit 36.
[0141] In the processing of step S20, the utterance storage unit 36 acquires an event ID and a playback start instruction. Upon acquiring the playback start instruction, the utterance storage unit 36 identifies the audio sampling data associated with the event ID. The utterance storage unit 36 starts transmitting the identified audio sampling data, i.e., playback data, to the sound collector 10.
[0142] In the process of step S21, the control unit 27 determines whether or not all of the playback data has been transmitted from the speech storage unit 36 to the sound collector 10. If the control unit 27 determines that all of the playback data has been transmitted (step S21: YES), the control unit 27 proceeds to the process of step S23. If the control unit 27 does not determine that all of the playback data has been transmitted (step S21: NO), the control unit 27 proceeds to the process of step S22.
[0143] In the process of step S22, the control unit 27 determines whether or not a request to stop playing an event has been received by receiving an input to an area 63 as shown in Fig. 8 via the input unit 22. If the control unit 27 determines that a request to stop playing an event has been received (step S22: YES), the control unit 27 proceeds to the process of step S23. If the control unit 27 does not determine that a request to stop playing an event has been received (step S22: NO), the control unit 27 returns to the process of step S21.
[0144] In the process of step S23, the control unit 27 sets the replay flag to False. After executing the process of step S23, the control unit 27 returns to the process of step S11.
[0145] 13, the utterance accumulation unit 32 determines whether or not an event ID and an output instruction have been acquired from the event detection unit 30. If the utterance accumulation unit 32 determines that an event ID and an output instruction have been acquired (step S24: YES), the utterance accumulation unit 32 proceeds to the processing of step S25. If the utterance accumulation unit 32 does not determine that an event ID and an output instruction have been acquired (step S24: NO), the utterance accumulation unit 32 proceeds to the processing of step S30.
[0146] In the process of step S25, the replay flag is set to True. This replay flag is set to True by the event detection unit 30 when the event detection unit 30 outputs the output instruction in the process of step S24 to the utterance accumulation unit 32.
[0147] In the process of step S26, the utterance accumulation unit 32 identifies an utterance ID that matches the event ID acquired in the process of step S24 from the voice sampling data accumulated in the data buffer 33. The utterance accumulation unit 32 acquires the voice sampling data corresponding to the identified utterance ID as playback data. The utterance accumulation unit 32 starts outputting the playback data from the data buffer 33 to the voice modulation unit 35.
[0148] In the process of step S27, the control unit 27 determines whether or not all of the reproduced data has been output from the data buffer 33 to the audio modulation unit 35. If the control unit 27 determines that all of the reproduced data has been output (step S27: YES), the control unit 27 proceeds to the process of step S29. If the control unit 27 does not determine that all of the reproduced data has been output (step S27: NO), the control unit 27 proceeds to the process of step S28.
[0149] In the process of step S28, the control unit 27 determines whether or not a replay stop has been received by receiving an input to an area 42 as shown in Fig. 6 from the input unit 22. If the control unit 27 determines that a replay stop has been received (step S28: YES), the control unit 27 proceeds to the process of step S29. If the control unit 27 does not determine that a replay stop has been received (step S28: NO), the control unit 27 returns to the process of step S27.
[0150] In the process of step S29, the control unit 27 sets the replay flag to False. After executing the process of step S29, the control unit 27 returns to the process of step S11 as shown in FIG.
[0151] In the process of step S30, the utterance accumulation unit 32 determines whether or not an event ID and a holding instruction have been acquired from the event detection unit 30. If the utterance accumulation unit 32 determines that an event ID and a holding instruction have been acquired (step S30: YES), the utterance accumulation unit 32 proceeds to the process of step S31. If the utterance accumulation unit 32 does not determine that an event ID and a holding instruction have been acquired (step S30: NO), the control unit 27 returns to the process of step S11 as shown in FIG.
[0152] In the process of step S31, the utterance accumulation unit 32 identifies an utterance ID that matches the event ID acquired in the process of step S30 from the voice sampling data accumulated in the data buffer 33. The utterance accumulation unit 32 outputs the voice sampling data associated with the identified utterance ID to the utterance holding unit 36 together with the event ID.
[0153] After executing the process of step S31, the control unit 27 returns to the process of step S11 as shown in FIG.
[0154] In this way, in the voice processing device 20, when the control unit 27 detects a voice that satisfies the set condition, it notifies the user that the voice that satisfies the set condition has been detected, according to the notification condition. In this embodiment, when the control unit 27 detects an utterance that includes a search word as a voice that satisfies the set condition, it notifies the user that the utterance has been detected, according to the priority set for the search word.
[0155] Here, the user may prefer to be notified of the detection of a voice that satisfies a set condition in some cases, or may not prefer to be notified of the detection in other cases, depending on the content of the voice. In this embodiment, the user can set a priority as a notification condition, thereby separating voices into those for which the user preferentially notifies of the detection and those for which the user does not preferentially notify of the detection. Therefore, the voice processing device 20 can improve user convenience.
[0156] Furthermore, when a voice that satisfies a set condition is detected, if the voice is simply played back, the user may miss the played back voice. By notifying the user that a voice has been detected, the voice processing device 20 can reduce the possibility that the user will miss the played back voice.
[0157] Therefore, according to the present embodiment, an improved voice processing device 20, voice processing method, and voice processing system 1 can be provided.
[0158] Furthermore, when the notification condition corresponding to the detected voice satisfies the first condition, the control unit 27 of the voice processing device 20 may notify the user that voice satisfying the set condition has been detected by playing a notification sound. As described above, by playing the notification sound, the user can immediately notice that voice has been detected.
[0159] Furthermore, when the notification condition corresponding to the detected sound satisfies the first condition, the control unit 27 of the sound processing device 20 may play a sound that satisfies the set condition after playing the notification sound. With this configuration, it is possible to improve the convenience for the user, as described above.
[0160] Furthermore, when the notification condition corresponding to the detected sound satisfies the second condition, the control unit 27 of the sound processing device 20 may notify the user that sound satisfying the set condition has been detected by presenting visual information to the user. As described above, the second condition has a lower priority for notifying the user than the first condition, the third condition, and the fourth condition. When the priority is low, a notification commensurate with the low priority can be given by presenting visual information to the user instead of playing a notification sound.
[0161] Furthermore, when the notification condition corresponding to the detected voice satisfies the second condition, the control unit 27 of the voice processing device 20 may present visual information to the user by presenting a notification list to the user. By looking at the notification list, the user can understand the date and time when the voice was detected and the circumstances under which the voice was detected.
[0162] Furthermore, the control unit 27 of the voice processing device 20 may play the detected voice based on the user's input if the notification condition corresponding to the detected voice satisfies the second condition. If the priority is low, it is highly likely that the user will want to check the detected voice later. This configuration can improve user convenience.
[0163] Furthermore, the control unit 27 of the voice processing device 20 may perform control so that, if the notification condition corresponding to the detected voice satisfies the third condition, a notification sound is played and playback of the utterance is started immediately after detecting the search word. With this configuration, if the priority is high, the user can immediately check the content of the utterance.
[0164] Furthermore, when the notification condition corresponding to the detected voice satisfies the fourth condition, the control unit 27 of the voice processing device 20 may perform control so that the notification sound is played and playback of the voice is started immediately after the voice ends. By starting playback of the voice immediately after the voice ends, the voice uttered in real time and the played voice do not overlap. This configuration allows the user to more accurately grasp the content of the played voice.
[0165] Furthermore, the control unit 27 of the voice processing device 20 may perform control so that audio data of an utterance section including the detected utterance is played back. As described above with reference to FIG. 9, an utterance section is a section in which audio data continues without a set time gap. By playing back the audio data of such an utterance section, a block of utterance including the search word is played back. This configuration allows the user to understand the meaning of the utterance including the search word.
[0166] Furthermore, regardless of the priority, the control unit 27 of the voice processing device 20 may display, from the information included in the notification list, the detection date and time when the event, i.e., the utterance, was detected and the search word included in the utterance on the display unit 23. With this configuration, the user can understand how the detected utterance was uttered.
[0167] (Other embodiments) 14 can provide a monitoring service for a baby, etc. The sound processing system 101 includes a sound collector 110 and a sound processing device 20.
[0168] The sound collector 110 and the sound processing device 20 are located farther apart from each other than the sound collector 10 and the sound processing device 20 shown in Fig. 1. For example, the sound collector 110 and the sound processing device 20 are located in separate rooms. The sound collector 110 is located in a room where a baby is present. The sound processing device 20 is located in a room where a user is present.
[0169] In another embodiment, the setting condition is a condition that the sound matches a predetermined sound characteristic. The user may input the sound characteristic that the user wants to set as the setting condition from a microphone of the input unit 22 of the sound processing device 20 as shown in FIG. 2 and set the sound characteristic as the setting condition in the sound processing device 20. For example, the user may set the sound characteristic of a baby's crying as the setting condition in the sound processing device 20.
[0170] 2, the sound collector 110 includes a microphone 11, a speaker 12, a communication unit 13, a storage unit 14, and a control unit 15. The sound collector 110 does not necessarily have to include the speaker 12.
[0171] The audio processing device 20 may further include a speaker 12 as shown in Fig. 2. The control unit 27 of the audio processing device 20 may further include an audio playback unit 17 and a storage unit 18 as shown in Fig. 2.
[0172] In another embodiment, the storage unit 26 as shown in Fig. 2 stores a search list in which data indicating voice characteristics is associated with a priority level, instead of the search list as shown in Fig. 3. The data indicating voice characteristics may be data on voice features that can be processed by a machine learning model used by the voice recognition unit 29. The voice features are, for example, Mel-Frequency Cepstral Coefficients (MFCC) or Perceptual Linear Prediction (PLP). For example, the storage unit 26 stores a search list in which data indicating a baby's crying is associated with a priority level of "high."
[0173] In another embodiment, the control unit 27 as shown in FIG. 2 detects, as a sound that satisfies a set condition, a sound whose characteristics match those of a sound that has been set in advance.
[0174] In the same or similar manner as in the above-described embodiment, the speech recognition unit 29 acquires information on the start time of the utterance, information on the end time of the utterance, the utterance ID, and speech sampling data from the section detection unit 28. In another embodiment, the speech recognition unit 29 determines whether or not the features of the speech in the utterance section match the features of the speech set in advance by speech recognition processing using a learning model generated by an arbitrary machine learning algorithm.
[0175] When the speech recognition unit 29 determines that the speech features in the speech section match the predetermined speech features, it outputs a result indicating the match as a speech recognition result, the utterance ID of the speech section, and data indicating the speech features to the event detection unit 30. When the speech recognition unit 29 determines that the speech features in the speech section do not match the predetermined speech features, it outputs a result indicating a mismatch as a speech recognition result and the utterance ID of the speech section to the event detection unit 30.
[0176] The event detection unit 30 may obtain from the speech recognition unit 29 a result indicating a match as a speech recognition result, an utterance ID, and data indicating predetermined speech characteristics. When the event detection unit 30 obtains a result indicating a match, it detects speech whose characteristics match the predetermined speech characteristics as an event. When the event detection unit 30 detects an event, it obtains the utterance ID obtained from the speech recognition unit 29 as an event ID. Furthermore, the event detection unit 30 refers to the search list and obtains a priority corresponding to the data indicating the speech characteristics obtained from the speech recognition unit 29. The event detection unit 30 executes notification processing according to the obtained priority in the same or similar manner as in the above-described embodiment.
[0177] The event detection unit 30 can obtain a result indicating a mismatch as a speech recognition result and an utterance ID from the speech recognition unit 29. When the event detection unit 30 obtains a result indicating a mismatch, the event detection unit 30 obtains the utterance ID obtained from the speech recognition unit 29 as a clear event ID. The event detection unit 30 outputs the clear event ID to the utterance accumulation unit 32.
[0178] The processing of the sound processing device 20 according to other embodiments is not limited to the processing described above. As another example, the control unit 27 may construct a classifier capable of classifying multiple types of sound. Furthermore, the control unit 27 may input sound data collected by the sound collector 110 into the constructed classifier, and based on the result, determine to which priority the sound collected by the sound collector 110 corresponds.
[0179] Other effects and configurations of the voice processing system 101 according to the other embodiment are the same as or similar to those of the voice processing system 1 shown in FIG.
[0180] While the present disclosure has been described based on various drawings and examples, it should be noted that those skilled in the art would easily be able to make various modifications and alterations based on the present disclosure. Therefore, it should be noted that these modifications and alterations are within the scope of the present disclosure. For example, the functions contained in each functional unit can be rearranged so as not to cause logical inconsistencies. Multiple functional units can be combined into one or separated. The above-described embodiments of the present disclosure are not limited to faithful implementation of each of the described embodiments, but can be implemented by combining features or omitting some features as appropriate. In other words, those skilled in the art can make various modifications and alterations based on the present disclosure. Therefore, these modifications and alterations are within the scope of the present disclosure. For example, in each embodiment, each functional unit, means, step, etc. can be added to other embodiments so as not to cause logical inconsistencies, or can be replaced with each functional unit, means, step, etc. of other embodiments. Furthermore, in each embodiment, multiple functional units, means, steps, etc. can be combined into one or separated. Furthermore, each of the above-described embodiments of the present disclosure is not limited to being implemented faithfully according to each of the described embodiments, but can also be implemented by combining each feature or omitting some of them as appropriate.
[0181] For example, the control unit 27 of the voice processing device 20 may detect, from one speech segment, multiple speeches that satisfy different setting conditions. In this case, the control unit 27 may notify the user that speech that satisfies the setting conditions has been detected, according to each of multiple notification conditions set for each of the different setting conditions. Alternatively, the control unit 27 may notify the user that speech that satisfies some of the multiple notification conditions set for each of the different setting conditions has been detected. Some of the multiple notification conditions may be notification conditions that satisfy a selection condition. The selection condition is a condition that is pre-selected from a first condition, a second condition, a third condition, and a fourth condition based on a user operation or the like. Alternatively, some of the multiple notification conditions may be notification conditions included in the first to Nth (N is an integer equal to or greater than 1) notification conditions, counting from the highest priority for notifying the user, as determined by the multiple notification conditions. The Nth priority may be set in advance based on a user operation or the like. For example, the control unit 27 may detect multiple different search words from one speech segment. In this case, the control unit 27 may execute a process according to each of the plurality of priorities set for each of the plurality of different search words. Alternatively, the control unit 27 may execute a process according to a part of the plurality of priorities set for each of the plurality of different search words. The part of the plurality of priorities may be, for example, the priorities included in the top N priorities counting from the highest priority.
[0182] For example, the control unit 27 of the voice processing device 20 may detect voice satisfying the same set condition multiple times from one speech section. In this case, the control unit 27 may execute the process of notifying the user according to the notification condition only once in one speech section, or may execute the process as many times as the voice satisfying the set condition is detected. For example, the control unit 27 may detect the same search word multiple times from one speech section. In this case, the control unit 27 may execute the process according to the priority only once in one speech section, or may execute the process as many times as the search word is detected.
[0183] For example, the section detection unit 28 shown in FIG. 2 may stop detecting speech sections while the replay flag is set to True.
[0184] For example, in the voice processing system 101 shown in Fig. 14, the voice characteristics preset as the setting conditions are described as being the characteristics of a baby's cry. However, the voice characteristics preset as the setting conditions are not limited to the characteristics of a baby's cry. Any voice characteristics may be set as the setting conditions depending on the usage situation of the voice processing system 101. As another example, the characteristics of a boss's voice, an intercom ringtone, or a telephone ringtone may be set as the setting conditions.
[0185] For example, in the above-described embodiment, the priority is described as being set in three stages including "high," "medium," and "low." However, the priority is not limited to being set in three stages. The priority may be set in multiple stages. For example, the priority may be set in two stages or multiple stages of four or more stages.
[0186] For example, in the above-described embodiment, the sound collector 10 and the sound processing device 20 have been described as separate devices. However, the sound collector 10 and the sound processing device 20 may be configured as a single device. An example of this will be described with reference to FIG. 15. A sound processing system 201 as shown in FIG. 15 includes a sound collector 210. The sound collector 210 is an earphone. The sound collector 210 is configured to execute the processing of the sound processing device 20. In other words, the sound collector 210 as an earphone constitutes the sound processing device of the present disclosure. The sound collector 210 includes a microphone 11, a speaker 12, a communication unit 13, a storage unit 14, and a control unit 15 as shown in FIG. 2. The control unit 15 of the sound collector 210 includes components corresponding to the control unit 27 of the sound processing device 20. The storage unit 14 of the sound collector 210 stores a notification list. The sound collector 210 may use another terminal device such as the user's smartphone to perform the screen display and light emission as notification means as shown in Fig. 5. For example, the control unit 15 of the sound collector 210 transmits the notification list in the memory unit 14 to the user's smartphone or the like via the communication unit 13, and displays it on the user's smartphone or the like.
[0187] For example, in the above-described embodiment, the voice processing device 20 has been described as performing the voice recognition processing. However, an external device other than the voice processing device 20 may perform the voice recognition processing. The control unit 27 of the voice processing device 20 may acquire the results of the voice recognition processing performed by the external device. The external device may be, for example, a dedicated computer configured to function as a server, a general-purpose personal computer, or a cloud computing system. In this case, the communication unit 13 of the sound collector 10 may be configured to further include at least one communication module that is the same as or similar to the communication unit 21 and can connect to any network including a mobile communication network and the Internet. In the sound collector 10, the control unit 15 may transmit voice sampling data to the external device via the network using the communication unit 13. Upon receiving the voice sampling data from the sound collector 10 via the network, the external device may perform the voice recognition processing. The external device may transmit the results of the voice recognition processing to the voice processing device 20 via the network. In the voice processing device 20, the control unit 27 may acquire the results of the voice recognition processing by receiving them from the external device via the network using the communication unit 21.
[0188] For example, in the above-described embodiment, the sound processing device 20 has been described as a terminal device. However, the sound processing device 20 is not limited to a terminal device. As another example, the sound processing device 20 may be a dedicated computer configured to function as a server, a general-purpose personal computer, a cloud computing system, or the like. In this case, the communication unit 13 of the sound collector 10 may be configured to further include at least one communication module that is the same as or similar to the communication unit 21 and can connect to any network including a mobile communication network and the Internet. The sound collector 10 and the sound processing device 20 may communicate via a network.
[0189] For example, an embodiment is also possible in which a general-purpose computer functions as the voice processing device 20 according to the above-described embodiment. Specifically, a program describing the processing content for realizing each function of the voice processing device 20 according to the above-described embodiment is stored in the memory of the general-purpose computer, and the program is read and executed by a processor. Therefore, the configuration according to the above-described embodiment can also be realized as a program executable by a processor or a non-transitory computer-readable medium storing the program. [Explanation of symbols]
[0190] 1,101,201 Voice Processing System 10,110,210 Sound collector 11. Mike 12 speakers 13 Communications Department 14 Storage section 15 Control Unit 16 Voice acquisition unit 17 Audio playback section 18 Storage Unit 20 Audio processing device 21 Communications Department 22 Input section 23 Display section 24 Vibration unit 25 Light-emitting part 26 Memory section 27 Control Unit 28 Section detection unit 29 Voice Recognition Unit 30 Event detection unit 31 Speech notification unit 32 Speech storage unit 33 Data Buffer 34 Ring Buffer 35 Audio modulation section 40 Main Screen 41,42,43,44 area 50 Settings screen 51,52,53,54,55,56 area 60 Notification screen 61,62,63,64 area
Claims
1. The storage unit stores sounds collected by the microphone around the user; When the control unit detects an utterance that satisfies a set condition in the surrounding sound, automatically outputting the ambient sound from a start point of an utterance section including the utterance when a notification condition corresponding to the utterance satisfies a first condition; outputting the ambient sound including the utterance based on an input from the user when a notification condition corresponding to the utterance satisfies a second condition. Sound processing methods.
2. a storage unit that stores sounds collected by a microphone around the user; a control unit; When an utterance satisfying a set condition is detected in the ambient sound, the control unit automatically outputting the ambient sound from a start point of an utterance section including the utterance when a notification condition corresponding to the utterance satisfies a first condition; outputting the ambient sound including the utterance based on an input from the user when a notification condition corresponding to the utterance satisfies a second condition; Sound processing device.
3. storing sounds around the user collected by a microphone in a storage unit; When an utterance that satisfies a set condition is detected in the surrounding sound, automatically outputting the ambient sound from a start point of an utterance section including the utterance when a notification condition corresponding to the utterance satisfies a first condition; outputting the ambient sound including the utterance based on an input from the user when a notification condition corresponding to the utterance satisfies a second condition; causing the control unit to execute Sound processing program.
4. The storage unit stores sounds collected by the microphone around the user; When the control unit detects a sound that satisfies a set condition in the surrounding sound, When a notification condition corresponding to the detected sound satisfies a first condition, the ambient sound is automatically output from a starting point of the ambient sound including the detected sound; When a notification condition corresponding to the detected sound satisfies a second condition, outputting the ambient sound including the detected sound based on an input from the user. Sound processing methods.
Citation Information
Patent Citations
Portable music reproducing device
JP2001256771A
Mobile apparatus
JP2007243493A
Information processing program, information processing method and program
JP2017069687A
Voice processor, voice processing method and voice processing program
JP2020156107A
Communication system
JP2021136668A