Speech processing device, speech processing method, and speech processing system

The audio processing device addresses inconsistent user notifications by employing prioritized audio detection and tailored notification methods, improving user experience through selective and timely alerts.

JP7894888B2Active Publication Date: 2026-07-24KYOCERA CORP
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
KYOCERA CORP
Filing Date
2023-01-10
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Conventional systems fail to provide user-preferred notification methods based on the content of detected voice, leading to inconsistent user experiences.

Method used

An audio processing device that selectively notifies users about detected audio based on pre-set conditions, using various notification methods such as sound, vibration, and visual cues, prioritized by the importance of the detected content.

Benefits of technology

Enhances user convenience by allowing tailored notifications based on audio content, ensuring immediate attention to high-priority events and enabling manual control over lower-priority alerts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007894888000001
    Figure 0007894888000001
  • Figure 0007894888000002
    Figure 0007894888000002
  • Figure 0007894888000003
    Figure 0007894888000003
Patent Text Reader

Abstract

This audio processing device is provided with a control unit. The control unit acquires the result of an audio recognition process for recognizing audio with respect to audio data. When the control unit detects, on the basis of the result of the audio recognition process, audio satisfying a setting condition set in advance in relation to audio, the control unit notifies a user that audio satisfying the setting condition has been detected, in accordance with a notification condition corresponding to the detected audio set in the setting condition.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-reference to related applications , , , ,

[0005] , , , ,

[0006] , , ,

[0004] , , , , , , , , , <00000​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​The system includes, when it detects audio that satisfies pre-set conditions based on the results of the speech recognition process, notifying the user that audio satisfying the conditions has been detected, according to the notification conditions corresponding to the detected audio set in the conditions.

[0007] A speech processing system according to one embodiment of this disclosure is A sound collector that gathers ambient sounds, The audio processing device includes: an audio processing device that obtains the result of an audio recognition process that recognizes audio from audio data collected by the sound collector, and when it detects an audio that satisfies pre-set setting conditions based on the result of the audio recognition process, it notifies the user that an audio that satisfies the setting conditions has been detected, according to the notification conditions corresponding to the detected audio set in the setting conditions. [Brief explanation of the drawing]

[0008] [Figure 1] This figure shows a schematic configuration of a speech processing system according to one embodiment of the present disclosure. [Figure 2] Figure 1 is a block diagram of the speech processing system shown. [Figure 3] This figure shows an example of a search list. [Figure 4] This figure shows an example of a notification sound list. [Figure 5] This diagram illustrates the notification methods and timings according to priority. [Figure 6] This figure shows an example of the main screen. [Figure 7] This figure shows an example of a settings screen. [Figure 8] This figure shows an example of a notification screen. [Figure 9] Figure 2 illustrates an example of the processing performed by the interval detection unit. [Figure 10] Figure 2 is a block diagram of the speech storage unit. [Figure 11] Figure 2 is a flowchart showing the operation of the event detection process performed by the audio processing device. [Figure 12] It is a flowchart showing the operation of the output process of playback data executed by the audio processing device shown in FIG. 2. [Figure 13] It is a flowchart showing the operation of the output process of playback data executed by the audio processing device shown in FIG. 2. [Figure 14] It is a diagram showing a schematic configuration of an audio processing system according to another embodiment of the present disclosure. [Figure 15] It is a diagram showing a schematic configuration of an audio processing system according to still another embodiment of the present disclosure.

Mode for Carrying Out the Invention

[0009] There is room for improvement in the conventional technology. For example, there are cases where the user wants to be preferentially notified that the voice has been detected depending on the content of the detected voice, and cases where the user does not want to be preferentially notified. According to an embodiment of the present disclosure, an improved audio processing device, audio processing method, and audio processing system can be provided.

[0010] In the present disclosure, "voice" includes any sound. For example, the voice includes a voice uttered by a person, a sound output by a machine, a cry uttered by an animal, and environmental sounds.

[0011] Hereinafter, embodiments according to the present disclosure will be described with reference to the drawings.

[0012] As shown in FIG. 1, the audio processing system 1 includes a microphone 10 and an audio processing device 20. The microphone 10 and the audio processing device 20 can communicate with each other via a communication line. The communication line includes at least one of wired and wireless.

[0013] In the present embodiment, the microphone 10 is an earphone. However, the microphone 10 is not limited to an earphone. The microphone 10 may be a headset or the like. The microphone 10 is worn by the user. The microphone 10 can output music or the like. The microphone 10 may include an earphone unit worn on the left ear of the user and an earphone unit worn on the right side of the user.

[0014] The sound collector 10 collects the sound around the sound collector 10. By being worn by the user, the sound collector 10 collects the sound around the user. Based on the control of the sound processing device 20, the sound collector 10 outputs the collected sound around the user. With such a configuration, the user can hear the sound around himself / herself while wearing the sound collector 10.

[0015] In this embodiment, the sound processing device 20 is a terminal device. The terminal device serving as the sound processing device 20 is, for example, a mobile phone, a smartphone, a tablet, or a personal computer (PC), etc. However, the sound processing device 20 is not limited to a terminal device.

[0016] The sound processing device 20 is operated by the user. The user can operate the sound processing device 20 to set the sound collector 10 and the like.

[0017] The sound processing device 20 controls the sound collector 10 to collect the sound around the user. When the sound processing device 20 detects a sound that satisfies a preset setting condition from the collected sound around the user, it notifies the user that a sound that satisfies the setting condition has been detected. The details of this process will be described later.

[0018] FIG. 2 is a block diagram of the sound processing system 1 shown in FIG. 1. In FIG. 2, the main flow of data and the like is shown by a solid line.

[0019] The sound collector 10 includes a microphone 11, a speaker 12, a communication unit 13, a storage unit 14, and a control unit 15.

[0020] The microphone

[0021] Speaker 12 is capable of outputting sound. Speaker 12 includes a left speaker and a right speaker. The left speaker may be included in the earphone unit attached to the user's left ear, which is included in the sound collector 10. The right speaker may be included in the earphone unit attached to the user's right side, which is included in the sound collector 10. For example, speaker 12 may be a stereo speaker.

[0022] The communication unit 13 comprises at least one communication module capable of communicating with the voice processing device 20 via a communication line. The communication module is a communication module that conforms to the communication line standard. The communication line standard is, for example, a wired communication standard or a short-range wireless communication standard including Bluetooth®, infrared and NFC (Near Field Communication).

[0023] The storage unit 14 is configured to include at least one semiconductor memory, at least one magnetic memory, at least one optical memory, or at least two combinations thereof. The semiconductor memory is, for example, RAM (Random Access Memory) or ROM (Read Only Memory). The RAM is, for example, SRAM (Static Random Access Memory) or DRAM (Dynamic Random Access Memory). The ROM is, for example, EEPROM (Electrically Erasable Programmable Read Only Memory). The storage unit 14 may function as a main memory, auxiliary memory, or cache memory. The storage unit 14 stores data used for the operation of the sound collector 10 and data obtained by the operation of the sound collector 10. For example, the storage unit 14 stores system programs, application programs, and embedded software.

[0024] The control unit 15 is comprised of at least one processor, at least one dedicated circuit, or a combination thereof. The processor is a general-purpose processor such as a CPU (Central Processing Unit) or GPU (Graphics Processing Unit), or a dedicated processor specialized for a specific process. The dedicated circuit is, for example, an FPGA (Field-Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). The control unit 15 controls each part of the sound collector 10 and executes processes related to the operation of the sound collector 10.

[0025] In this embodiment, the control unit 15 includes an audio acquisition unit 16, an audio playback unit 17, and a storage unit 18. The storage unit 18 is configured to include the same or similar components as the memory unit 14. At least a part of the storage unit 18 may be a part of the memory unit 14. The operation of the storage unit 18 is performed by the processor or the like of the control unit 15.

[0026] The audio acquisition unit 16 acquires digital audio data from the analog audio data collected by the microphone 11. In this embodiment, the audio acquisition unit 16 acquires audio sampling data as digital audio data by sampling the analog audio data at a preset sampling rate.

[0027] The audio acquisition unit 16 outputs the audio sampling data to the audio playback unit 17. The audio acquisition unit 16 also transmits the audio sampling data to the audio processing device 20 via the communication unit 13.

[0028] If the microphone 11 includes a left microphone and a right microphone, the audio acquisition unit 16 may acquire left-side audio sampling data from the analog data of the audio collected by the left microphone. The audio acquisition unit 16 may also acquire right-side audio sampling data from the analog data of the audio collected by the right microphone. The audio acquisition unit 16 may transmit the left-side audio sampling data and the right-side audio sampling data to the audio processing device 20 via the communication unit 13. Hereinafter, when there is no particular distinction between left-side audio sampling data and right-side audio sampling data, they will simply be referred to as "audio sampling data".

[0029] The audio playback unit 17 acquires audio sampling data from the audio acquisition unit 16. The audio playback unit 17 receives a replay flag from the audio processing device 20 via the communication unit 13.

[0030] The replay flag is set to either True or False. When the replay flag is False, the audio processing system 1 operates in pass-through mode. Pass-through mode is a mode in which the audio data collected by the sound collector 10 is output from the sound collector 10 without going through the audio processing device 20. When the replay flag is True, the audio processing system 1 operates in playback mode. Playback mode is a mode in which the sound collector 10 outputs the playback data acquired from the audio processing device 20. The conditions under which the replay flag is set to True or False will be described later.

[0031] When the replay flag is False, i.e., when the audio processing system 1 is in pass-through mode, the audio playback unit 17 outputs the audio sampling data acquired from the audio acquisition unit 16 to the speaker 12.

[0032] When the replay flag is True, i.e., when the audio processing system 1 is in playback mode, the audio playback unit 17 outputs the playback data stored in the storage unit 18 to the speaker 12.

[0033] The audio playback unit 17 receives a notification sound file from the audio processing unit 20 via the communication unit 13. The notification sound file is transmitted from the audio processing unit 20 to the audio playback unit 17 when the audio processing unit 20 detects an audio that meets the set conditions. Upon receiving the notification sound file, the audio playback unit 17 outputs a notification sound to the speaker 12. With this configuration, the user can be notified that an audio that meets the set conditions has been detected.

[0034] The storage unit 18 stores playback data. Playback data is data transmitted from the audio processing device 20 to the sound collector 10. When the control unit 15 receives playback data from the audio processing device 20 via the communication unit 13, it stores the received playback data in the storage unit 18. The control unit 15 can receive playback stop instructions and replay stop instructions from the audio processing device 20 via the communication unit 13. When the control unit 15 receives a playback stop instruction or a replay stop instruction, it erases the playback data stored in the storage unit 18.

[0035] The control unit 15 may receive playback data for the left channel and playback data for the right channel from the audio processing device 20 and store them in the storage unit 18. In this case, the audio playback unit 17 may output the playback data for the left channel stored in the storage unit 18 to the left speaker of speaker 12, and output the playback data for the right channel stored in the storage unit 18 to the right speaker of speaker 12.

[0036] The audio processing device 20 includes a communication unit 21, an input unit 22, a display unit 23, a vibration unit 24, a storage unit 26, and a control unit 27.

[0037] The communication unit 21 comprises at least one communication module capable of communicating with the sound collector 10 via a communication line. The communication module is a communication module that conforms to the communication line standard. The communication line standard is, for example, a wired communication standard or a short-range wireless communication standard including Bluetooth®, infrared, and NFC.

[0038] The communication unit 21 may further include at least one communication module capable of connecting to any network, including mobile communication networks and the Internet. The communication module is, for example, a communication module compatible with mobile communication standards such as LTE (Long Term Evolution), 4G (4th Generation), or 5G (5th Generation).

[0039] The input unit 22 is capable of receiving input from the user. The input unit 22 is configured to include at least one input interface capable of receiving input from the user. The input interface may be, for example, a physical key, a capacitive key, a pointing device, a touchscreen integrated with a display, or a microphone.

[0040] The display unit 23 is capable of displaying data. The display unit 23 is, for example, a display. The display is, for example, an LCD (Liquid Crystal Display) or an organic EL (Electro Luminescence) display.

[0041] The vibration unit 24 is capable of vibrating the sound processing device 20. The vibration unit 24 is composed of a vibration element, which is, for example, a piezoelectric element.

[0042] The light-emitting part 25 is capable of emitting light. The light-emitting part 25 is, for example, an LED (Light Emitting Diode).

[0043] The storage unit 26 is configured to include at least one semiconductor memory, at least one magnetic memory, at least one optical memory, or at least two combinations thereof. The semiconductor memory is, for example, RAM or ROM. The RAM is, for example, SRAM or DRAM. The ROM is, for example, EEPROM. The storage unit 26 may function as a main memory, auxiliary memory, or cache memory. The storage unit 26 stores data used for the operation of the voice processing device 20 and data obtained by the operation of the voice processing device 20. For example, the storage unit 26 stores system programs, application programs, and embedded software.

[0044] The memory unit 26 stores, for example, a search list as shown in Figure 3 below, and a notification sound list and notification sound files as shown in Figure 4 below. The memory unit 26 stores, for example, the notification list below.

[0045] The control unit 27 is configured to include at least one processor, at least one dedicated circuit, or a combination thereof. The processor is a general-purpose processor such as a CPU or GPU, or a dedicated processor specialized for a specific process. The dedicated circuit is, for example, an FPGA or ASIC. The control unit 27 controls each part of the voice processing device 20 and executes processes related to the operation of the voice processing device 20.

[0046] The control unit 27 performs speech recognition processing to recognize speech in the audio data. However, the control unit 27 may also obtain the results of speech recognition processing performed by an external device. If the control unit 27 detects speech that satisfies the set conditions based on the results of the speech recognition processing, it notifies the user that speech that satisfies the set conditions has been detected. The set conditions are conditions that have been set in advance regarding speech. The control unit 27 notifies the user that speech that satisfies the set conditions has been detected according to the notification conditions corresponding to the detected speech set in the set conditions.

[0047] The notification conditions are criteria for determining the priority of notifying the user that sound has been detected. The higher the priority, the earlier the user should be notified. The higher the priority, the more easily the user can be notified using a notification method that is easy for the user to notice. For example, as mentioned above, the user is wearing earphones, which are sound collectors 10. Therefore, if sound such as a notification sound is output from sound collectors 10, the user can immediately notice the sound. In other words, a notification method using sound, such as a notification sound, should have a higher priority than a notification method using the presentation of visual information, etc. As will be described later, the user may be notified that sound has been detected by playing the detected sound. In this case, the higher the priority, the earlier the detected sound should be played. Also, if the priority is low, the detected sound may be played at any time.

[0048] The notification conditions include the first and second conditions. If the notification conditions meet the second condition, the priority for notifying the user is lower than if the notification conditions meet the first condition. The first condition includes the third and fourth conditions. If the notification conditions meet the fourth condition, the priority for notifying the user is lower than if the notification conditions meet the third condition.

[0049] In this embodiment, notification conditions are set by priority. Priority indicates the priority order in which audio that meets the set conditions is notified to the user. Priority may be set in multiple levels. The higher the priority, the higher the priority for notification to the user. In this embodiment, as shown in Figures 3 and 4, priority is set in three levels, including "high," "medium," and "low." "High" is the highest priority among the three priority levels. "Medium" is the middle priority among the three priority levels. "Low" is the lowest priority among the three priority levels.

[0050] In this embodiment, the notification condition is satisfied if the first condition is met, meaning that the priority is "high" or "medium". The notification condition is satisfied if the second condition is met, meaning that the priority is "low". The notification condition is satisfied if the third condition is met, meaning that the priority is "high". The notification condition is satisfied if the fourth condition is met, meaning that the priority is "medium".

[0051] The control unit 27 may notify the user that a sound meeting the set conditions has been detected by playing a notification sound if the notification conditions corresponding to the detected sound satisfy the first condition, that is, if the priority corresponding to the detected sound is "medium" or "high". With this configuration, the user can immediately notice that a sound has been detected.

[0052] The control unit 27 may play a notification sound followed by an audio that satisfies the set conditions if the notification condition corresponding to the detected audio satisfies the first condition, i.e., if the priority corresponding to the detected audio is "medium" or "high". With this configuration, if the priority is "medium" or "high", the detected audio is automatically played. When the priority is "medium" or "high", the user is likely to want to immediately check the content of the detected audio. By automatically playing an audio that satisfies the set conditions when the notification condition satisfies the first condition, the user can immediately check the detected audio. Therefore, user convenience can be improved.

[0053] The control unit 27 may notify the user that an audio meeting the set conditions has been detected by presenting visual information if the notification condition corresponding to the detected audio satisfies the second condition, that is, if the priority corresponding to the detected audio is "low". When the priority is "low", instead of playing a notification sound, the control unit 27 can provide a notification appropriate to the low priority by presenting visual information to the user.

[0054] In this embodiment, the setting condition is that the utterance includes a pre-set search word. Each search word is assigned a priority. The search word consists of, for example, at least one of letters and numbers. The search word may be any information as long as it can be processed as text data. In this embodiment, when the control unit 27 detects an utterance that includes a search word as an audio that satisfies the setting condition, it notifies the user that the utterance has been detected.

[0055] Figure 3 shows the search list. The search list associates search words with the priority assigned to them. The search word "Flight 153" is assigned a priority of "High". The search word "Hello" is assigned a priority of "Medium". The search word "Good morning" is assigned a priority of "Low". For example, the control unit 27 generates the search list based on user input to the settings screen 50, as shown in Figure 7, which will be described later.

[0056] Figure 4 shows the notification sound list. The notification sound list associates priority levels with the notification sounds assigned to those priorities. Notification sounds are used to notify the user that speech has been detected. In this embodiment, notification sounds are also used when speech corresponding to a "low" priority level is detected. However, notification sounds do not have to be used when speech corresponding to a "low" priority level is detected. In Figure 4, priority levels are associated with notification sound files. Notification sound files are files used to store notification sounds on the computer. The notification sound file "ring.wav" is associated with a "high" priority level. The notification sound file "alert.wav" is associated with a "medium" priority level. The notification sound file "notify.wav" is associated with a "low" priority level.

[0057] The notification according to priority according to this embodiment will be explained with reference to Figure 5. In Figure 5, the notification means is a means for notifying the user that an utterance containing a search word has been detected. The notification timing is the timing at which the user is notified that an utterance containing a search word has been detected. The playback timing is the timing at which the detected utterance is played back.

[0058] As shown in Figure 5, when the control unit 27 detects an utterance corresponding to the priority of "High," it uses a notification sound and vibration by the vibration unit 24 as notification means. The control unit 27 sets the notification timing to the moment immediately after the search word is detected. The control unit 27 sets the playback timing to the moment immediately after notifying that an utterance has been detected. In other words, when the notification condition satisfies the third condition, the control unit 27 controls the system so that the notification sound is played and the playback of the utterance begins immediately after the search word is detected. For example, suppose the search word "Flight 153" is set to the priority of "High." In this case, the control unit 27 sets the notification timing to the moment immediately after the search word "Flight 153" is detected. In other words, the control unit 27 vibrates the vibration unit 24 and plays the notification sound immediately after the search word "Flight 153" is detected. Furthermore, the control unit 27 sets the playback timing to the moment immediately after the notification sound is played so that the utterance "Flight 153 is scheduled to depart 20 minutes late" is played. With this configuration, as soon as the search term "Flight 153" is detected, a notification sound is played, and the announcement "Flight 153 is scheduled to depart 20 minutes late" begins to play. Additionally, announcements containing the search term are played automatically.

[0059] As shown in Figure 5, when the control unit 27 detects an utterance corresponding to the priority level of "medium," it uses a notification sound and vibration by the vibration unit 24 as notification means. The control unit 27 sets the notification timing to the moment immediately after the utterance containing the search word has finished. The control unit 27 sets the playback timing to the moment immediately after notifying that an utterance has been detected. In other words, when the notification condition satisfies the fourth condition, the control unit 27 controls the system so that the notification sound is played and the playback of the utterance begins immediately after the utterance containing the search word has finished. For example, suppose the search word "Flight 153" is set to the priority level of "medium." In this case, the control unit 27 sets the notification timing to the moment immediately after the utterance "Flight 153 is scheduled to depart 20 minutes late" has finished, vibrates the vibration unit 24, and plays the notification sound. The control unit 27 also sets the playback timing to the moment immediately after the notification sound is played, and controls the system so that the utterance "Flight 153 is scheduled to depart 20 minutes late" is played. With this configuration, immediately after the statement "Flight 153 is scheduled to depart 20 minutes late" finishes, a notification sound is played, and the statement "Flight 153 is scheduled to depart 20 minutes late" begins playing again. In addition, statements containing search terms are played automatically.

[0060] As shown in Figure 5, when the control unit 27 detects an utterance corresponding to a priority of "low," it uses screen display by the display unit 23 and light emission by the light emission unit 25 as notification means. Screen display and light emission are examples of notification means that present visual information to the user. The control unit 27 displays a notification list on the display unit 23 as a screen display. The notification list is a list of audio information that satisfies the detected setting conditions. In this embodiment, the notification list is a list of event information. An event is an utterance that contains a search word. Details of the notification list will be described later. The control unit 27 sets the playback timing to immediately after the user instructs playback of the utterance. In other words, if the notification condition satisfies the second condition, the control unit 27 plays back the utterance containing the search word based on the user's input. With this configuration, the utterance containing the search word is played back manually.

[0061] <Input / Output Processing> The control unit 27 receives user input via the input unit 22. Based on the input received by the input unit 22, the control unit 27 selects a screen to display on the display unit 23. For example, based on the input received by the input unit 22, the control unit 27 displays a screen as shown in Figures 6, 7, or 8. In the configurations shown in Figures 6 to 8, the input unit 22 is a touchscreen integrated with the display of the display unit 23.

[0062] The main screen 40, as shown in Figure 6, includes areas 41, 42, 43, and 44.

[0063] Area 41 displays the status of the sound collector 10. In Figure 6, area 41 displays the information "Replaying…" indicating that the sound collector 10 is in the process of replaying.

[0064] If the audio processing system 1 is in pass-through mode, the words "Start Replay" are displayed in area 42. If the audio processing system 1 is in playback mode, the words "Stop Replay" are displayed in area 42. The control unit 27 can receive input to area 42 via the input unit 22.

[0065] When the words "Start Replay" are displayed in area 42, that is, when the voice processing system 1 is in pass-through mode, the control unit 27 can accept the start of replay by receiving input to area 42 via the input unit 22. When the control unit 27 accepts the start of replay, it sets the replay flag to True and outputs a replay instruction to the speech storage unit 32, which will be described later.

[0066] When the words "Stop Replay" are displayed in area 42, that is, when the audio processing system 1 is in playback mode, the control unit 27 can accept the replay stop request by receiving input to area 42 via the input unit 22. When the control unit 27 accepts the replay stop request, it sets the replay flag to False and transmits a replay stop instruction to the sound collector 10 via the communication unit 21.

[0067] The words "Notification List" are displayed in area 43. The control unit 27 can receive input for area 43 via the input unit 22. When the control unit 27 receives input for area 43 via the input unit 22, it displays a notification screen 60, as shown in Figure 8, on the display unit 23.

[0068] The word "Settings" is displayed in area 44. The control unit 27 can receive input for area 44 via the input unit 22. When the control unit 27 receives input for area 44 via the input unit 22, it displays a settings screen 50, as shown in Figure 7, on the display unit 23.

[0069] The settings screen 50 shown in Figure 7 is a screen for the user to make various settings. The settings screen 50 includes areas 51, 52, 53, 54, 55, and 56.

[0070] The words "Add Search Term" are displayed in area 51. The control unit 27 can receive input for area 51 via the input unit 22. The control unit 27 receives the input of a search term and the corresponding priority from area 51.

[0071] Area 52 displays the set search words. In Figure 7, the search words "Flight 153", "Hello", and "Good morning" are displayed in Area 52. The control unit 27 can receive input for Area 52 via the input unit 22. When the control unit 27 receives input for Area 52 via the input unit 22, it displays a search list on the display unit 23 as shown in Figure 3.

[0072] Area 53 displays the words "Recording Buffer Settings". Area 53 is used to set the length of the recording time for recording the sound collected by the sound collector 10. In this embodiment, the audio sampling data for the recording time is stored in a ring buffer 34 as shown in Figure 10, which will be described later. The control unit 27 can receive input to area 53 via the input unit 22. The control unit 27 accepts input of recording times such as 5 seconds, 10 seconds, and 15 seconds. The control unit 27 stores the received recording time information in the storage unit 26.

[0073] Area 54 displays the words "Speed ​​Setting". Area 54 is used to set the playback speed of the sound output from the sound collector 10. The control unit 27 can receive input to area 54 via the input unit 22. The control unit 27 accepts input of sound speeds such as 1x speed, 1.1x speed, and 1.2x speed. The control unit 27 stores the received sound speed information in the storage unit 26.

[0074] Area 55 displays the words "Speech Threshold Setting". Area 55 is used to set the speech threshold that will be cut as noise from the sound collected by the sound collector 10. In this embodiment, speech below the speech threshold is cut as noise. The control unit 27 can receive input to area 55 via the input unit 22. The control unit 27 accepts input of speech thresholds from, for example, -50 [dBA] to -5 [dBA]. The control unit 27 stores the received speech threshold information in the storage unit 26.

[0075] The words "Setting complete" are displayed in area 56. The control unit 27 can receive input for area 56 via the input unit 22. When the control unit 27 receives input for area 56 via the input unit 22, it displays the main screen 40 shown in Figure 6 on the display unit 23.

[0076] The notification screen 60 shown in Figure 8 is a screen for notifying the user of various information. The notification screen 60 includes areas 61, 62, 63, and 64.

[0077] Area 61 displays the notification list. As described above, the notification list is a list of event information. As described above, an event is an utterance containing a search word. The control unit 27 displays in area 61 the event information with a "low" priority from among the events included in the notification list. However, the control unit 27 may display all event information included in the notification list in area 61 regardless of priority. The control unit 27 can receive input for each event in the notification list displayed in area 61 via the input unit 22. The control unit 27 accepts the selection of an event in the notification list by receiving input for each event in the notification list from area 61 via the input unit 22.

[0078] Area 62 displays the words "Detailed Display". The control unit 27 can receive input for area 62 via the input unit 22. The control unit 27 can receive the selection of an event included in the notification list from area 61, and can further receive input for area 62 via the input unit 22. In this case, the control unit 27 displays the details of the event information selected from area 61 in area 61 via the display unit 23. For example, the control unit 27 displays the left-hand speech recognition result and the right-hand speech recognition result, which will be described later, as details of the event information.

[0079] Area 63 displays the words "Start Playback / Stop Playback". The control unit 27 can accept a start or stop of playback by receiving input to area 63 via the input unit 22. When speech is not being played back, the control unit 27 accepts the selection of an event included in the notification list from area 61, and further accepts the start of event playback by receiving input to area 63 via the input unit 22. When the start of event playback is accepted, the control unit 27 controls the system so that the selected event, i.e., speech, from area 61 is played back. In this embodiment, the control unit 27 refers to the notification list in the storage unit 26 and obtains the event ID of the selected event from area 61, as described below. The control unit 27 outputs the event ID and a start playback instruction to the speech holding unit 36, as described below, and controls the system so that the speech is played back. Also, when speech is being played back, the control unit 27 accepts a stop of event playback by receiving input to area 63 via the input unit 22. When the control unit 27 receives a request to stop playback of an event, it controls the playback of the speech to stop. In this embodiment, the communication unit 21 transmits a playback stop instruction to the sound collector 10, and controls the playback of the speech to stop.

[0080] The word "Back" is displayed in area 64. The control unit 27 can receive input for area 64 via the input unit 22. When the control unit 27 receives input for area 64 via the input unit 22, it displays the main screen 40 shown in Figure 6 on the display unit 23.

[0081] <Audio Processing> As shown in Figure 2, the control unit 27 includes a section detection unit 28, a speech recognition unit 29, an event detection unit 30, a speech notification unit 31, a speech storage unit 32, a speech modulation unit 35, and a speech holding unit 36. The speech holding unit 36 ​​is configured to include the same or similar components as the storage unit 26. At least a part of the speech holding unit 36 ​​may be a part of the storage unit 26. The operation of the speech holding unit 36 ​​is performed by the processor of the control unit 27, etc.

[0082] The section detection unit 28 receives audio sampling data from the sound collector 10 via the communication unit 21. The section detection unit 28 detects speech sections from the audio sampling data. A speech section is a section in which the speech state continues. The section detection unit 28 can also detect non-speaking sections by detecting speech sections from the audio sampling data. A non-speaking section is a section in which the non-speaking state continues. The starting point of a speech section is also referred to as the "start of speech." The starting point of a speech section is the ending point of a non-speaking section. The ending point of a speech section is also referred to as the "end of speech." The ending point of a speech section is the starting point of a non-speaking section.

[0083] An example of the processing of the interval detection unit 28 will be described with reference to Figure 9. However, the processing of the interval detection unit 28 is not limited to the processing described with reference to Figure 9. The interval detection unit 28 may detect speech intervals from the speech sampling data by any method. As another example, the interval detection unit 28 may detect speech intervals from the speech sampling data using a machine learning model generated using any machine learning algorithm.

[0084] In Figure 9, the horizontal axis represents time. The audio sampling data shown in Figure 9 is acquired by the audio acquisition unit 16 of the sound collector 10. The section detection unit 28 acquires audio section detection data from the audio sampling data. The audio section detection data is data obtained by averaging the power of the audio sampling data over a preset time width. The time width of the audio section detection data may be set based on the specifications of the audio processing device 20, etc. In Figure 9, one audio section detection data is shown as one rectangle. The time width of this rectangle, that is, the time width of one audio section detection data, is, for example, 200 [ms].

[0085] The interval detection unit 28 obtains audio threshold information from the storage unit 26 and classifies the audio interval detection data into audio data and non-audio data. In Figure 9, audio data is the data with a dark color among the audio interval detection data shown as squares. Non-audio data is the data with a white outline among the audio interval detection data shown as squares. The interval detection unit 28 classifies the audio interval detection data as non-audio data if the value of the audio interval detection data is null. The interval detection unit 28 classifies the audio interval detection data as non-audio data if the value of the audio interval detection data is not null and the value of the audio interval detection data is equal to or greater than the audio threshold. The interval detection unit 28 classifies the audio interval detection data as audio data if the value of the audio interval detection data is not null and the value of the audio interval detection data is equal to or greater than the audio threshold.

[0086] The section detection unit 28 detects a section in which voice data continues without a set time interval as a speech section. The set time may be set based on the language processed by the voice processing device 20. If the language processed is Japanese, for example, the set time is 500 [ms]. In Figure 9, when the section detection unit 28 detects voice data after non-voice data has continued for longer than the set time, it identifies the time when the voice data was detected as the start of speech. For example, the section detection unit 28 identifies time t1 as the start of speech. After identifying the start of speech, if the section detection unit 28 determines that non-voice data has continued for longer than the set time, it identifies the time when that determination was made as the end of speech. For example, the section detection unit 28 identifies time t2 as the end of speech. The section detection unit 28 detects the section from the start of speech to the end of speech as a speech section.

[0087] The section detection unit 28 may receive left-channel audio sampling data and right-channel audio sampling data from the sound collector 10. In this case, the section detection unit 28 may identify the time when audio data is detected in either the left-channel or right-channel audio sampling data after non-audio data has continued for a set time in both the left-channel and right-channel audio sampling data for longer than the set time, as the start time of speech. Alternatively, if the section detection unit 28 determines that non-audio data has continued for longer than the set time in both the left-channel and right-channel audio sampling data, it may identify the time when this determination is made as the end time of speech.

[0088] The section detection unit 28 identifies the start time of speech from the speech sampling data and generates an utterance ID. Each utterance ID is a uniquely identifiable piece of identification information. The section detection unit 28 outputs the information about the start time of speech and the utterance ID to the speech recognition unit 29 and the utterance storage unit 32, respectively.

[0089] When the section detection unit 28 identifies the end of speech from the speech sampling data, it outputs information about the end of speech to the speech recognition unit 29 and the speech storage unit 32, respectively.

[0090] The section detection unit 28 sequentially outputs the audio sampling data received from the sound collector 10 to the speech recognition unit 29 and the speech storage unit 32, respectively.

[0091] The speech recognition unit 29 obtains information about the start of speech and an utterance ID from the section detection unit 28. Once the speech recognition unit 29 obtains the information about the start of speech, it performs speech recognition processing to recognize speech on the speech sampling data that is sequentially obtained from the section detection unit 28. In this embodiment, the speech recognition unit 29 recognizes speech by converting the speech data contained in the speech sampling data into text data through speech recognition processing.

[0092] The speech recognition unit 29 outputs the speech start time information and speech ID obtained from the section detection unit 28 to the event detection unit 30. After outputting the speech start time information etc. to the event detection unit 30, the speech recognition unit 29 sequentially outputs the text data as the speech recognition result to the event detection unit 30.

[0093] The speech recognition unit 29 obtains information about the end of speech from the section detection unit 28. Once the speech recognition unit 29 obtains the information about the end of speech, it terminates the speech recognition process. The speech recognition unit 29 outputs the information about the end of speech obtained from the section detection unit 28 to the event detection unit 30. After this, the speech recognition unit 29 may obtain information about the start of a new speech and an utterance ID from the section detection unit 28. Once the speech recognition unit 29 obtains the information about the start of a new speech, it performs speech recognition processing again on the speech sampling data sequentially obtained from the section detection unit 28.

[0094] The speech recognition unit 29 may obtain left-channel speech sampling data and right-channel speech sampling data from the interval detection unit 28. In this case, the speech recognition unit 29 may convert each of the left-channel speech sampling data and right-channel speech sampling data into text data. Hereinafter, the text data obtained from the left-channel speech sampling data will also be referred to as "left-channel text data" or "left-channel speech recognition result." The text data obtained from the right-channel speech sampling data will also be referred to as "right-channel text data" or "right-channel speech recognition result."

[0095] The event detection unit 30 obtains information about the start of speech and an utterance ID from the speech recognition unit 29. After obtaining the information about the start of speech, the event detection unit 30 sequentially obtains text data from the speech recognition unit 29. The event detection unit 30 refers to a search list as shown in Figure 3 and determines whether any of the search words in the search list are included in the text data sequentially obtained from the speech recognition unit 29.

[0096] The event detection unit 30 detects an utterance containing a search word as an event if it determines that the text data contains the search word. When the event detection unit 30 detects an event, it obtains the utterance ID acquired from the speech recognition unit 29 as the event ID. Furthermore, the event detection unit 30 refers to a search list as shown in Figure 3 and obtains the priority corresponding to the search word contained in the text data. Once the event detection unit 30 obtains the priority, it executes notification processing according to the priority.

[0097] If the priority is "high," the event detection unit 30, upon determining that the text data contains a search word, outputs an event ID and an output instruction to the speech storage unit 32, and outputs the priority level "high" to the speech notification unit 31. The output instruction instructs the speech storage unit 32 to output the speech sampling data corresponding to the event ID to the speech modulation unit 35 as playback data. When the event detection unit 30 outputs the output instruction, it sets the replay flag to True. In this way, if the priority is "high," the output instruction and other information are output to the speech storage unit 32, etc., immediately after the search word in the text data is detected. With this configuration, as shown in Figure 5, if the priority is "high," a notification sound is played and speech playback begins immediately after the search word is detected.

[0098] When the priority is "medium," the event detection unit 30 obtains information about the end of speech from the speech recognition unit 29, outputs the event ID and output instruction to the speech storage unit 32, and outputs the priority level of "medium" to the speech notification unit 31. When the event detection unit 30 outputs the output instruction, it sets the replay flag to True. In this way, when the priority is "medium," the output instruction and other information are output to the speech storage unit 32, etc., as soon as the speech ends. With this configuration, as shown in Figure 5, when the priority is "medium," a notification sound is played and speech playback begins immediately after the end of speech containing the search word.

[0099] If the priority is "low," the event detection unit 30 obtains information about the end of speech from the speech recognition unit 29, outputs the event ID and a hold instruction to the speech storage unit 32, and outputs the priority of "low" to the speech notification unit 31. The hold instruction is an instruction to the speech storage unit 32 to output the speech sampling data corresponding to the event ID to the speech storage unit 36. The speech sampling data stored in the speech storage unit 36 ​​is played back when the user instructs playback, as described above, as shown in Figure 8. With this configuration, as shown in Figure 5, if the priority is "low," the utterance containing the search word is played back immediately after the user instructs playback.

[0100] The event detection unit 30 updates the notification list stored in the storage unit 26 based on the event ID, priority, detection date and time when the event was detected, and the search word contained in the text data. The notification list in the storage unit 26 includes, for example, the association between the event ID, priority, detection date and time when the event was detected, the search word, and the text data. As an example of the update process, the event detection unit 30 associates the event ID, priority, detection date and time, the search word, and the text data. The event detection unit 30 updates the notification list by including this association in the notification list.

[0101] The event detection unit 30 determines whether the text data contains the search word until it obtains information from the speech recognition unit 29 regarding the end of the utterance. If the event detection unit 30 determines that the text data obtained sequentially from the speech recognition unit 29 does not contain the search word when it obtains information regarding the end of the utterance, it obtains the utterance ID obtained from the speech recognition unit 29 as the clear event ID. The event detection unit 30 outputs the clear event ID to the utterance storage unit 32.

[0102] The event detection unit 30 can obtain information about the start of a new utterance and an utterance ID from the speech recognition unit 29. When the event detection unit 30 obtains information about the start of a new utterance, it determines whether or not any of the search words in the search list are included in the text data newly obtained sequentially from the speech recognition unit 29.

[0103] The event detection unit 30 may obtain left-hand text data and right-hand text data from the speech recognition unit 29. In this case, if the event detection unit 30 determines that the search word is included in either the left-hand text data or the right-hand text data, it may detect the utterance containing the search word as an event. If the event detection unit 30 determines that the search word is not included in either the left-hand text data or the right-hand text data, it may obtain the utterance IDs corresponding to those text data as clear event IDs.

[0104] The speech notification unit 31 obtains a priority from the event detection unit 30. The speech notification unit 31 obtains a notification sound file corresponding to the priority from the storage unit 26. The speech notification unit 31 transmits the obtained notification sound file to the sound collector 10 via the communication unit 21.

[0105] If the priority is "high", the speech notification unit 31 refers to the notification sound list shown in Figure 4 and retrieves the notification sound file "ring.wav" associated with the priority of "high" from the storage unit 26. The speech notification unit 31 then transmits the retrieved notification sound file to the sound collector 10 via the communication unit 21.

[0106] If the priority is "medium," the speech notification unit 31 refers to the notification sound list shown in Figure 4 and retrieves the notification sound file "alert.wav" associated with the priority of "medium" from the storage unit 26. The speech notification unit 31 then transmits the retrieved notification sound file to the sound collector 10 via the communication unit 21.

[0107] If the priority is "low", the speech notification unit 31 refers to the notification sound list shown in Figure 4 and retrieves the notification sound file "notify.wav" associated with the priority of "low" from the storage unit 26. The speech notification unit 31 then transmits the retrieved notification sound file to the sound collector 10 via the communication unit 21.

[0108] As shown in Figure 10, the speech storage unit 32 has a data buffer 33 and a ring buffer 34. The data buffer 33 and the ring buffer 34 are configured to include the same or similar components as the storage unit 26. At least a portion of the data buffer 33 and the ring buffer 34 may be a portion of the storage unit 26. The operation of the speech storage unit 32 is performed by the processor of the control unit 27, etc.

[0109] The speech storage unit 32 obtains information about the start of speech and a speech ID from the section detection unit 28. When the speech storage unit 32 obtains information about the start of speech, etc., it stores the voice sampling data obtained sequentially from the section detection unit 28 in the data buffer 33, associating it with the speech ID. When the speech storage unit 32 obtains new information about the start of speech and a new speech ID from the section detection unit 28, it stores the voice sampling data obtained sequentially from the section detection unit 28 in the data buffer 33, associating it with the new speech ID. In Figure 10, the data buffer 33 stores multiple voice sampling data corresponding to speech ID 1, multiple voice sampling data corresponding to speech ID 2, and multiple voice sampling data corresponding to speech ID 3.

[0110] The speech storage unit 32 receives speech sampling data from the sound collector 10 via the communication unit 21. The speech storage unit 32 stores the speech sampling data received from the sound collector 10 in the ring buffer 34. The speech storage unit 32 refers to the recording time information stored in the memory unit 26 and stores speech sampling data for the recording time in the ring buffer 34. The speech storage unit 32 sequentially stores the speech sampling data in the ring buffer 34 in chronological order.

[0111] The speech storage unit 32 may obtain a clear event ID from the event detection unit 30. If the speech storage unit 32 obtains a clear event ID, it deletes the speech sampling data stored in the data buffer 33 that is associated with a speech ID that matches the clear event ID.

[0112] The speech storage unit 32 can obtain an event ID and an output instruction from the event detection unit 30. When the speech storage unit 32 obtains an output instruction, it identifies a speech ID that matches the event ID obtained along with the output instruction from among the speech sampling data stored in the data buffer 33. The speech storage unit 32 outputs the speech sampling data corresponding to the identified speech ID as playback data to the speech modulation unit 35. The speech storage unit 32 outputs the speech sampling data to the speech modulation unit 35 so that playback begins from the first speech sampling data. The first speech sampling data is the oldest time-series speech sampling data among multiple speech sampling data arranged in chronological order.

[0113] The speech storage unit 32 can obtain an event ID and a hold instruction from the event detection unit 30. When the speech storage unit 32 obtains a hold instruction, it identifies an utterance ID from the audio sampling data stored in the data buffer 33 that matches the event ID obtained along with the hold instruction. The speech storage unit 32 outputs the audio sampling data associated with the identified utterance ID, along with the event ID, to the speech holding unit 36.

[0114] The speech storage unit 32 can acquire a replay instruction. When the speech storage unit 32 acquires a replay instruction, it outputs the audio sampling data stored in the ring buffer 34 to the audio modulation unit 35 as playback data so that playback begins from the first audio sampling data.

[0115] As shown in Figure 2, the audio modulation unit 35 acquires playback data from the speech storage unit 32. If the replay flag is True, the audio modulation unit 35 refers to the audio speed information stored in the storage unit 26 and modulates the playback data so that it is played back as audio at that audio speed. The audio modulation unit 35 transmits the modulated playback data to the sound collector 10 via the communication unit 21.

[0116] The speech retention unit 36 ​​acquires an event ID and voice sampling data from the speech storage unit 32. The speech retention unit 36 ​​stores the acquired voice sampling data in association with the acquired event ID.

[0117] The speech retention unit 36 ​​can acquire an event ID and a playback start instruction. When the speech retention unit 36 ​​acquires a playback start instruction, it identifies the audio sampling data associated with the event ID. The speech retention unit 36 ​​transmits the identified audio sampling data as playback data to the sound collector 10 via the communication unit 21.

[0118] Figure 11 is a flowchart illustrating the operation of the event detection process performed by the audio processing device 20 shown in Figure 2. This operation corresponds to an example of the audio processing method according to this embodiment. For example, when the transmission of audio sampling data from the sound collector 10 to the audio processing device 20 begins, the audio processing device 20 starts the process of step S1.

[0119] The section detection unit 28 receives audio sampling data from the sound collector 10 via the communication unit 21 (step S1).

[0120] In step S2, the interval detection unit 28 sequentially outputs the speech sampling data acquired in step S1 to the speech recognition unit 29 and the speech storage unit 32, respectively.

[0121] In step S2, the section detection unit 28 identifies the start time of speech from the speech sampling data acquired in step S1. Once the section detection unit 28 identifies the start time of speech, it generates a speech ID. The section detection unit 28 outputs the information about the start time of speech and the speech ID to the speech recognition unit 29 and the speech storage unit 32, respectively.

[0122] In step S2, the section detection unit 28 identifies the end of speech from the speech sampling data acquired in step S1. Once the section detection unit 28 identifies the end of speech, it outputs information about the end of speech to the speech recognition unit 29 and the speech storage unit 32, respectively.

[0123] In the process of step S3, when the speech recognition unit 29 obtains information such as the start time of speech from the section detection unit 28, it sequentially converts the speech sampling data obtained sequentially from the section detection unit 28 into text data. When the speech recognition unit 29 outputs information such as the start time of speech to the event detection unit 30, it sequentially outputs the text data as the speech recognition result to the event detection unit 30. When the speech recognition unit 29 obtains information about the end time of speech from the section detection unit 28, it terminates the speech recognition process. However, when the speech recognition unit 29 obtains new information such as the start time of speech from the section detection unit 28, it sequentially converts the speech sampling data obtained sequentially from the section detection unit 28 into text data.

[0124] In step S4, the event detection unit 30 refers to the search list shown in Figure 3 and determines whether or not any of the search words in the search list are included in the text data sequentially acquired from the speech recognition unit 29.

[0125] If the event detection unit 30 determines that the search word is not included in the sequentially acquired text data when it obtains information about the end of speech from the speech recognition unit 29 (step S4: NO), it proceeds to step S5. If the event detection unit 30 determines that the search word is included in the sequentially acquired text data from the speech recognition unit 29 before obtaining information about the end of speech (step S4: YES), it proceeds to step S6.

[0126] In step S5, the event detection unit 30 obtains the speech ID acquired from the speech recognition unit 29 as a clear event ID. The event detection unit 30 outputs the clear event ID to the speech storage unit 32.

[0127] In the processing of step S6, the event detection unit 30 detects an utterance containing the search word as an event.

[0128] In step S7, the event detection unit 30 obtains the utterance ID acquired from the speech recognition unit 29 as the event ID. Furthermore, the event detection unit 30 refers to a search list as shown in Figure 3 and obtains the priority corresponding to the search word contained in the text data.

[0129] In step S8, the event detection unit 30 executes notification processing according to the priority obtained in step S7.

[0130] In step S9, the event detection unit 30 updates the notification list stored in the storage unit 26 based on the event ID, priority, detection date and time when the event was detected, and the search word included in the text data.

[0131] Figures 12 and 13 are flowcharts showing the operation of the output processing of playback data performed by the audio processing device 20 shown in Figure 2. This operation corresponds to an example of the audio processing method according to this embodiment. For example, when the transmission of audio sampling data from the sound collector 10 to the audio processing device 20 begins, the audio processing device 20 starts the process of step S11 as shown in Figure 12.

[0132] In step S11, the audio processing device 20 operates in pass-through mode. In the sound collector 10, the audio playback unit 17 outputs the audio sampling data acquired from the audio acquisition unit 16 to the speaker 12. In step S11, the replay flag is set to False.

[0133] In step S12, the control unit 27 receives input to the region 42 as shown in Figure 6 from the input unit 22 and determines whether or not it has received a replay start notification. If the control unit 27 determines that it has received a replay start notification (step S12: YES), it proceeds to step S13. If the control unit 27 does not determine that it has received a replay start notification (step S12: NO), it proceeds to step S18.

[0134] In step S13, the control unit 27 sets the replay flag to True and outputs a replay instruction to the speech storage unit 32.

[0135] In step S14, the speech storage unit 32 receives a replay instruction. Upon receiving the replay instruction, the speech storage unit 32 starts outputting playback data from the ring buffer 34 to the speech modulation unit 35.

[0136] In step S15, the control unit 27 determines whether all of the playback data has been output from the ring buffer 34 to the audio modulation unit 35. If the control unit 27 determines that all of the playback data has been output (step S15: YES), it proceeds to step S17. If the control unit 27 does not determine that all of the playback data has been output (step S15: NO), it proceeds to step S16.

[0137] In step S16, the control unit 27 receives input to the region 42 as shown in Figure 6 from the input unit 22 and determines whether or not it has received a replay stop request. If the control unit 27 determines that it has received a replay stop request (step S16: YES), it proceeds to step S17. If the control unit 27 does not determine that it has received a replay stop request (step S16: NO), it returns to step S15.

[0138] In step S17, the control unit 27 sets the replay flag to False. After executing step S17, the control unit 27 returns to step S11.

[0139] In step S18, the control unit 27 receives input to the region 63 as shown in Figure 8 via the input unit 22 and determines whether or not it has received a request to start event playback. If the control unit 27 determines that it has received a request to start event playback (step S18: YES), it proceeds to step S19. If the control unit 27 does not determine that it has received a request to start event playback (step S18: NO), it proceeds to step S24 as shown in Figure 13.

[0140] In step S19, the control unit 27 sets the replay flag to True. The control unit 27 also refers to the notification list in the storage unit 26 and obtains the event ID of the selected event from the area 61 as shown in Figure 8. The control unit 27 outputs the event ID and a playback start instruction to the speech holding unit 36.

[0141] In step S20, the speech retention unit 36 ​​obtains an event ID and a playback start instruction. Upon obtaining the playback start instruction, the speech retention unit 36 ​​identifies the audio sampling data associated with the event ID. The speech retention unit 36 ​​then begins transmitting the identified audio sampling data, i.e., the playback data, to the sound collector 10.

[0142] In step S21, the control unit 27 determines whether all of the playback data has been transmitted from the speech holding unit 36 ​​to the sound collector 10. If the control unit 27 determines that all of the playback data has been transmitted (step S21: YES), it proceeds to step S23. If the control unit 27 does not determine that all of the playback data has been transmitted (step S21: NO), it proceeds to step S22.

[0143] In step S22, the control unit 27 receives input to the region 63 as shown in Figure 8 via the input unit 22 and determines whether or not it has received a request to stop event playback. If the control unit 27 determines that it has received a request to stop event playback (step S22: YES), it proceeds to step S23. If the control unit 27 does not determine that it has received a request to stop event playback (step S22: NO), it returns to step S21.

[0144] In step S23, the control unit 27 sets the replay flag to False. After executing step S23, the control unit 27 returns to step S11.

[0145] In the process of step S24 as shown in Figure 13, the speech storage unit 32 determines whether or not it has obtained an event ID and an output instruction from the event detection unit 30. If the speech storage unit 32 determines that it has obtained an event ID and an output instruction (step S24: YES), it proceeds to the process of step S25. If the speech storage unit 32 does not determine that it has obtained an event ID and an output instruction (step S24: NO), it proceeds to the process of step S30.

[0146] In step S25, the replay flag is set to True. This replay flag is set to True by the event detection unit 30 when it outputs an output instruction in step S24 to the speech storage unit 32.

[0147] In step S26, the speech storage unit 32 identifies an utterance ID from the audio sampling data stored in the data buffer 33 that matches the event ID obtained in step S24. The speech storage unit 32 acquires the audio sampling data corresponding to the identified utterance ID as playback data. The speech storage unit 32 starts outputting the playback data from the data buffer 33 to the audio modulation unit 35.

[0148] In step S27, the control unit 27 determines whether all of the playback data has been output from the data buffer 33 to the audio modulation unit 35. If the control unit 27 determines that all of the playback data has been output (step S27: YES), it proceeds to step S29. If the control unit 27 does not determine that all of the playback data has been output (step S27: NO), it proceeds to step S28.

[0149] In step S28, the control unit 27 receives input to the region 42 as shown in Figure 6 from the input unit 22 and determines whether or not it has received a replay stop request. If the control unit 27 determines that it has received a replay stop request (step S28: YES), it proceeds to step S29. If the control unit 27 does not determine that it has received a replay stop request (step S28: NO), it returns to step S27.

[0150] In step S29, the control unit 27 sets the replay flag to False. After executing step S29, the control unit 27 returns to the process of step S11 as shown in Figure 12.

[0151] In step S30, the speech storage unit 32 determines whether or not it has obtained an event ID and a retention instruction from the event detection unit 30. If the speech storage unit 32 determines that it has obtained an event ID and a retention instruction (step S30: YES), it proceeds to step S31. If the speech storage unit 32 does not determine that it has obtained an event ID and a retention instruction (step S30: NO), the control unit 27 returns to the process of step S11 as shown in Figure 12.

[0152] In step S31, the speech storage unit 32 identifies an utterance ID from the audio sampling data stored in the data buffer 33 that matches the event ID obtained in step S30. The speech storage unit 32 outputs the audio sampling data associated with the identified utterance ID, along with the event ID, to the speech holding unit 36.

[0153] After executing the process in step S31, the control unit 27 returns to the process in step S11 as shown in Figure 12.

[0154] In this way, the voice processing device 20, when the control unit 27 detects voice that satisfies the set conditions, notifies the user that voice that satisfies the set conditions has been detected, according to the notification conditions. In this embodiment, when the control unit 27 detects an utterance containing a search word as voice that satisfies the set conditions, it notifies the user that an utterance has been detected, according to the priority set for the search word.

[0155] Here, users may have preferences regarding whether they want to be notified preferentially when an audio that meets the set conditions is detected, depending on the content of the audio, or whether they do not want to be notified preferentially. In this embodiment, by setting a priority as a notification condition, the user can differentiate between audio that should be notified preferentially when detected and audio that should not be notified preferentially when detected. Therefore, the audio processing device 20 can improve user convenience.

[0156] Furthermore, if a sound that meets the set conditions is detected, simply playing that sound may cause the user to miss the played sound. The sound processing device 20 can reduce the possibility of the user missing the played sound by notifying the user that a sound has been detected.

[0157] Therefore, according to this embodiment, an improved voice processing device 20, a voice processing method, and a voice processing system 1 can be provided.

[0158] Furthermore, the control unit 27 of the audio processing device 20 may notify the user that an audio meeting the set conditions has been detected by playing a notification sound if the notification conditions corresponding to the detected audio satisfy the first condition. As described above, the user can immediately notice that an audio has been detected when a notification sound is played.

[0159] Furthermore, if the notification condition corresponding to the detected voice satisfies the first condition, the control unit 27 of the voice processing device 20 may play a notification sound, and then play a voice that satisfies the set condition. With this configuration, as described above, user convenience can be improved.

[0160] Furthermore, if the notification condition corresponding to the detected sound satisfies the second condition, the control unit 27 of the sound processing device 20 may notify the user that a sound satisfying the set conditions has been detected by presenting visual information to the user. As described above, the second condition has a lower priority for notifying the user than the first, third, and fourth conditions. When the priority is low, instead of playing a notification sound, the user can be notified in a manner appropriate to the low priority by presenting visual information to the user.

[0161] Furthermore, the control unit 27 of the voice processing device 20 may present the user with visual information by displaying a notification list to the user if the notification conditions corresponding to the detected voice satisfy the second condition. By viewing the notification list, the user can understand the date and time the voice was detected, and the circumstances under which the voice was detected.

[0162] Furthermore, the control unit 27 of the voice processing device 20 may play the detected voice based on user input if the notification condition corresponding to the detected voice satisfies the second condition. If the priority is low, the user is likely to want to review the detected voice later. This configuration can improve user convenience.

[0163] Furthermore, the control unit 27 of the voice processing device 20 may be controlled so that, if the notification conditions corresponding to the detected voice satisfy the third condition, a notification sound is played and the playback of the utterance begins immediately after the search word is detected. With this configuration, if the priority is high, the user can immediately confirm the content of the utterance.

[0164] Furthermore, the control unit 27 of the voice processing device 20 may control the system so that, if the notification condition corresponding to the detected voice satisfies the fourth condition, a notification sound is played immediately after the utterance ends, and playback of the utterance begins. By starting playback of the utterance immediately after the utterance ends, the utterance spoken in real time and the playback utterance do not overlap. With this configuration, the user can more accurately grasp the content of the playback utterance.

[0165] Furthermore, the control unit 27 of the voice processing device 20 may be controlled to play back the voice data of the utterance section containing the detected utterance. An utterance section is a section in which voice data continues without a set time gap, as described above with reference to Figure 9. By playing back the voice data of such an utterance section, a group of utterances containing the search word is played back. With this configuration, the user can understand the meaning of the utterance containing the search word.

[0166] Furthermore, the control unit 27 of the voice processing device 20 may display on the display unit 23, regardless of priority, the detection date and time when an event, i.e., an utterance, was detected, and the search words contained in that utterance, from among the information included in the notification list. With this configuration, the user can understand the circumstances under which the detected utterance was spoken.

[0167] (Other embodiments) The voice processing system 101 shown in Figure 14 can provide a monitoring service for babies and other infants. The voice processing system 101 includes a sound collector 110 and a voice processing device 20.

[0168] The sound collector 110 and the sound processing device 20 are located further apart from each other than the sound collector 10 and sound processing device 20 shown in Figure 1. For example, the sound collector 110 and the sound processing device 20 are located in separate rooms. The sound collector 110 is located in the room where the baby is. The sound processing device 20 is located in the room where the user is.

[0169] In other embodiments, the setting condition is that the sound matches a pre-set characteristic. The user may input the characteristic of the sound they wish to set as the setting condition from the microphone of the input unit 22 of the sound processing device 20, as shown in Figure 2, and set it as the setting condition in the sound processing device 20. For example, the user may set the characteristics of a baby crying as the setting condition in the sound processing device 20.

[0170] The sound collector 110 includes a microphone 11, a speaker 12, a communication unit 13, a storage unit 14, and a control unit 15, as shown in Figure 2. The sound collector 110 does not necessarily have to include the speaker 12.

[0171] The audio processing device 20 may further include a speaker 12 as shown in Figure 2. The control unit 27 of the audio processing device 20 may further include an audio playback unit 17 and a storage unit 18 as shown in Figure 2.

[0172] In other embodiments, the memory unit 26, as shown in Figure 2, stores a search list that associates data representing speech features with priority levels, instead of the search list shown in Figure 3. The data representing speech features may be data of speech features that can be processed by the machine learning model used by the speech recognition unit 29. Speech features include, for example, Mel-frequency cepstral coefficients (MFCC) or PLP (Perceptual Linear Prediction). For example, the memory unit 26 stores a search list that associates data representing a baby crying with a priority level of "high".

[0173] In another embodiment, the control unit 27 shown in Figure 2 detects audio that satisfies the setting conditions, specifically audio whose characteristics match those of a preset audio.

[0174] The speech recognition unit 29 acquires information about the start of speech, information about the end of speech, speech ID, and speech sampling data from the section detection unit 28, in the same or similar manner as in the embodiments described above. In other embodiments, the speech recognition unit 29 determines whether the speech features in the speech section match pre-set speech features by performing speech recognition processing using a learning model generated by an arbitrary machine learning algorithm.

[0175] If the speech recognition unit 29 determines that the speech characteristics in the utterance section match the pre-set speech characteristics, it outputs a result indicating the match as a speech recognition result, the utterance ID of that utterance section, and data indicating the speech characteristics to the event detection unit 30. If the speech recognition unit 29 determines that the speech characteristics in the utterance section do not match the pre-set speech characteristics, it outputs a result indicating the mismatch as a speech recognition result, and the utterance ID of that utterance section to the event detection unit 30.

[0176] The event detection unit 30 can obtain from the speech recognition unit 29 a result indicating a match as a speech recognition result, an utterance ID, and data indicating pre-set speech characteristics. When the event detection unit 30 obtains a result indicating a match, it detects speech that matches the pre-set speech characteristics as an event. When the event detection unit 30 detects an event, it obtains the utterance ID obtained from the speech recognition unit 29 as the event ID. Furthermore, the event detection unit 30 refers to a search list and obtains a priority corresponding to the data indicating speech characteristics obtained from the speech recognition unit 29. The event detection unit 30 executes notification processing according to the obtained priority, in the same or similar manner as in the embodiment described above.

[0177] The event detection unit 30 may obtain a result indicating a mismatch in the speech recognition result and an utterance ID from the speech recognition unit 29. If the event detection unit 30 obtains a result indicating a mismatch, it obtains the utterance ID obtained from the speech recognition unit 29 as a clear event ID. The event detection unit 30 outputs the clear event ID to the utterance storage unit 32.

[0178] The processing of the audio processing device 20 according to other embodiments is not limited to the processing described above. As another example, the control unit 27 may construct a classifier capable of classifying multiple types of sounds. Furthermore, the control unit 27 may determine which priority level the sounds collected by the sound collector 110 correspond to, based on the result of inputting the audio data collected by the sound collector 110 into the constructed classifier.

[0179] Other effects and configurations of the voice processing system 101 according to other embodiments are the same as or similar to those of the voice processing system 1 shown in Figure 1.

[0180] While this disclosure has been described based on the drawings and embodiments, it should be noted that those skilled in the art will find it easy to make various modifications or alterations based on this disclosure. Therefore, it should be noted that these modifications or alterations are within the scope of this disclosure. For example, the functions, etc., included in each functional part can be rearranged in a logically consistent manner. Multiple functional parts, etc., may be combined into one or divided. The embodiments relating to this disclosure described above are not limited to being implemented strictly according to the respective embodiments, but can be implemented by combining features or omitting some as appropriate. In other words, the contents of this disclosure can be modified and altered in various ways based on this disclosure by those skilled in the art. Therefore, these modifications and alterations are within the scope of this disclosure. For example, in each embodiment, each functional part, each means, each step, etc., can be added to other embodiments in a logically consistent manner, or replaced with each functional part, each means, each step, etc., from other embodiments. Also, in each embodiment, multiple functional parts, each means, each step, etc., can be combined into one or divided. Furthermore, the embodiments of this disclosure described above are not limited to being implemented strictly according to the respective embodiments described, but can also be implemented by combining the features or omitting some of them as appropriate.

[0181] For example, the control unit 27 of the voice processing device 20 may detect multiple voices from a single utterance segment that satisfy different setting conditions. In this case, the control unit 27 may notify the user that voices satisfying the setting conditions have been detected, according to each of the multiple notification conditions set for each of the different setting conditions. Alternatively, the control unit 27 may notify the user that voices satisfying the setting conditions have been detected, according to some of the multiple notification conditions set for each of the different setting conditions. Some of these multiple notification conditions may be notification conditions that satisfy selection conditions. The selection conditions are conditions that have been pre-selected from the first, second, third, and fourth conditions based on user operation, etc. Alternatively, some of these multiple notification conditions may be notification conditions that are included in the Nth (N is an integer of 1 or more) notification conditions, counting from the highest priority for notifying the user, as determined by each of the multiple notification conditions. The Nth may be pre-set based on user operation, etc. For example, the control unit 27 may detect multiple different search words from a single utterance segment. In this case, the control unit 27 may perform processing according to each of the multiple priorities set for each of the multiple different search words. Alternatively, the control unit 27 may perform processing according to some of the multiple priorities set for each of the multiple different search words. Some of the multiple priorities may be, for example, the priorities that are included in the Nth priority, counting from the highest priority.

[0182] For example, the control unit 27 of the speech processing device 20 may detect multiple instances of speech that satisfy the same setting conditions from a single utterance. In this case, the control unit 27 may execute the process of notifying the user according to the notification conditions only once in a single utterance, or it may execute it as many times as it detects speech that satisfies the setting conditions. For example, the control unit 27 may detect the same search word multiple times from a single utterance. In this case, the control unit 27 may execute the process according to priority only once in a single utterance, or it may execute it as many times as it detects the search word.

[0183] For example, the section detection unit 28 shown in Figure 2 may stop detecting speech sections while the replay flag is set to True.

[0184] For example, in the voice processing system 101 shown in Figure 14, the pre-set voice characteristics as setting conditions are described as being those of a baby crying. However, the pre-set voice characteristics as setting conditions are not limited to those of a baby crying. Depending on the usage of the voice processing system 101, any voice characteristics may be set as setting conditions. Other examples include the voice of a supervisor, the ringtone of an intercom, or the ringtone of a telephone.

[0185] For example, in the embodiment described above, the priority was described as being set to three levels, including "high," "medium," and "low." However, the priority is not limited to being set to three levels. The priority may be set to multiple levels. For example, the priority may be set to two levels or four or more levels.

[0186] For example, in the embodiments described above, the sound collector 10 and the voice processing device 20 were described as separate devices. However, the sound collector 10 and the voice processing device 20 may be configured as a single device. An example of this will be described with reference to Figure 15. The voice processing system 201 shown in Figure 15 includes a sound collector 210. The sound collector 210 is an earphone. The sound collector 210 is configured to perform processing of the voice processing device 20. In other words, the sound collector 210 as an earphone becomes the voice processing device of this disclosure. The sound collector 210 includes a microphone 11, a speaker 12, a communication unit 13, a storage unit 14, and a control unit 15, as shown in Figure 2. The control unit 15 of the sound collector 210 includes components corresponding to the control unit 27 of the voice processing device 20. The storage unit 14 of the sound collector 210 stores a notification list. The sound collector 210 may use the user's smartphone or other terminal device to perform screen display and illumination as notification means, as shown in Figure 5. For example, the control unit 15 of the sound collector 210 transmits the notification list from the storage unit 14 to the user's smartphone or the like via the communication unit 13, and displays it on the user's smartphone or the like.

[0187] For example, in the embodiment described above, the voice processing device 20 was described as performing voice recognition processing. However, an external device other than the voice processing device 20 may perform voice recognition processing. The control unit 27 of the voice processing device 20 may acquire the results of the voice recognition processing performed by the external device. The external device may be, for example, a dedicated computer configured to function as a server, a general-purpose personal computer, or a cloud computing system. In this case, the communication unit 13 of the sound collector 10 may further include at least one communication module that can connect to any network, including mobile communication networks and the Internet, in the same or similar manner as the communication unit 21. In the sound collector 10, the control unit 15 may transmit voice sampling data to the external device via the network using the communication unit 13. When the external device receives voice sampling data from the sound collector 10 via the network, it may perform voice recognition processing. The external device may transmit the results of the voice recognition processing to the voice processing device 20 via the network. In the voice processing device 20, the control unit 27 may acquire the results of the voice recognition processing from the external device via the network by receiving them using the communication unit 21.

[0188] For example, in the embodiments described above, the voice processing device 20 was described as a terminal device. However, the voice processing device 20 is not limited to a terminal device. As another example, the voice processing device 20 may be a dedicated computer configured to function as a server, a general-purpose personal computer, or a cloud computing system. In this case, the communication unit 13 of the sound collector 10 may further include at least one communication module that can connect to any network, including mobile communication networks and the Internet, in the same or similar manner as the communication unit 21. The sound collector 10 and the voice processing device 20 may communicate via a network.

[0189] For example, it is also possible to use a general-purpose computer as the voice processing device 20 according to the above-described embodiment. Specifically, a program describing the processing content that realizes each function of the voice processing device 20 according to the above-described embodiment is stored in the memory of the general-purpose computer, and the processor reads and executes the program. Therefore, the configuration according to the above-described embodiment can also be realized as a program that can be executed by a processor or as a non-temporary computer-readable medium that stores the program. [Explanation of symbols]

[0190] 1,101,201 Voice Processing System 10,110,210 Sound collector 11 Mike 12 speakers 13 Communications Department 14 Storage section 15 Control Unit 16. Voice acquisition unit 17 Audio playback unit 18 Storage section 20. Voice processing device 21 Communications Department 22 Input section 23 Display section 24 Vibration section 25 Light-emitting part 26 Memory section 27 Control Unit 28 Section detection unit 29. Voice Recognition Unit 30 Event detection unit 31. Speech Notification Unit 32 Speech storage unit 33 Data Buffer 34 Ring Buffer 35. Audio Modulation Section 40 Main screen 41,42,43,44 area 50 Settings screen 51,52,53,54,55,56 area 60 Notification screen 61,62,63,64 area

Claims

1. A storage unit that stores a list of keywords and their corresponding priorities, It comprises a control unit and, The control unit, By referring to the above list, it is determined whether or not the utterance recognized by the speech recognition process contains the above keyword. If it is determined that the utterance contains the keyword, the priority associated with the keyword is obtained by referring to the list. If the acquired priority is the first priority, immediately after detecting the keyword, a notification sound is played and the playback of the utterance is started. Voice processing device.

2. The control unit is If the acquired priority is a second priority that is lower than the first priority, Immediately after the aforementioned utterance ends, a notification sound is played, and the playback of the aforementioned utterance is started. The audio processing device according to claim 1.

3. The control unit is If the acquired priority is a third priority that is lower than the second priority, By presenting visual information to the user, the system notifies the user that an utterance containing the aforementioned keyword has been detected. The audio processing device according to claim 2.

4. The control unit is A list of utterance information containing the detected keyword is presented to the user. The audio processing device according to claim 3.

5. The control unit is Based on the user input, the system plays back an utterance containing the keyword. The audio processing device according to claim 3.

6. The control unit plays back the audio data of the speech segment including the detected speech, The aforementioned speech interval is a section in which audio data continues without a set time interval. The audio processing device according to claim 1.

7. A speech processing method performed by a speech processing device that stores a list of keywords and priorities associated with each other, By referring to the aforementioned list, it is determined whether or not the utterance recognized by the speech recognition process contains the keyword, If it is determined that the utterance contains the keyword, the priority associated with the keyword is obtained by referring to the list, If the acquired priority is the first priority, the system includes playing a notification sound and starting playback of the utterance immediately after detecting the keyword. Audio processing methods.

8. An acquisition device for acquiring ambient sound, A memory device that stores a list of keywords and their corresponding priorities, Equipped with a voice processing device, The aforementioned audio processing device is Referring to the above list, it is determined whether the keyword is included in the utterance recognized by speech recognition processing on the voice data acquired by the acquisition device. If it is determined that the utterance contains the keyword, the priority associated with the keyword is obtained by referring to the list. If the acquired priority is the first priority, immediately after detecting the keyword, a notification sound is played and the playback of the utterance is started. Voice processing system.

9. When it is determined that an utterance recognized by speech recognition processing contains a pre-set keyword, the control unit plays a notification sound and starts playback of the utterance immediately after detecting the keyword if the priority associated with the keyword is the first priority. Voice processing device.

10. A speech processing method performed by a speech processing device, The process involves determining whether the utterance recognized by speech recognition processing contains pre-set keywords, and If it is determined that the utterance contains the keyword, the priority associated with the keyword is obtained, If the acquired priority is the first priority, the system includes playing a notification sound and starting playback of the utterance immediately after detecting the keyword. Audio processing methods.

11. An acquisition device for acquiring ambient sound, The voice processing device includes, when it determines that a pre-set keyword is included in the utterance recognized by speech recognition processing on the voice data acquired by the acquisition device, and the priority associated with the keyword is first priority, it plays a notification sound and starts playback of the utterance immediately after detecting the keyword. Voice processing system.