Sound processing method, sound processing apparatus, and sound processing program

The audio processing device addresses the challenge of selectively notifying users of detected voices by employing speech recognition and customizable priority settings, thereby enhancing user convenience and reducing distractions.

JP7692098B2Active Publication Date: 2025-06-12KYOCERA CORP
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2024118778
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-01-21
Filing Date
2024-07-24
Publication Date
2025-06-12
Estimated Expiration
2043-01-10

AI Technical Summary

Technical Problem

Conventional audio processing technologies do not effectively allow users to selectively prioritize notifications based on the content of detected voices, leading to potential missed notifications or unnecessary distractions.

Method used

An audio processing device and method that utilize speech recognition to detect voices satisfying preset conditions, allowing users to set notification priorities for specific voices, with notifications being delivered through a combination of sounds and visual cues based on predefined priority levels.

Benefits of technology

Enhances user convenience by allowing tailored notification preferences, reducing the likelihood of missed notifications while minimizing unnecessary distractions, and improving the overall user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007692098000001
    Figure 0007692098000001
  • Figure 0007692098000002
    Figure 0007692098000002
  • Figure 0007692098000003
    Figure 0007692098000003
Patent Text Reader

Abstract

To provide an improved speech processing apparatus, speech processing method, and speech processing system.SOLUTION: A speech processing apparatus comprises a control unit. The control unit acquires a result of speech recognition processing of recognizing a speech from speech data. When the control unit detects a speech satisfying a set condition set in advance to a speech on the basis of the result of the speech recognition processing, it notifies a user of the detection of the speech satisfying the set condition according to a notification condition set in the set condition and corresponding to the detected speech.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Cross-reference to related applications

[0001] This application claims the priority of Japanese Patent Application No. 2022-008227 filed in Japan on January 21, 2022, and incorporates the entire disclosure of the previous application herein by reference.

Technical Field

[0002] The present disclosure relates to an audio processing device, an audio processing method, and an audio processing system.

Background Art

[0003] Conventionally, there is known a technique that enables a user to hear ambient sounds while wearing an audio output device such as headphones or earphones. In such a technique, there is known a portable music playback device including a notification means for notifying from headphones that an external sound matches a predetermined phrase when they match (Patent Document 1).

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

[0005] An audio processing device according to an embodiment of the present disclosure obtains a result of speech recognition processing for recognizing speech in speech data, and when detecting speech that satisfies preset setting conditions regarding the speech based on the result of the speech recognition processing, according to a notification condition corresponding to the detected speech set in the setting conditions, a control unit is provided that notifies the user that speech satisfying the setting conditions has been detected.

[0006] An audio processing method according to an embodiment of the present disclosure obtains a result of speech recognition processing for recognizing speech in speech data, and When a voice that satisfies preset setting conditions regarding voice is detected based on the result of the voice recognition process, the user is notified that a voice that satisfies the setting conditions has been detected according to the notification conditions corresponding to the detected voice set in the setting conditions.

[0007] A voice processing system according to an embodiment of the present disclosure includes a microphone that collects ambient sound, a voice processing device that obtains a result of a voice recognition process for recognizing voice with respect to the voice data collected by the microphone, and when a voice that satisfies preset setting conditions regarding voice is detected based on the result of the voice recognition process, notifies the user that a voice that satisfies the setting conditions has been detected according to the notification conditions corresponding to the detected voice set in the setting conditions.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

DETAILED DESCRIPTION OF THE INVENTION

[0009] There is room for improvement in the conventional technology. For example, depending on the content of the detected voice, the user may want to be preferentially notified that the voice has been detected, or may not want to be preferentially notified. According to one embodiment of the present disclosure, an improved audio processing apparatus, audio processing method, and audio processing system can be provided.

[0010] In the present disclosure, "voice" includes any sound. For example, the voice includes a voice uttered by a person, a sound output by a machine, a cry uttered by an animal, and environmental sounds.

[0011] Hereinafter, embodiments according to the present disclosure will be described with reference to the drawings.

[0012] As shown in FIG. 1, the audio processing system 1 includes a microphone 10 and an audio processing apparatus 20. The microphone 10 and the audio processing apparatus 20 can communicate with each other via a communication line. The communication line includes at least one of wired and wireless.

[0013] In the present embodiment, the microphone 10 is an earphone. However, the microphone 10 is not limited to an earphone. The microphone 10 may be a headphone or the like. The microphone 10 is worn by the user. The microphone 10 can output music or the like. The microphone 10 may include an earphone unit worn on the left ear of the user and an earphone unit worn on the right side of the user.

[0014] The sound collector 10 collects the sound around the sound collector 10. By being worn by the user, the sound collector 10 collects the sound around the user. The sound collector 10 outputs the collected sound around the user based on the control of the sound processing device 20. With such a configuration, the user can listen to the sound around himself / herself while wearing the sound collector 10.

[0015] In this embodiment, the sound processing device 20 is a terminal device. The terminal device serving as the sound processing device 20 is, for example, a mobile phone, a smartphone, a tablet, or a personal computer (PC), etc. However, the sound processing device 20 is not limited to a terminal device.

[0016] The sound processing device 20 is operated by the user. The user can operate the sound processing device 20 to set the sound collector 10 and the like.

[0017] The sound processing device 20 controls the sound collector 10 to collect the sound around the user. When the sound processing device 20 detects a sound that satisfies a preset setting condition from the collected sound around the user, it notifies the user that a sound that satisfies the setting condition has been detected. The details of this process will be described later.

[0018] FIG. 2 is a block diagram of the sound processing system 1 shown in FIG. 1. In FIG. 2, the main flow of data and the like is indicated by a solid line.

[0019] The sound collector 10 includes a microphone 11, a speaker 12, a communication unit 13, a storage unit 14, and a control unit 15.

[0020] The microphone 11 can collect the sound around the sound collector 10. The microphone 11 includes a left microphone and a right microphone. The left microphone may be included in an earphone part worn on the left ear of the user included in the sound collector 10. The right microphone may be included in an earphone part worn on the right side of the user included in the sound collector 10. For example, the microphone 11 is a stereo microphone or the like.

[0021] Speaker 12 can output sound. Speaker 12 includes a left speaker and a right speaker. The left speaker may be included in an earphone unit worn on the left ear part of the user included in the sound collector 10. The right speaker may be included in an earphone unit worn on the right side of the user included in the sound collector 10. For example, Speaker 12 is a stereo speaker or the like.

[0022] The communication unit 13 is configured to include at least one communication module capable of communicating with the voice processing device 20 via a communication line. The communication module is a communication module corresponding to the standard of the communication line. The standard of the communication line is, for example, a wired communication standard or a short-range wireless communication standard including Bluetooth (registered trademark), infrared rays, NFC (Near Field Communication), etc.

[0023] The storage unit 14 is configured to include at least one semiconductor memory, at least one magnetic memory, at least one optical memory, or a combination of at least two of these. The semiconductor memory is, for example, RAM (Random Access Memory) or ROM (Read Only Memory), etc. The RAM is, for example, SRAM (Static Random Access Memory) or DRAM (Dynamic Random Access Memory), etc. The ROM is, for example, EEPROM (Electrically Erasable Programmable Read Only Memory), etc. The storage unit 14 may function as a main storage device, an auxiliary storage device, or a cache memory. The storage unit 14 stores data used for the operation of the sound collector 10 and data obtained by the operation of the sound collector 10. For example, the storage unit 14 stores system programs, application programs, embedded software, etc.

[0024] The control unit 15 is configured to include at least one processor, at least one dedicated circuit, or a combination thereof. The processor is a general-purpose processor such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit), or a dedicated processor specialized for specific processing. The dedicated circuit is, for example, an FPGA (Field-Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). The control unit 15 executes processes related to the operation of the microphone 10 while controlling each part of the microphone 10.

[0025] In this embodiment, the control unit 15 includes an audio acquisition unit 16, an audio playback unit 17, and a storage unit 18. The storage unit 18 is configured to include the same or similar components as the storage unit 14. At least a part of the storage unit 18 may be a part of the storage unit 14. The operation of the storage unit 18 is executed by the processor of the control unit 15 or the like.

[0026] The audio acquisition unit 16 acquires digital audio data from the analog data of the audio collected by the microphone 11. In this embodiment, the audio acquisition unit 16 samples the analog audio data at a preset sampling rate to acquire audio sampling data as the digital audio data.

[0027] The audio acquisition unit 16 outputs the audio sampling data to the audio playback unit 17. Also, the audio acquisition unit 16 transmits the audio sampling data to the audio processing device 20 through the communication unit 13.

[0028] When the microphone 11 includes a left microphone and a right microphone, the audio acquisition unit 16 may acquire left audio sampling data from the analog data of the audio collected by the left microphone. Further, the audio acquisition unit 16 may acquire right audio sampling data from the analog data of the audio collected by the right microphone. The audio acquisition unit 16 may transmit the left audio sampling data and the right audio sampling data to the audio processing device 20 through the communication unit 13. Hereinafter, when the left audio sampling data and the right audio sampling data are not particularly distinguished, they are simply referred to as "audio sampling data".

[0029] The audio playback unit 17 acquires audio sampling data from the audio acquisition unit 16. The audio playback unit 17 receives a replay flag from the audio processing device 20 through the communication unit 13.

[0030] The replay flag is set to True or False. When the replay flag is False, the audio processing system 1 operates in the through mode. The through mode is a mode in which the audio data collected by the sound collector 10 is output from the sound collector 10 without passing through the audio processing device 20. When the replay flag is True, the audio processing system 1 operates in the playback mode. The playback mode is a mode in which the playback data acquired by the sound collector 10 from the audio processing device 20 is output. The conditions for setting the replay flag to True or False will be described later.

[0031] When the replay flag is False, that is, when the audio processing system 1 is in the through mode, the audio playback unit 17 causes the speaker 12 to output the audio sampling data acquired from the audio acquisition unit 16.

[0032] When the replay flag is True, that is, when the audio processing system 1 is in the playback mode, the audio playback unit 17 causes the speaker 12 to output the playback data stored in the storage unit 18.

[0033] The voice playback unit 17 receives the notification sound file from the voice processing device 20 via the communication unit 13. The notification sound file is transmitted from the voice processing device 20 to the voice playback unit 17 when the voice processing device 20 detects a voice that satisfies the set conditions. When the voice playback unit 17 receives the notification sound file, it outputs the notification sound to the speaker 12. With such a configuration, the user can know that a voice that satisfies the set conditions has been detected.

[0034] The storage unit 18 stores the playback data. The playback data is data transmitted from the voice processing device 20 to the microphone 10. When the control unit 15 receives the playback data from the voice processing device 20 via the communication unit 13, it stores the received playback data in the storage unit 18. The control unit 15 can receive a playback stop instruction and a replay stop instruction, which will be described later, from the voice processing device 20 via the communication unit 13. When the control unit 15 receives the playback stop instruction or the replay stop instruction, it deletes the playback data stored in the storage unit 18.

[0035] The control unit 15 may receive the left-side playback data and the right-side playback data from the voice processing device 20 and store them in the storage unit 18. In this case, the voice playback unit 17 may output the left-side playback data stored in the storage unit 18 to the left speaker of the speaker 12 and output the right-side playback data stored in the storage unit 18 to the right speaker of the speaker 12.

[0036] The voice processing device 20 includes a communication unit 21, an input unit 22, a display unit 23, a vibration unit 24, a storage unit 26, and a control unit 27.

[0037] The communication unit 21 includes at least one communication module capable of communicating with the microphone 10 via a communication line. The communication module is a communication module corresponding to the standard of the communication line. The standard of the communication line is, for example, a wired communication standard or a short-range wireless communication standard including Bluetooth (registered trademark), infrared rays, and NFC.

[0038] The communication unit 21 may further include at least one communication module that can be connected to an arbitrary network including a mobile communication network and the Internet or the like. The communication module is, for example, a communication module corresponding to a mobile communication standard such as LTE (Long Term Evolution), 4G (4th Generation), or 5G (5th Generation).

[0039] The input unit 22 can receive an input from a user. The input unit 22 includes at least one input interface capable of receiving an input from a user. The input interface is, for example, a physical key, a capacitive key, a pointing device, a touch screen provided integrally with a display, or a microphone or the like.

[0040] The display unit 23 can display data. The display unit 23 is, for example, a display or the like. The display is, for example, an LCD (Liquid Crystal Display) or an organic EL (Electro Luminescence) display or the like.

[0041] The vibration unit 24 can vibrate the voice processing device 20. The vibration unit 24 includes a vibration element. The vibration element is, for example, a piezoelectric element or the like.

[0042] The light emitting unit 25 can emit light. The light emitting unit 25 is, for example, an LED (Light Emitting Diode) or the like.

[0043] The storage unit 26 is configured to include at least one semiconductor memory, at least one magnetic memory, at least one optical memory, or a combination of at least two of these. The semiconductor memory is, for example, a RAM or a ROM, etc. The RAM is, for example, an SRAM or a DRAM, etc. The ROM is, for example, an EEPROM, etc. The storage unit 26 may function as a main memory device, an auxiliary memory device, or a cache memory. The storage unit 26 stores data used for the operation of the audio processing device 20 and data obtained by the operation of the audio processing device 20. For example, the storage unit 26 stores system programs, application programs, embedded software, etc.

[0044] The storage unit 26 stores, for example, a search list as shown in FIG. 3 described later, a notification sound list and a notification sound file as shown in FIG. 4 described later. The storage unit 26 stores, for example, a notification list described later.

[0045] The control unit 27 is configured to include at least one processor, at least one dedicated circuit, or a combination of these. The processor is a general-purpose processor such as a CPU or a GPU, or a dedicated processor specialized for specific processing. The dedicated circuit is, for example, an FPGA or an ASIC, etc. The control unit 27 executes processes related to the operation of the audio processing device 20 while controlling each part of the audio processing device 20.

[0046] The control unit 27 executes an audio recognition process for recognizing audio in the audio data. However, the control unit 27 may obtain the result of the audio recognition process executed by an external device. When the control unit 27 detects an audio that satisfies the set conditions based on the result of the audio recognition process, it notifies the user that an audio that satisfies the set conditions has been detected. The set conditions are conditions preset for the audio. The control unit 27 notifies the user that an audio that satisfies the set conditions has been detected according to the notification conditions corresponding to the detected audio set in the set conditions.

[0047] The notification condition is a condition for determining the priority of notifying the user that voice has been detected. The higher the priority, the earlier the timing of notifying the user may be. The higher the priority, the more noticeable the notification means for notifying the user may be. For example, as described above, the user is wearing earphones that are the microphone 10. Therefore, when voice such as a notification sound is output from the microphone 10, the user can immediately notice the voice. That is, the notification means by voice such as a notification sound has a higher priority than the notification means by presenting visual information or the like. As will be described later, the user may be notified that voice has been detected by playing the detected voice. In this case, the higher the priority, the earlier the timing of playing the detected voice may be. Also, when the priority is low, the detected voice may be played at an arbitrary timing.

[0048] The notification condition includes a first condition and a second condition. When the notification condition satisfies the second condition, the priority of notifying the user is lower than when the notification condition satisfies the first condition. The first condition includes a third condition and a fourth condition. When the notification condition satisfies the fourth condition, the priority of notifying the user is lower than when the notification condition satisfies the third condition.

[0049] In the present embodiment, the notification condition is set according to the priority. The priority indicates the priority of notifying the user of the voice that satisfies the setting condition. The priority may be set in a plurality of levels. The higher the priority, the higher the priority of notifying the user. In the present embodiment, as shown in FIGS. 3 and 4, the priority is set in three levels including "high", "medium", and "low". The priority of "high" is the highest priority among the three levels of priority. The priority of "medium" is the middle priority among the three levels of priority. The priority of "low" is the lowest priority among the three levels of priority.

[0050] In this embodiment, that the notification condition satisfies the first condition means satisfying the condition that the priority is "high" or "medium". That the notification condition satisfies the second condition means satisfying the condition that the priority is "low". That the notification condition satisfies the third condition means satisfying the condition that the priority is "high". That the notification condition satisfies the fourth condition means satisfying the condition that the priority is "medium".

[0051] When the notification condition corresponding to the detected voice satisfies the first condition, that is, when the priority corresponding to the detected voice is "medium" or "high", the control unit 27 may notify the user that a voice satisfying the set condition has been detected by playing a notification sound. With such a configuration, the user can immediately notice that a voice has been detected.

[0052] When the notification condition corresponding to the detected voice satisfies the first condition, that is, when the priority corresponding to the detected voice is "medium" or "high", after the notification sound is played, the control unit 27 may play the voice satisfying the set condition. With such a configuration, when the priority is "medium" or "high", the detected voice is automatically played. When the priority is "medium" or "high", there is a high possibility that the user wants to immediately check the content of the detected voice. By automatically playing the voice satisfying the set condition when the notification condition satisfies the first condition, the user can immediately check the detected voice. Therefore, the convenience of the user can be improved.

[0053] When the notification condition corresponding to the detected voice satisfies the second condition, that is, when the priority corresponding to the detected voice is "low", the control unit 27 may notify the user that a voice satisfying the set condition has been detected by presenting visual information to the user. When the priority is "low", instead of playing a notification sound, by presenting visual information to the user, a notification commensurate with the low priority can be made.

[0054] In this embodiment, the setting condition is a condition that includes a preset search word. A priority is set for each search word. The search word is composed of, for example, at least one of characters and numbers. The search word may be any information as long as it can be processed as text data. In this embodiment, when the control unit 27 detects a speech including a search word as a voice that satisfies the setting condition, it notifies the user that the speech has been detected.

[0055] FIG. 3 shows a search list. The search list associates a search word with the priority set for the search word. A priority of "high" is set for the search word "flight 153". A priority of "medium" is set for the search word "hello". A priority of "low" is set for the search word "good morning". For example, the control unit 27 generates the search list based on the user's input to a setting screen 50 as shown in FIG. 7 described later.

[0056] FIG. 4 shows a notification sound list. The notification sound list associates a priority with the notification sound set for the priority. The notification sound is used when notifying the user that a speech has been detected. In this embodiment, the notification sound is also used when a speech corresponding to the priority of "low" is detected. However, when a speech corresponding to the priority of "low" is detected, the notification sound may not be used. In FIG. 4, the priority is associated with a notification sound file. The notification sound file is a file for storing the notification sound on a computer. A notification sound file of "ring.wav" is associated with the priority of "high". A notification sound file of "alert.wav" is associated with the priority of "medium". A notification sound file of "notify.wav" is associated with the priority of "low".

[0057] A notification according to the priority according to the present embodiment will be described with reference to FIG. 5. In FIG. 5, the notification means is a means for notifying the user that an utterance including a search word has been detected. The notification timing is the timing for notifying the user that an utterance including a search word has been detected. The reproduction timing is the timing for reproducing the detected utterance.

[0058] As shown in FIG. 5, when the control unit 27 detects an utterance corresponding to the priority of "high", as the notification means, a notification sound and vibration by the vibration unit 24 are used. The control unit 27 sets the notification timing to the timing immediately after the search word is detected. The control unit 27 sets the reproduction timing to the timing immediately after notifying that the utterance has been detected. That is, when the notification condition satisfies the third condition, the control unit 27 controls so that the notification sound is reproduced immediately after the search word is detected and the reproduction of the utterance is started. For example, it is assumed that the priority of "high" is set for the search word of "flight 153". In this case, the control unit 27 sets the timing immediately after the search word of "flight 153" is detected as the notification timing. That is, the control unit 27 vibrates the vibration unit 24 and reproduces the notification sound immediately after the search word of "flight 153" is detected. Further, the control unit 27 sets the timing immediately after the notification sound is reproduced as the reproduction timing, and controls so that the utterance of "flight 153 is scheduled to depart 20 minutes late" is reproduced. With such a configuration, the notification sound is reproduced immediately after the search word of "flight 153" is detected, and the reproduction of the utterance of "flight 153 is scheduled to depart 20 minutes late" is started. Further, the utterance including the search word is automatically reproduced.

[0059] As shown in FIG. 5, when the control unit 27 detects a speech corresponding to the priority of "medium", as a notification means, a notification sound and vibration by the vibration unit 24 are used. The control unit 27 sets the notification timing to immediately after the speech including the search word ends. The control unit 27 sets the reproduction timing to immediately after notifying that the speech has been detected. That is, when the notification condition satisfies the fourth condition, the control unit 27 controls so that the notification sound is reproduced and the reproduction of the speech is started immediately after the speech including the search word ends. For example, it is assumed that the priority of "medium" is set for the search word of "flight 153". In this case, the control unit 27 sets the timing immediately after the speech of "Flight 153 is scheduled to depart 20 minutes late" ends as the notification timing, vibrates the vibration unit 24, and reproduces the notification sound. Further, the control unit 27 sets the timing immediately after the notification sound is reproduced as the reproduction timing, and controls so that the speech of "Flight 153 is scheduled to depart 20 minutes late" is reproduced. With such a configuration, the notification sound is reproduced immediately after the speech of "Flight 153 is scheduled to depart 20 minutes late" ends, and the reproduction of the speech of "Flight 153 is scheduled to depart 20 minutes late" is started. Also, the speech including the search word is automatically reproduced.

[0060] As shown in FIG. 5, when the control unit 27 detects a speech corresponding to the priority of "low", as a notification means, screen display by the display unit 23 and light emission by the light emission unit 25 are used. The screen display and the light emission are an example of a notification means for presenting visual information to the user. The control unit 27 causes the display unit 23 to display a notification list as the screen display. The notification list is a list of information of voices that satisfy the detected setting conditions. In the present embodiment, the notification list is a list of event information. An event is a speech including a search word. Details of the notification list will be described later. The control unit 27 sets the reproduction timing to immediately after the user instructs the reproduction of the speech. That is, when the notification condition satisfies the second condition, the control unit 27 reproduces the speech including the search word based on the user's input. With such a configuration, the speech including the search word is manually reproduced.

[0061] <Input / Output Processing> The control unit 27 receives an input from the user via the input unit 22. Based on the input received by the input unit 22, the control unit 27 selects, for example, the screen to be displayed on the display unit 23. For example, the control unit 27 causes a screen as shown in FIG. 6, FIG. 7, or FIG. 8 to be displayed based on the input received by the input unit 22. In the configurations shown in FIGS. 6 to 8, the input unit 22 is a touch screen provided integrally with the display of the display unit 23.

[0062] The main screen 40 as shown in FIG. 6 includes an area 41, an area 42, an area 43, and an area 44.

[0063] The state of the microphone 10 is displayed in the area 41. In FIG. 6, the information "Replaying…" indicating that the microphone 10 is in the replay mode is displayed in the area 41.

[0064] When the voice processing system 1 is in the through mode, the characters "Start Replay" are displayed in the area 42. When the voice processing system 1 is in the playback mode, the characters "Stop Replay" are displayed in the area 42. The control unit 27 can receive an input to the area 42 via the input unit 22.

[0065] When the characters "Start Replay" are displayed in the area 42, that is, when the voice processing system 1 is in the through mode, the control unit 27 can receive the start of replay by receiving an input to the area 42 via the input unit 22. When the control unit 27 receives the start of replay, it sets the replay flag to True and outputs a replay instruction to the speech accumulation unit 32 described later.

[0066] When the characters "Stop Replay" are displayed in the area 42, that is, when the voice processing system 1 is in the playback mode, the control unit 27 can receive the stop of replay by receiving an input to the area 42 via the input unit 22. When the control unit 27 receives the stop of replay, it sets the replay flag to False and transmits a replay stop instruction to the microphone 10 via the communication unit 21.

[0067] In area 43, the characters "Notification List" are displayed. The control unit 27 can receive an input to area 43 by the input unit 22. When the control unit 27 receives an input to area 43 by the input unit 22, it causes the display unit 23 to display a notification screen 60 as shown in FIG. 8.

[0068] In area 44, the characters "Settings" are displayed. The control unit 27 can receive an input to area 44 by the input unit 22. When the control unit 27 receives an input to area 44 by the input unit 22, it causes the display unit 23 to display a settings screen 50 as shown in FIG. 7.

[0069] The settings screen 50 as shown in FIG. 7 is a screen for the user to perform various settings. The settings screen 50 includes an area 51, an area 52, an area 53, an area 54, an area 55, and an area 56.

[0070] In area 51, the characters "Add Search Word" are displayed. The control unit 27 can receive an input to area 51 by the input unit 22. The control unit 27 receives an input of a search word and an input of a priority corresponding to the search word from area 51.

[0071] In area 52, the set search words are displayed. In FIG. 7, the search words "Flight 153", "Hello", and "Good Morning" are displayed in area 52. The control unit 27 can receive an input to area 52 by the input unit 22. When the control unit 27 receives an input to area 52 by the input unit 22, it causes the display unit 23 to display a search list as shown in FIG. 3.

[0072] In area 53, the characters "Recording Buffer Setting" are displayed. Area 53 is used to set the length of the recording time for recording the sound collected by the microphone 10. In the present embodiment, the voice sampling data for the recording time is stored in the ring buffer 34 as shown in FIG. 10 described later. The control unit 27 can receive an input to area 53 by the input unit 22. The control unit 27 receives inputs of recording times such as 5 seconds, 10 seconds, and 15 seconds, for example. The control unit 27 causes the storage unit 26 to store the received recording time information.

[0073] In area 54, the characters "Speed Setting" are displayed. Area 54 is used to set the playback speed of the sound output from the microphone 10. The control unit 27 can receive an input to area 54 by the input unit 22. The control unit 27 receives inputs of voice speeds such as 1x speed, 1.1x speed, and 1.2x speed, for example. The control unit 27 causes the storage unit 26 to store the received voice speed information.

[0074] In area 55, the characters "Voice Threshold Setting" are displayed. Area 55 is used to set a voice threshold for cutting out the sound collected by the microphone 10 as noise. In the present embodiment, the sound below the voice threshold is cut out as noise. The control unit 27 can receive an input to area 55 by the input unit 22. The control unit 27 receives inputs of voice thresholds from -50 [dBA] to -5 [dBA], for example. The control unit 27 causes the storage unit 26 to store the received voice threshold information.

[0075] In area 56, the characters "Setting Completed" are displayed. The control unit 27 can receive an input to area 56 by the input unit 22. When the control unit 27 receives an input to area 56 by the input unit 22, the control unit 27 causes the display unit 23 to display the main screen 40 as shown in FIG. 6.

[0076] The notification screen 60 as shown in FIG. 8 is a screen for notifying various information to the user. The notification screen 60 includes an area 61, an area 62, an area 63, and an area 64.

[0077] In area 61, a notification list is displayed. As described above, the notification list is a list of event information. An event, as described above, is a speech including a search word. The control unit 27 causes the event information with a "low" priority among the events included in the notification list to be displayed in area 61. However, the control unit 27 may cause all the event information included in the notification list to be displayed in area 61 regardless of the priority. The control unit 27 can accept an input for each event in the notification list displayed in area 61 by the input unit 22. The control unit 27 accepts the selection of an event in the notification list by accepting an input for each event in the notification list from area 61 by the input unit 22.

[0078] In area 62, the characters "Detailed Display" are displayed. The control unit 27 can accept an input for area 62 by the input unit 22. The control unit 27 can accept the selection of an event included in the notification list from area 61 and further accept an input for area 62 by the input unit 22. In this case, the control unit 27 causes the details of the event information selected from area 61 to be displayed in area 61 by the display unit 23. For example, the control unit 27 causes the voice recognition result for the left side and the voice recognition result for the right side, which will be described later, to be displayed as the details of the event information.

[0079] In area 63, the characters "Play start / Play stop" are displayed. The control unit 27 can accept play start or play stop by receiving an input to area 63 by the input unit 22. When the speech is not being played, the control unit 27 accepts the selection of an event included in the notification list from area 61, and further accepts an input to area 63 by the input unit 22, thereby accepting the start of event playback. When accepting the start of event playback, the control unit 27 controls so that the event selected from area 61, that is, the speech, is played. In the present embodiment, the control unit 27 refers to the notification list in the storage unit 26 and acquires the event ID of the event selected from area 61, which will be described later. The control unit 27 outputs the event ID and the play start instruction to the speech holding unit 36, which will be described later, and controls so that the speech is played. Also, when the speech is being played, the control unit 27 accepts the stop of event playback by receiving an input to area 63 by the input unit 22. When accepting the stop of event playback, the control unit 27 controls so that the playback of the speech stops. In the present embodiment, a stop playback instruction is transmitted to the microphone 10 by the communication unit 21, and control is performed so that the playback of the speech stops.

[0080] In area 64, the characters "Return" are displayed. The control unit 27 can accept an input to area 64 by the input unit 22. When the control unit 27 accepts an input to area 64 by the input unit 22, the main screen 40 as shown in FIG. 6 is displayed on the display unit 23.

[0081] <Voice processing> As shown in FIG. 2, the control unit 27 includes an interval detection unit 28, a speech recognition unit 29, an event detection unit 30, a speech notification unit 31, a speech accumulation unit 32, a voice modulation unit 35, and a speech holding unit 36. The speech holding unit 36 is configured to include the same or similar components as the storage unit 26. At least a part of the speech holding unit 36 may be a part of the storage unit 26. The operation of the speech holding unit 36 is executed by a processor or the like of the control unit 27.

[0082] The section detection unit 28 receives voice sampling data from the microphone 10 via the communication unit 21. The section detection unit 28 detects a speaking section from the voice sampling data. The speaking section is a section in which the speaking state continues. By detecting the speaking section from the voice sampling data, the section detection unit 28 can also detect a non-speaking section. The non-speaking section is a section in which the non-speaking state continues. The start point of the speaking section is also referred to as the "speaking start time". The start point of the speaking section is the end point of the non-speaking section. The end point of the speaking section is also referred to as the "speaking end time". The end point of the speaking section is the start point of the non-speaking section.

[0083] An example of the processing of the section detection unit 28 will be described with reference to FIG. 9. However, the processing of the section detection unit 28 is not limited to the processing described with reference to FIG. 9. The section detection unit 28 may detect a speaking section from the voice sampling data by any method. As another example, the section detection unit 28 may detect a speaking section from the voice sampling data by a machine learning model generated using any machine learning algorithm.

[0084] In FIG. 9, the horizontal axis represents time. The voice sampling data as shown in FIG. 9 is obtained by the voice acquisition unit 16 of the microphone 10. The section detection unit 28 acquires voice section detection data from the voice sampling data. The voice section detection data is data obtained by averaging the power of the voice sampling data over a preset time width. The time width of the voice section detection data may be set based on the specifications of the voice processing apparatus 20 and the like. In FIG. 9, one piece of voice section detection data is shown as one square. The time width of this one square, that is, the time width of one piece of voice section detection data, is, for example, 200 [ms].

[0085] The section detection unit 28 acquires information on the voice threshold from the storage unit 26 and classifies the voice section detection data into voice data and non-voice data. In FIG. 9, the voice data is the data with a dark color among the voice section detection data shown as squares. Also, the non-voice data is the data with a white background among the voice section detection data shown as squares. When the value of the voice section detection data is null, the section detection unit 28 classifies the voice section detection data as non-voice data. When the value of the voice section detection data is not null and the voice section detection data is less than the voice threshold, the section detection unit 28 classifies the voice section detection data as non-voice data. When the value of the voice section detection data is not null and the value of the voice section detection data is equal to or greater than the voice threshold, the section detection unit 28 classifies the voice section detection data as voice data.

[0086] The section detection unit 28 detects, as a speaking section, a section in which voice data continues without a set time interval. The set time interval may be set based on the language processed by the voice processing device 20. When the language to be processed is Japanese, the set time interval is, for example, 500 [ms]. In FIG. 9, when the section detection unit 28 detects voice data after non-voice data has continued for longer than the set time interval, the section detection unit 28 specifies the time point at which the voice data is detected as the speaking start time point. For example, the section detection unit 28 specifies time t1 as the speaking start time point. After specifying the speaking start time point, when the section detection unit 28 determines that non-voice data has continued for longer than the set time interval, the section detection unit 28 specifies the time point at which the determination is made as the speaking end time point. For example, the section detection unit 28 specifies time t2 as the speaking end time point. The section detection unit 28 detects, as a speaking section, the section from the speaking start time point to the speaking end time point.

[0087] The section detection unit 28 may receive left-side audio sampling data and right-side audio sampling data from the microphone 10. In this case, when non-audio data continues for more than a set time in both the left-side and right-side audio sampling data and then audio data is detected in either the left-side or right-side data, the section detection unit 28 may specify the time point at which the audio data is detected as the start time point of speech. Also, when it is determined that non-audio data has continued for more than a set time in both the left-side and right-side data, the section detection unit 28 may specify the time point at which the determination is made as the end time point of speech.

[0088] When the section detection unit 28 specifies the start time point of speech from the audio sampling data, it generates a speech ID. The speech ID is identification information that can be uniquely identified. The section detection unit 28 outputs the information on the start time point of speech and the speech ID to the speech recognition unit 29 and the speech storage unit 32, respectively.

[0089] When the section detection unit 28 specifies the end time point of speech from the audio sampling data, it outputs the information on the end time point of speech to the speech recognition unit 29 and the speech storage unit 32, respectively.

[0090] The section detection unit 28 sequentially outputs the audio sampling data received from the microphone 10 to the speech recognition unit 29 and the speech storage unit 32, respectively.

[0091] The speech recognition unit 29 acquires the information on the start time point of speech and the speech ID from the section detection unit 28. When the speech recognition unit 29 acquires information such as the start time point of speech, it executes a speech recognition process for recognizing speech on the audio sampling data sequentially acquired from the section detection unit 28. In the present embodiment, the speech recognition unit 29 recognizes speech by converting the audio data included in the audio sampling data into text data by the speech recognition process.

[0092] The voice recognition unit 29 outputs the information on the start time of speech and the speech ID acquired from the section detection unit 28 to the event detection unit 30. When the voice recognition unit 29 outputs the information on the start time of speech and the like to the event detection unit 30, it sequentially outputs text data as the voice recognition result to the event detection unit 30.

[0093] The voice recognition unit 29 acquires the information on the end time of speech from the section detection unit 28. When the voice recognition unit 29 acquires the information on the end time of speech, it ends the voice recognition process. The voice recognition unit 29 outputs the information on the end time of speech acquired from the section detection unit 28 to the event detection unit 30. After that, the voice recognition unit 29 can acquire the information on the start time of a new speech and the speech ID from the section detection unit 28. When the voice recognition unit 29 acquires the information on the start time of a new speech and the like, it executes the voice recognition process again on the voice sampling data sequentially acquired from the section detection unit 28.

[0094] The voice recognition unit 29 may acquire the left voice sampling data and the right voice sampling data from the section detection unit 28. In this case, the voice recognition unit 29 may convert each of the left voice sampling data and the right voice sampling data into text data. Hereinafter, the text data acquired from the left voice sampling data is also described as "left text data" or "left voice recognition result". The text data acquired from the right voice sampling data is also described as "right text data" or "right voice recognition result".

[0095] The event detection unit 30 acquires the information on the start time of speech and the speech ID from the voice recognition unit 29. After the event detection unit 30 acquires the information on the start time of speech and the like, it sequentially acquires text data from the voice recognition unit 29. The event detection unit 30 refers to a search list as shown in FIG. 3 and determines whether any of the search words in the search list is included in the text data sequentially acquired from the voice recognition unit 29.

[0096] When the event detection unit 30 determines that the search word is included in the text data, it detects the utterance including the search word as an event. When the event detection unit 30 detects an event, it acquires the utterance ID obtained from the speech recognition unit 29 as the event ID. Further, the event detection unit 30 refers to a search list as shown in FIG. 3, and acquires the priority corresponding to the search word included in the text data. When the event detection unit 30 acquires the priority, it executes a notification process according to the priority.

[0097] When the priority is "high", when the event detection unit 30 determines that the search word is included in the text data, it outputs the event ID and the output instruction to the utterance accumulation unit 32, and outputs the priority of "high" to the utterance notification unit 31. The output instruction is an instruction to cause the utterance accumulation unit 32 to output the voice sampling data corresponding to the event ID to the voice modulation unit 35 as reproduction data. When the event detection unit 30 outputs the output instruction, it sets the replay flag to True. Thus, when the priority is "high", immediately after the search word included in the text data is detected, the output instruction and the like are output to the utterance accumulation unit 32 and the like. With such a configuration, as shown in FIG. 5, when the priority is "high", immediately after the search word is detected, the notification sound is reproduced and the reproduction of the utterance is started.

[0098] When the priority is "medium", when the event detection unit 30 acquires the information at the end of the utterance from the speech recognition unit 29, it outputs the event ID and the output instruction to the utterance accumulation unit 32, and outputs the priority of "medium" to the utterance notification unit 31. When the event detection unit 30 outputs the output instruction, it sets the replay flag to True. Thus, when the priority is "medium", at the time when the utterance ends, the output instruction and the like are output to the utterance accumulation unit 32 and the like. With such a configuration, as shown in FIG. 5, when the priority is "medium", immediately after the utterance including the search word ends, the notification sound is reproduced and the reproduction of the utterance is started.

[0099] When the priority is "low", when the event detection unit 30 acquires information on the end point of speech from the speech recognition unit 29, it outputs an event ID and a holding instruction to the speech accumulation unit 32, and outputs the priority of "low" to the speech notification unit 31. The holding instruction is an instruction to cause the speech accumulation unit 32 to output the speech sampling data corresponding to the event ID to the speech holding unit 36. The speech sampling data held in the speech holding unit 36 is reproduced when the user instructs reproduction as described above with reference to FIG. 8. With such a configuration, when the priority is "low" as shown in FIG. 5, the speech including the search word is reproduced immediately after the user instructs reproduction.

[0100] The event detection unit 30 updates the notification list stored in the storage unit 26 based on the event ID, the priority, the detection date and time when the event is detected, and the search word included in the text data. The notification list in the storage unit 26 includes, for example, the association of the event ID, the priority, the detection date and time when the event is detected, the search word, and the text data. As an example of the update process, the event detection unit 30 associates the event ID, the priority, the detection date and time, the search word, and the text data. The event detection unit 30 updates the notification list by including this association in the notification list.

[0101] The event detection unit 30 determines whether or not the text data includes a search word until it acquires information on the end point of speech from the speech recognition unit 29. When the event detection unit 30 determines that the text data sequentially acquired from the speech recognition unit 29 does not include the search word at the time when it acquires the information on the end point of speech, the event detection unit 30 acquires the speech ID acquired from the speech recognition unit 29 as the clear event ID. The event detection unit 30 outputs the clear event ID to the speech accumulation unit 32.

[0102] The event detection unit 30 can acquire information on the start point of a new speech and the speech ID from the speech recognition unit 29. When the event detection unit 30 acquires information such as the start point of a new speech, it determines whether or not any of the search words in the search list is included in the text data newly and sequentially acquired from the speech recognition unit 29.

[0103] The event detection unit 30 may acquire left text data and right text data from the speech recognition unit 29. In this case, when the event detection unit 30 determines that a search word is included in either the left text data or the right text data, the utterance including the search word may be detected as an event. When the event detection unit 30 determines that the search word is not included in both the left text data and the right text data, the utterance ID corresponding to the text data may be acquired as a clear event ID.

[0104] The utterance notification unit 31 acquires the priority from the event detection unit 30. The utterance notification unit 31 acquires a notification sound file corresponding to the priority from the storage unit 26. The utterance notification unit 31 transmits the acquired notification sound file to the sound collector 10 through the communication unit 21.

[0105] When the priority is "high", the utterance notification unit 31 refers to a notification sound list as shown in FIG. 4, and acquires a notification sound file "ring.wav" associated with the "high" priority from the storage unit 26. The utterance notification unit 31 transmits the acquired notification sound file to the sound collector 10 through the communication unit 21.

[0106] When the priority is "medium", the utterance notification unit 31 refers to a notification sound list as shown in FIG. 4, and acquires a notification sound file "alert.wav" associated with the "medium" priority from the storage unit 26. The utterance notification unit 31 transmits the acquired notification sound file to the sound collector 10 through the communication unit 21.

[0107] When the priority is "low", the utterance notification unit 31 refers to a notification sound list as shown in FIG. 4, and acquires a notification sound file "notify.wav" associated with the "low" priority from the storage unit 26. The utterance notification unit 31 transmits the acquired notification sound file to the sound collector 10 through the communication unit 21.

[0108] As shown in FIG. 10, the speech accumulation unit 32 includes a data buffer 33 and a ring buffer 34. The data buffer 33 and the ring buffer 34 are configured to include the same or similar components as the storage unit 26. At least a part of the data buffer 33 and the ring buffer 34 may be a part of the storage unit 26. The operation of the speech accumulation unit 32 is executed by a processor or the like of the control unit 27.

[0109] The speech accumulation unit 32 acquires information on the start time of speech and a speech ID from the section detection unit 28. When the speech accumulation unit 32 acquires information such as the start time of speech, it accumulates the voice sampling data sequentially acquired from the section detection unit 28 in the data buffer 33 in association with the speech ID. When the speech accumulation unit 32 acquires information on a new start time of speech and a new speech ID from the section detection unit 28, it accumulates the voice sampling data sequentially acquired from the section detection unit 28 in the data buffer 33 in association with the new speech ID. In FIG. 10, a plurality of voice sampling data corresponding to speech ID1, a plurality of voice sampling data corresponding to speech ID2, and a plurality of voice sampling data corresponding to speech ID3 are stored in the data buffer 33.

[0110] The speech accumulation unit 32 receives voice sampling data from the microphone 10 through the communication unit 21. The speech accumulation unit 32 accumulates the voice sampling data received from the microphone 10 in the ring buffer 34. The speech accumulation unit 32 refers to the recording time information stored in the storage unit 26 and accumulates voice sampling data for the recording time in the ring buffer 34. The speech accumulation unit 32 sequentially accumulates voice sampling data in the ring buffer 34 in time series.

[0111] The speech accumulation unit 32 may acquire a clear event ID from the event detection unit 30. When the speech accumulation unit 32 acquires the clear event ID, it deletes the voice sampling data associated with the speech ID that matches the clear event ID among the voice sampling data stored in the data buffer 33.

[0112] The speech accumulation unit 32 can acquire an event ID and an output instruction from the event detection unit 30. When the speech accumulation unit 32 acquires the output instruction, it identifies a speech ID that matches the event ID acquired together with the output instruction from among the voice sampling data stored in the data buffer 33. The speech accumulation unit 32 outputs the voice sampling data corresponding to the identified speech ID as reproduction data to the voice modulation unit 35. The speech accumulation unit 32 outputs the voice sampling data to the voice modulation unit 35 so that it is reproduced from the first voice sampling data. The first of the voice sampling data is the voice sampling data at the oldest time among a plurality of voice sampling data along the time series.

[0113] The speech accumulation unit 32 can acquire an event ID and a holding instruction from the event detection unit 30. When the speech accumulation unit 32 acquires the holding instruction, it identifies a speech ID that matches the event ID acquired together with the holding instruction from among the voice sampling data stored in the data buffer 33. The speech accumulation unit 32 outputs the voice sampling data associated with the identified speech ID to the speech holding unit 36 together with the event ID.

[0114] The speech accumulation unit 32 can acquire a replay instruction. When the speech accumulation unit 32 acquires the replay instruction, it outputs the voice sampling data stored in the ring buffer 34 as reproduction data to the voice modulation unit 35 so that it is reproduced from the first voice sampling data.

[0115] As shown in FIG. 2, the voice modulation unit 35 acquires reproduction data from the speech accumulation unit 32. When the replay flag is True, the voice modulation unit 35 refers to the voice speed information stored in the storage unit 26 and modulates the reproduction data so that the reproduction data is reproduced as voice at that voice speed. The voice modulation unit 35 transmits the modulated reproduction data to the sound collector 10 via the communication unit 21.

[0116] The speech holding unit 36 acquires the event ID and voice sampling data from the speech accumulation unit 32. The speech holding unit 36 holds the acquired voice sampling data in association with the acquired event ID.

[0117] The speech holding unit 36 can acquire the event ID and the playback start instruction. When the speech holding unit 36 acquires the playback start instruction, it specifies the voice sampling data associated with the event ID. The speech holding unit 36 transmits the specified voice sampling data as playback data to the microphone 10 via the communication unit 21.

[0118] FIG. 11 is a flowchart showing the operation of the event detection process executed by the voice processing apparatus 20 shown in FIG. 2. This operation corresponds to an example of the voice processing method according to the present embodiment. For example, when the transmission of voice sampling data from the microphone 10 to the voice processing apparatus 20 is started, the voice processing apparatus 20 starts the process of step S1.

[0119] The section detection unit 28 receives voice sampling data from the microphone 10 via the communication unit 21 (step S1).

[0120] In the process of step S2, the section detection unit 28 sequentially outputs the voice sampling data acquired in the process of step S1 to the voice recognition unit 29 and the speech accumulation unit 32, respectively.

[0121] In the process of step S2, the section detection unit 28 specifies the speech start time from the voice sampling data acquired in the process of step S1. When the section detection unit 28 specifies the speech start time, it generates a speech ID. The section detection unit 28 outputs the information on the speech start time and the speech ID to the voice recognition unit 29 and the speech accumulation unit 32, respectively.

[0122] In the process of step S2, the section detection unit 28 identifies the end point of speech from the voice sampling data obtained in the process of step S1. When the section detection unit 28 identifies the end point of speech, it outputs the information of the end point of speech to the voice recognition unit 29 and the speech storage unit 32 respectively.

[0123] In the process of step S3, when the voice recognition unit 29 obtains information such as the start point of speech from the section detection unit 28, it sequentially converts the voice sampling data sequentially obtained from the section detection unit 28 into text data. When the voice recognition unit 29 outputs information such as the start point of speech to the event detection unit 30, it sequentially outputs the text data as the voice recognition result to the event detection unit 30. When the voice recognition unit 29 obtains the information of the end point of speech from the section detection unit 28, it ends the voice recognition process. However, when the voice recognition unit 29 obtains new information such as the start point of speech from the section detection unit 28, it sequentially converts the voice sampling data sequentially obtained from the section detection unit 28 into text data.

[0124] In the process of step S4, the event detection unit 30 refers to the search list as shown in FIG. 3 and determines whether any of the search words in the search list is included in the text data sequentially obtained from the voice recognition unit 29.

[0125] When the event detection unit 30 determines that the search word is not included in the sequentially obtained text data when it obtains the information of the end point of speech from the voice recognition unit 29 (step S4: NO), it proceeds to the process of step S5. When the event detection unit 30 determines that the search word is included in the text data sequentially obtained from the voice recognition unit 29 before obtaining the information of the end point of speech (step S4: YES), it proceeds to the process of step S6.

[0126] In the process of step S5, the event detection unit 30 obtains the speech ID obtained from the voice recognition unit 29 as the clear event ID. The event detection unit 30 outputs the clear event ID to the speech storage unit 32.

[0127] In the process of step S6, the event detection unit 30 detects an utterance including the search word as an event.

[0128] In the process of step S7, the event detection unit 30 acquires the utterance ID obtained from the speech recognition unit 29 as the event ID. Further, the event detection unit 30 refers to a search list as shown in FIG. 3 and acquires the priority corresponding to the search word included in the text data.

[0129] In the process of step S8, the event detection unit 30 executes a notification process according to the priority acquired in the process of step S7.

[0130] In the process of step S9, the event detection unit 30 updates the notification list stored in the storage unit 26 based on the event ID, the priority, the detection date and time when the event is detected, and the search word included in the text data.

[0131] FIG. 12 and FIG. 13 are flowcharts showing the operation of the output process of the reproduction data executed by the audio processing apparatus 20 shown in FIG. 2. This operation corresponds to an example of the audio processing method according to the present embodiment. For example, when the transmission of the audio sampling data from the microphone 10 to the audio processing apparatus 20 is started, the audio processing apparatus 20 starts the process of step S11 as shown in FIG. 12.

[0132] In the process of step S11, the audio processing apparatus 20 operates in the through mode. In the microphone 10, the audio reproduction unit 17 causes the speaker 12 to output the audio sampling data acquired from the audio acquisition unit 16. In the process of step S11, the replay flag is set to False.

[0133] In the process of step S12, the control unit 27 determines whether to accept the start of replay by receiving an input for the area 42 as shown in FIG. 6 from the input unit 22. If the control unit 27 determines that the start of replay has been accepted (step S12: YES), it proceeds to the process of step S13. If the control unit 27 determines that the start of replay has not been accepted (step S12: NO), it proceeds to the process of step S18.

[0134] In the process of step S13, the control unit 27 sets the replay flag to True and outputs a replay instruction to the speech accumulation unit 32.

[0135] In the process of step S14, the speech accumulation unit 32 acquires the replay instruction. When the speech accumulation unit 32 acquires the replay instruction, it starts outputting playback data from the ring buffer 34 to the voice modulation unit 35.

[0136] In the process of step S15, the control unit 27 determines whether all the playback data has been output from the ring buffer 34 to the voice modulation unit 35. If the control unit 27 determines that all the playback data has been output (step S15: YES), it proceeds to the process of step S17. If the control unit 27 determines that all the playback data has not been output (step S15: NO), it proceeds to the process of step S16.

[0137] In the process of step S16, the control unit 27 determines whether to accept the stop of replay by receiving an input for the area 42 as shown in FIG. 6 from the input unit 22. If the control unit 27 determines that the stop of replay has been accepted (step S16: YES), it proceeds to the process of step S17. If the control unit 27 determines that the stop of replay has not been accepted (step S16: NO), it returns to the process of step S15.

[0138] In the process of step S17, the control unit 27 sets the replay flag to False. After executing the process of step S17, the control unit 27 returns to the process of step S11.

[0139] In the process of step S18, the control unit 27 determines whether it has received the start of event playback by receiving an input to the area 63 as shown in FIG. 8 by the input unit 22. When the control unit 27 determines that it has received the start of event playback (step S18: YES), it proceeds to the process of step S19. When the control unit 27 determines that it has not received the start of event playback (step S18: NO), it proceeds to the process of step S24 as shown in FIG. 13.

[0140] In the process of step S19, the control unit 27 sets the replay flag to True. Also, the control unit 27 refers to the notification list in the storage unit 26 and acquires the event ID of the event selected from the area 61 as shown in FIG. 8. The control unit 27 outputs the event ID and the playback start instruction to the speech holding unit 36.

[0141] In the process of step S20, the speech holding unit 36 acquires the event ID and the playback start instruction. When the speech holding unit 36 acquires the playback start instruction, it specifies the voice sampling data associated with the event ID. The speech holding unit 36 starts transmitting the specified voice sampling data, that is, the playback data, to the microphone 10.

[0142] In the process of step S21, the control unit 27 determines whether all the playback data has been transmitted from the speech holding unit 36 to the microphone 10. When the control unit 27 determines that all the playback data has been transmitted (step S21: YES), it proceeds to the process of step S23. When the control unit 27 determines that not all the playback data has been transmitted (step S21: NO), it proceeds to the process of step S22.

[0143] In the process of step S22, the control unit 27 determines whether it has received a stop of event playback by receiving an input for the area 63 as shown in FIG. 8 through the input unit 22. If the control unit 27 determines that it has received a stop of event playback (step S22: YES), it proceeds to the process of step S23. If the control unit 27 determines that it has not received a stop of event playback (step S22: NO), it returns to the process of step S21.

[0144] In the process of step S23, the control unit 27 sets the replay flag to False. After executing the process of step S23, the control unit 27 returns to the process of step S11.

[0145] In the process of step S24 as shown in FIG. 13, the speech accumulation unit 32 determines whether it has acquired an event ID and an output instruction from the event detection unit 30. If the speech accumulation unit 32 determines that it has acquired an event ID and an output instruction (step S24: YES), it proceeds to the process of step S25. If the speech accumulation unit 32 determines that it has not acquired an event ID and an output instruction (step S24: NO), it proceeds to the process of step S30.

[0146] In the process of step S25, the replay flag is set to True. This replay flag is set to True by the event detection unit 30 when the event detection unit 30 outputs the output instruction in the process of step S24 to the speech accumulation unit 32.

[0147] In the process of step S26, the speech accumulation unit 32 identifies a speech ID that matches the event ID acquired in the process of step S24 from among the voice sampling data stored in the data buffer 33. The speech accumulation unit 32 acquires the voice sampling data corresponding to the identified speech ID as playback data. The speech accumulation unit 32 starts outputting the playback data from the data buffer 33 to the voice modulation unit 35.

[0148] In the process of step S27, the control unit 27 determines whether all the playback data has been output from the data buffer 33 to the voice modulation unit 35. If the control unit 27 determines that all the playback data has been output (step S27: YES), it proceeds to the process of step S29. If the control unit 27 does not determine that all the playback data has been output (step S27: NO), it proceeds to the process of step S28.

[0149] In the process of step S28, the control unit 27 determines whether a replay stop has been received by receiving an input from the input unit 22 for the area 42 as shown in FIG. 6. If the control unit 27 determines that a replay stop has been received (step S28: YES), it proceeds to the process of step S29. If the control unit 27 does not determine that a replay stop has been received (step S28: NO), it returns to the process of step S27.

[0150] In the process of step S29, the control unit 27 sets the replay flag to False. After executing the process of step S29, the control unit 27 returns to the process of step S11 as shown in FIG. 12.

[0151] In the process of step S30, the speech accumulation unit 32 determines whether it has acquired an event ID and a holding instruction from the event detection unit 30. If the speech accumulation unit 32 determines that it has acquired an event ID and a holding instruction (step S30: YES), it proceeds to the process of step S31. If the speech accumulation unit 32 does not determine that it has acquired an event ID and a holding instruction (step S30: NO), the control unit 27 returns to the process of step S11 as shown in FIG. 12.

[0152] In the process of step S31, the speech accumulation unit 32 identifies a speech ID that matches the event ID acquired in the process of step S30 from among the voice sampling data stored in the data buffer 33. The speech accumulation unit 32 outputs the voice sampling data associated with the identified speech ID to the speech holding unit 36 together with the event ID.

[0153] After executing the process of step S31, the control unit 27 returns to the process of step S11 as shown in FIG. 12.

[0154] As described above, in the voice processing apparatus 20, when the control unit 27 detects a voice that satisfies the set conditions, it notifies the user that a voice that satisfies the set conditions has been detected according to the notification conditions. In the present embodiment, when the control unit 27 detects an utterance including a search word as a voice that satisfies the set conditions, it notifies the user that the utterance has been detected according to the priority set for the search word.

[0155] Here, there are cases where the user wants to be preferentially notified that a voice that satisfies the set conditions has been detected according to the content of the voice, and cases where the user does not want to be preferentially notified. In the present embodiment, the user can divide the voices that are preferentially notified that they have been detected and the voices that are not preferentially notified that they have been detected by setting the priority as the notification condition. Therefore, the voice processing apparatus 20 can improve the convenience for the user.

[0156] In addition, when a voice that satisfies the set conditions is detected, if the voice is simply played back, the user may miss the played-back voice. In the voice processing apparatus 20, by notifying the user that a voice has been detected, the possibility that the user misses the played-back voice can be reduced.

[0157] Therefore, according to the present embodiment, an improved voice processing apparatus 20, a voice processing method, and a voice processing system 1 can be provided.

[0158] Furthermore, when the notification condition corresponding to the detected voice satisfies the first condition, the control unit 27 of the voice processing apparatus 20 may notify the user that a voice that satisfies the set conditions has been detected by playing a notification sound. As described above, by playing the notification sound, the user can immediately notice that a voice has been detected.

[0159] Further, when the notification condition corresponding to the detected voice satisfies the first condition, the control unit 27 of the voice processing device 20 may play a voice that satisfies the setting condition after the notification sound is played. With such a configuration, as described above, the convenience of the user can be improved.

[0160] Also, when the notification condition corresponding to the detected voice satisfies the second condition, the control unit 27 of the voice processing device 20 may notify the user that a voice satisfying the setting condition has been detected by presenting visual information to the user. As described above, the second condition has a lower priority for notifying the user than the first condition, the third condition, and the fourth condition. When the priority is low, instead of playing a notification sound, a notification commensurate with the low priority can be made by presenting visual information to the user.

[0161] Also, when the notification condition corresponding to the detected voice satisfies the second condition, the control unit 27 of the voice processing device 20 may present visual information to the user by presenting a notification list to the user. By looking at the notification list, the user can grasp the date and time when the voice was detected and the circumstances under which the voice was detected.

[0162] Also, when the notification condition corresponding to the detected voice satisfies the second condition, the control unit 27 of the voice processing device 20 may play the detected voice based on the user's input. When the priority is low, it is highly likely that the user wants to check the detected voice later. With such a configuration, the convenience of the user can be improved.

[0163] Also, when the notification condition corresponding to the detected voice satisfies the third condition, the control unit 27 of the voice processing device 20 may control such that the notification sound is played immediately after the search word is detected and the playback of the utterance is started. With such a configuration, when the priority is high, the user can immediately check the content of the utterance.

[0164] Further, when the notification condition corresponding to the detected voice satisfies the fourth condition, the control unit 27 of the voice processing device 20 may control such that a notification sound is reproduced and the reproduction of the utterance is started immediately after the utterance ends. By starting the reproduction of the utterance immediately after the utterance ends, the utterance uttered in real time and the reproduced utterance do not overlap. With such a configuration, the user can more accurately grasp the content of the reproduced utterance.

[0165] Further, the control unit 27 of the voice processing device 20 may control such that the voice data in the utterance section including the detected utterance is reproduced. As described above with reference to FIG. 9, the utterance section is a section in which voice data continues without a set time interval. By reproducing the voice data in such an utterance section, a collection of utterances including the search word is reproduced. With such a configuration, the user can understand the meaning of the utterance including the search word.

[0166] Further, regardless of the priority, the control unit 27 of the voice processing device 20 may cause the display unit 23 to display the detection date and time when an event, that is, an utterance, is detected and the search word included in the utterance among the information included in the notification list. With such a configuration, the user can grasp how the detected utterance was uttered.

[0167] (Other Embodiments) A voice processing system 101 as shown in FIG. 14 can provide a monitoring service for a baby or the like. The voice processing system 101 includes a microphone 110 and a voice processing device 20.

[0168] The microphone 110 and the voice processing device 20 are located farther apart from each other than the microphone 10 and the voice processing device 20 as shown in FIG. 1. For example, the microphone 110 and the voice processing device 20 are located in separate rooms. The microphone 110 is located in the room where the baby is. The voice processing device 20 is located in the room where the user is.

[0169] In other embodiments, the setting condition is a condition that it matches the preset voice feature. The user may input the feature of the voice that the user wants to set as the setting condition from the microphone of the input unit 22 of the voice processing device 20 as shown in FIG. 2, and set it as the setting condition in the voice processing device 20. For example, the user sets the feature of the baby's crying voice as the setting condition in the voice processing device 20.

[0170] The sound collector 110 includes a microphone 11, a speaker 12, a communication unit 13, a storage unit 14, and a control unit 15 as shown in FIG. 2. The sound collector 110 may not include the speaker 12.

[0171] The voice processing device 20 may further include a speaker 12 as shown in FIG. 2. The control unit 27 of the voice processing device 20 may further include a voice playback unit 17 and an accumulation unit 18 as shown in FIG. 2.

[0172] In other embodiments, the storage unit 26 as shown in FIG. 2 stores a search list in which data indicating the feature of the voice is associated with the priority instead of the search list as shown in FIG. 3. The data indicating the feature of the voice may be data of a feature amount of the voice that can be processed by the machine learning model used by the voice recognition unit 29. The feature amount of the voice is, for example, a Mel-Frequency Cepstral Coefficient (MFCC) or PLP (Perceptual Linear Prediction). For example, the storage unit 26 stores a search list in which data indicating the baby's crying voice is associated with the priority of "high".

[0173] In other embodiments, the control unit 27 as shown in FIG. 2 detects a voice whose feature matches the preset voice feature as the voice that satisfies the setting condition.

[0174] The voice recognition unit 29 obtains information on the start time of speech, information on the end time of speech, a speech ID, and voice sampling data from the section detection unit 28 in the same or similar manner as in the above-described embodiment. In other embodiments, the voice recognition unit 29 determines whether the features of the voice in the speech section match the preset features of the voice by performing voice recognition processing using a learning model generated by an arbitrary machine learning algorithm.

[0175] When the voice recognition unit 29 determines that the features of the voice in the speech section match the preset features of the voice, it outputs a result indicating a match as a voice recognition result, the speech ID of the speech section, and data indicating the features of the voice to the event detection unit 30. When the features of the voice in the speech section do not match the preset features of the voice, the voice recognition unit 29 outputs a result indicating a mismatch as a voice recognition result and the speech ID of the speech section to the event detection unit 30.

[0176] The event detection unit 30 can obtain a result indicating a match as a voice recognition result, a speech ID, and data indicating the preset features of the voice from the voice recognition unit 29. When the event detection unit 30 obtains a result indicating a match, it detects the voice whose features match the preset features of the voice as an event. When the event detection unit 30 detects an event, it obtains the speech ID obtained from the voice recognition unit 29 as an event ID. Further, the event detection unit 30 refers to the search list and obtains the priority corresponding to the data indicating the features of the voice obtained from the voice recognition unit 29. The event detection unit 30 executes notification processing according to the obtained priority in the same or similar manner as in the above-described embodiment.

[0177] The event detection unit 30 can obtain a result indicating a mismatch as a voice recognition result and a speech ID from the voice recognition unit 29. When the event detection unit 30 obtains a result indicating a mismatch, it obtains the speech ID obtained from the voice recognition unit 29 as a clear event ID. The event detection unit 30 outputs the clear event ID to the speech storage unit 32.

[0178] The processing of the voice processing apparatus 20 according to other embodiments is not limited to the processing described above. As another example, the control unit 27 may construct a classifier capable of classifying multiple types of voices. Further, the control unit 27 may determine which priority the voice collected by the microphone 110 corresponds to based on the result of inputting the voice data collected by the microphone 110 into the constructed classifier.

[0179] Other effects and configurations of the voice processing system 101 according to other embodiments are the same as or similar to those of the voice processing system 1 as shown in FIG. 1.

[0180] Although the present disclosure has been described based on the drawings and examples, it should be noted that those skilled in the art can easily make various modifications or corrections based on the present disclosure. Therefore, it should be noted that these modifications or corrections are included in the scope of the present disclosure. For example, the functions and the like included in each functional unit can be rearranged so as not to be logically contradictory. A plurality of functional units and the like may be combined into one or divided. Each of the embodiments according to the present disclosure described above is not limited to being faithfully implemented in each of the described embodiments, and can be implemented by appropriately combining each feature or omitting a part thereof. That is, those skilled in the art can make various modifications and corrections based on the present disclosure. Therefore, these modifications and corrections are included in the scope of the present disclosure. For example, in each embodiment, each functional unit, each means, each step, etc. can be added to other embodiments so as not to be logically contradictory, or replaced with each functional unit, each means, each step, etc. of other embodiments. Also, in each embodiment, a plurality of each functional unit, each means, each step, etc. can be combined into one or divided. Also, each of the embodiments of the present disclosure described above is not limited to being faithfully implemented in each of the described embodiments, and can also be implemented by appropriately combining each feature or omitting a part thereof.

[0181] For example, the control unit 27 of the voice processing device 20 may detect a plurality of voices that satisfy different setting conditions from one utterance section. In this case, the control unit 27 may notify the user that a voice that satisfies the setting condition has been detected according to each of a plurality of notification conditions set for each of the different setting conditions. Alternatively, the control unit 27 may notify the user that a voice that satisfies the setting condition has been detected according to a part of the plurality of notification conditions set for each of the different setting conditions. A part of the plurality of notification conditions may be notification conditions that satisfy a selection condition. The selection condition is a condition that is pre-selected based on a user operation or the like from among the first condition, the second condition, the third condition, and the fourth condition. Alternatively, a part of the plurality of notification conditions may be notification conditions included up to the Nth (N is an integer of 1 or more) counted from the ones with higher priority for notifying the user determined by the plurality of notification conditions. The Nth may be preset based on a user operation or the like. For example, the control unit 27 may detect a plurality of different search words from one utterance section. In this case, the control unit 27 may execute processing according to each of a plurality of priorities set for each of the plurality of different search words. Alternatively, the control unit 27 may execute processing according to a part of the plurality of priorities set for each of the plurality of different search words. A part of the plurality of priorities may be, for example, priorities included up to the Nth counted from the ones with higher priority.

[0182] For example, the control unit 27 of the voice processing device 20 may detect voices that satisfy the same setting condition a plurality of times from one utterance section. In this case, in one utterance section, the control unit 27 may execute the process of notifying the user according to the notification condition only once, or may execute it the number of times the voice that satisfies the setting condition is detected. For example, the control unit 27 may detect the same search word a plurality of times from one utterance section. In this case, in one utterance section, the control unit 27 may execute the process according to the priority only once, or may execute it the number of times the search word is detected.

[0183] For example, the section detection unit 28 as shown in FIG. 2 may stop detecting the utterance section while the replay flag is set to True.

[0184] For example, in the voice processing system 101 as shown in FIG. 14, the voice feature preset as a setting condition has been described as the feature of a baby's crying sound. However, the voice feature preset as a setting condition is not limited to the feature of a baby's crying sound. Depending on the usage situation of the voice processing system 101, any voice feature may be set as a setting condition. As another example, the feature of a supervisor's voice, an intercom call sound, or a telephone incoming call sound may be set as a setting condition.

[0185] For example, in the above-described embodiment, the priority has been described as being set in three levels including "high", "medium", and "low". However, the priority is not limited to being set in three levels. The priority may be set in a plurality of levels. For example, the priority may be set in two levels or a plurality of levels of four levels or more.

[0186] For example, in the above-described embodiment, the microphone 10 and the voice processing device 20 were described as separate devices. However, the microphone 10 and the voice processing device 20 may be configured as one device. An example of this will be described with reference to FIG. 15. A voice processing system 201 as shown in FIG. 15 includes a microphone 210. The microphone 210 is an earphone. The microphone 210 is configured to execute the processing of the voice processing device 20. That is, the microphone 210 as an earphone serves as the voice processing device of the present disclosure. The microphone 210 includes a microphone 11, a speaker 12, a communication unit 13, a storage unit 14, and a control unit 15 as shown in FIG. 2. The control unit 15 of the microphone 210 includes a component corresponding to the control unit 27 of the voice processing device 20. The storage unit 14 of the microphone 210 stores a notification list. The microphone 210 may execute screen display and light emission as notification means as shown in FIG. 5 using another terminal device such as the user's smartphone. For example, the control unit 15 of the microphone 210 transmits the notification list in the storage unit 14 to the user's smartphone or the like via the communication unit 13 and causes it to be displayed on the user's smartphone or the like.

[0187] For example, in the above-described embodiment, the voice processing device 20 has been described as performing voice recognition processing. However, an external device other than the voice processing device 20 may perform the voice recognition processing. The control unit 27 of the voice processing device 20 may acquire the result of the voice recognition processing executed by the external device. The external device may be, for example, a dedicated computer configured to function as a server, a general-purpose personal computer, or a cloud computing system, etc. In this case, the communication unit 13 of the microphone 10 may be further configured to include at least one communication module that can be connected to an arbitrary network including a mobile communication network and the Internet, etc., in the same or similar manner as the communication unit 21. In the microphone 10, the control unit 15 may transmit the voice sampling data to an external device via the network by the communication unit 13. When the external device receives the voice sampling data from the microphone 10 via the network, it may execute voice recognition processing. The external device may transmit the result of the voice recognition processing to the voice processing device 20 via the network. In the voice processing device 20, the control unit 27 may acquire the result of the voice recognition processing from the external device via the network by receiving it by the communication unit 21.

[0188] For example, in the above-described embodiment, the voice processing device 20 has been described as a terminal device. However, the voice processing device 20 is not limited to a terminal device. As another example, the voice processing device 20 may be a dedicated computer configured to function as a server, a general-purpose personal computer, or a cloud computing system, etc. In this case, the communication unit 13 of the microphone 10 may be further configured to include at least one communication module that can be connected to an arbitrary network including a mobile communication network and the Internet, etc., in the same or similar manner as the communication unit 21. The microphone 10 and the voice processing device 20 may communicate with each other via the network.

[0189] For example, an embodiment is also possible in which a general-purpose computer functions as the voice processing apparatus 20 according to the above-described embodiment. Specifically, a program describing the processing content for realizing each function of the voice processing apparatus 20 according to the above-described embodiment is stored in the memory of a general-purpose computer, and the processor reads and executes the program. Therefore, the configuration according to the above-described embodiment can also be realized as a program executable by a processor or a non-transitory computer-readable medium storing the program.

Explanation of Signs

[0190] 1,101,201 Voice processing system 10,110,210 Microphone 11 Mic 12 Speaker 13 Communication unit 14 Storage unit 15 Control unit 16 Voice acquisition unit 17 Voice playback unit 18 Accumulation unit 20 Voice processing apparatus 21 Communication unit 22 Input unit 23 Display unit 24 Vibration unit 25 Light-emitting unit 26 Storage unit 27 Control unit 28 Section detection unit 29 Voice recognition unit 30 Event detection unit 31 Speech notification unit 32 Speech accumulation unit 33 Data buffer 34 Ring buffer 35 Voice modulation unit 40 Main screen 41,42,43,44 Areas 50 Setting screen 51,52,53,54,55,56 Areas 60 Notification screen 61,62,63,64 Areas

Claims

1. A first step of acquiring a first sound, the first sound being a sound emitted around the user, by a microphone; The first sound acquired by the microphone is output by a speaker, and a second sound is stored in an utterance storage unit. Tep and The control unit, after the first sound is output from the speaker in the second step, outputs the first sound stored in the utterance storage unit as When a notification condition corresponding to the first sound satisfies a first condition, the first sound is automatically output; a third step of outputting the first sound based on an input from the user when the notification condition corresponding to the first sound satisfies a second condition; Includes Sound processing methods.

2. In the third step, the control unit outputs a notification sound before automatically outputting the first sound. The sound processing method according to claim 1 .

3. In the third step, when the notification condition corresponding to the first sound satisfies a second condition, the control unit presents visual information of the first sound. The sound processing method according to claim 1 .

4. The first sound includes a speech sound, In the third step, the control unit outputs the first sound from a start point of the utterance. The sound processing method according to claim 1 .

5. A microphone that captures a first sound that is a sound generated around the user; a speaker that outputs the first sound acquired by the microphone; a speech storage unit that stores the first sound; After the first sound is output from the speaker, the first sound stored in the speech storage unit is When a notification condition corresponding to the first sound satisfies a first condition, the first sound is automatically output; a control unit that outputs the first sound based on an input from the user when a notification condition corresponding to the first sound satisfies a second condition; A sound processing device comprising:

6. A microphone that captures a first sound, the first sound being a sound emitted around a user; a speaker that outputs the first sound acquired by the microphone; a speech storage unit that stores the first sound; A sound processing program for a sound processing device comprising: The sound processing program After the first sound is output from the speaker, the first sound stored in the speech storage unit is When a notification condition corresponding to the first sound satisfies a first condition, the first sound is automatically output; causing the sound processing device to output the first sound based on an input from the user when a notification condition corresponding to the first sound satisfies a second condition; Sound processing program.

Citation Information

Patent Citations

  • Portable music reproducing device

    JP2001256771A

  • Information obtaining apparatus

    JP2009017514A

  • Mobile terminal, mobile system, and warning method

    JP2012074976A

  • Hearing aid

    JP2012134919A

  • Information processing program, information processing method and program

    JP2017069687A