Speech processing system, speech processing method, and speech processing program

The voice processing system addresses the inability to individually control multiple voice devices by using pre-registered sounds to manage mute/unmute and settings, enhancing user convenience.

JP2026072192APending Publication Date: 2026-05-01SHARP KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SHARP KK
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Conventional voice processing systems cannot individually mute or unmute multiple voice devices, leading to low convenience in setting adjustments.

Method used

A voice processing system that acquires input sounds from multiple devices and performs predetermined processing, allowing individual setting changes based on pre-registered sounds, such as tapping patterns, to control mute, unmute, or switch settings like speaker mode.

Benefits of technology

Enables individual control of voice device settings, improving convenience by allowing personalized mute/unmute and setting changes without affecting other devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026072192000001_ABST
    Figure 2026072192000001_ABST
Patent Text Reader

Abstract

The present invention provides a voice processing system, a voice processing method, and a voice processing program that allow for the individual modification of settings for multiple voice devices. [Solution] The audio processing device 1 includes an acquisition processing unit 111 that acquires input sound input to the microphone of audio device 2A among a plurality of audio devices 2, and a setting processing unit 116 that, when the input sound acquired by the acquisition processing unit 111 is a pre-registered registered sound, changes the setting content of a predetermined setting item of audio device 2A to a setting content that has been pre-registered in association with the registered sound, but does not change the setting content of predetermined setting items of audio devices 2B to 2C among the plurality of audio devices 2.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a technique for controlling voices when multiple users conduct conversations using voice devices individually.

Background Art

[0002] Conventionally, a system that enables multiple users to conduct conversations using voice devices each equipped with a microphone and a speaker is known. For example, a system includes a plurality of voice devices (personal communication devices) and a hub device installed in a conference space that allows the plurality of voice devices to be simultaneously connected to a local network, enabling conversations using the voice devices, and capable of setting the microphones of each voice device to mute when a mute button provided on the hub device is pressed (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In a conventional system, in a hub device to which a plurality of voice devices are connected, all the voice devices are collectively muted or unmuted, and it is not possible to individually mute or unmute the plurality of voice devices. Thus, in the conventional technology, there is a problem that it is difficult to individually change various settings for a plurality of voice devices, resulting in low convenience.

[0005] An object of the present disclosure is to provide a voice processing system, a voice processing method, and a voice processing program capable of individually changing settings of a plurality of voice devices.

Means for Solving the Problems

[0006] An audio processing system according to one aspect of this disclosure is a system that acquires input sounds input to the microphones of each of a plurality of audio devices and performs predetermined audio processing. The audio processing system comprises an acquisition processing unit and a setting processing unit. The acquisition processing unit acquires input sounds input to the microphone of a first audio device among the plurality of audio devices. If the input sound acquired by the acquisition processing unit is a registered sound that has been registered in advance, the setting processing unit changes the setting content of a predetermined setting item of the first audio device to a setting content that has been registered in advance in association with the registered sound, but does not change the setting content of the predetermined setting item of the other audio devices among the plurality of audio devices except for the first audio device.

[0007] Another aspect of the present disclosure relates to a sound processing method that acquires input sounds input to the microphones of multiple sound devices and performs predetermined sound processing. The sound processing method is performed by one or more processors, which acquires input sounds input to the microphone of a first sound device among the multiple sound devices, and, if the acquired input sound is a pre-registered sound, changes the setting content of a predetermined setting item of the first sound device to a setting content that has been pre-registered in association with the said registered sound, while not changing the setting content of the predetermined setting item of the other sound devices among the multiple sound devices except for the first sound device.

[0008] Another aspect of the present disclosure relates to a speech processing program that acquires input sounds input to the microphones of multiple speech devices and performs predetermined speech processing. The speech processing program causes one or more processors to perform the following actions: acquire input sounds input to the microphone of a first speech device among the multiple speech devices; and, if the acquired input sound is a pre-registered sound, change the setting content of a predetermined setting item of the first speech device to a setting content that has been pre-registered in association with the said registered sound, while not changing the setting content of the predetermined setting item of the other speech devices among the multiple speech devices except for the first speech device. [Effects of the Invention]

[0009] According to this disclosure, it is possible to provide a voice processing system, a voice processing method, and a voice processing program that can individually change the settings of multiple voice devices. [Brief explanation of the drawing]

[0010] [Figure 1] Figure 1 shows an example of the application of the voice processing system according to the embodiment of this disclosure. [Figure 2] Figure 2 is a block diagram showing the configuration of the voice processing system according to the embodiment of this disclosure. [Figure 3] Figure 3 is a schematic diagram illustrating an example of the application of the audio processing device according to the present disclosure. [Figure 4] Figure 4 shows an example of a configuration information registration list stored in an audio processing device according to the embodiment of this disclosure. [Figure 5] Figure 5 is a flowchart illustrating an example of the procedure for voice control processing performed in the voice processing device according to Embodiment 1 of this disclosure. [Figure 6] Figure 6 is a flowchart illustrating an example of the procedure for voice control processing performed in the voice processing device according to Embodiment 2 of this disclosure. [Figure 7] Figure 7 is a flowchart illustrating an example of the procedure for voice control processing performed in the voice processing device according to Embodiment 3 of this disclosure. [Figure 8] Figure 8 is a flowchart illustrating an example of the procedure for voice control processing performed in the voice processing device according to Embodiment 4 of this disclosure. [Figure 9] Figure 9 is a flowchart illustrating an example of the procedure for voice control processing performed in the voice processing device according to Embodiment 5 of this disclosure. [Figure 10] Figure 10 shows an example of a user information registration list stored in the voice processing device according to Embodiment 5 of this disclosure. [Modes for carrying out the invention]

[0011] The embodiments of this disclosure will be described below with reference to the attached drawings. Note that the following embodiments are merely examples of the embodiments of this disclosure and do not limit the technical scope of this disclosure.

[0012] The voice processing system disclosed herein can be applied, for example, to a case where multiple users in the same space (e.g., a conference room) use voice devices equipped with microphones and speakers to converse (conference) with users in other spaces. Furthermore, the voice processing system can also be applied to a case where multiple users in one space converse using separate voice devices. Additionally, the voice processing system can be applied to a case where one user in one space converses with users in other spaces using a voice device.

[0013] Figure 1 shows an example of the application of the voice processing system 100 according to this embodiment. As shown in Figure 1, users A to D participate in a meeting in conference room R1, and other users (not shown) participate in a meeting in conference room R2. Users A to D each use neckband-type voice devices 2A to 2D that can be worn around the neck to communicate. Users in conference room R2 may use voice device 2, or they may use a single microphone speaker device installed in conference room R2. Voice devices 2A to 2D may be the same model or different models. Also, voice devices 2A to 2D may be voice devices equipped only with a microphone and not with a speaker. Furthermore, voice devices 2A to 2D may be known general-purpose voice devices. For example, voice device 2 may be a pin-type, gooseneck-type, handheld-type, or desktop-type microphone device.

[0014] Each voice device 2 in conference room R1 is wirelessly connected (connected via Bluetooth (registered trademark)) to the voice processing device 1. The voice input to the microphone of each voice device 2 is output (played back) from the speaker of the voice device 2 (or microphone speaker device) of the user in conference room R2 via the voice processing device 1, the conference terminal 3, and the conference server 4. Also, the voice input to the microphone of the voice device 2 (or microphone speaker device) in conference room R2 is played back from the speaker of the voice device 2 of each user in conference room R1 via the conference server 4, the conference terminal 3, and the voice processing device 1.

[0015] In this way, the voice processing system 100 is a system that enables a plurality of users to individually use voice devices 2 to conduct conversations in the same space (conference room R1 in FIG. 1). The voice processing system 100 may include a display device ⑤ that can be used in the conference. On the display device ⑤, conference information such as the camera images of conference participants and conference materials is displayed by a conference application, or the recognition result (text information) obtained by converting voice into text by voice recognition processing is displayed.

[0016] As shown in FIG. 1, the voice processing system 100 includes a voice processing device 1, voice devices 2, a conference terminal 3, and a conference server 4. The voice device 2 is a wireless connection type acoustic device equipped with a microphone and a speaker. Note that the voice device 2 may have functions such as an AI speaker or a smart speaker. The voice processing system 100 includes a plurality of voice devices 2 and is a system that transmits and receives voice data of the user's spoken voice between the plurality of voice devices 2. The voice processing system 100 is an example of the voice processing system of the present disclosure.

[0017] The voice processing device 1 controls the voice (input voice, output voice, etc.) of the voice device 2, and for example, when a meeting starts in a meeting room, it executes a process of transmitting and receiving voice with a plurality of voice devices 2. For example, the voice processing device 1 controls the voice of a plurality of voice devices 2 arranged in the same space. Further, the voice processing device 1 accumulates the voice acquired from the voice device 2 as recording voice, or executes a process (voice recognition process) of converting the acquired voice into text. Note that the voice processing device 1 alone may constitute the voice processing system of the present disclosure.

[0018] Further, the voice processing system of the present disclosure may include various servers that provide various services such as a meeting service, a subtitle (transcription) service by voice recognition, a translation service, a minutes of meeting service, etc. In the present embodiment, it includes a meeting server 4 that provides a meeting service. The meeting server 4 provides an online meeting service of a meeting application that is one of general-purpose software. For example, the meeting application is installed in the meeting terminal 3. By starting and logging in to the meeting terminal 3, it becomes possible to hold an online meeting (for example, an online meeting in meeting rooms R1 and R2) using the meeting application.

[0019] [Voice processing device 1] As shown in FIG. 2, the voice processing device 1 is a device including a control unit 11, a storage unit 12, a communication unit 13, etc. For example, the voice processing device 1 is connected to a plurality of voice devices 2 and is composed of a device (for example, a mixer box) having a function of mixing or splitting the voice input from the plurality of voice devices 2 or the meeting terminal 3.

[0020] The communication unit 13 is a communication unit for connecting the voice processing device 1 to a communication network by wire or wirelessly and executing data communication according to a predetermined communication protocol with external devices such as the voice device 2 and the meeting terminal 3 via the communication network. For example, the communication unit 13 executes pairing processing by the Bluetooth method to wirelessly connect to each voice device 2.

[0021] The storage unit 12 is a non-volatile storage unit such as an HDD (Hard Disk Drive), SSD (Solid State Drive), or flash memory that stores various types of information. Specifically, the storage unit 12 may store data such as information that can identify the audio device 2 (device number, device ID, etc.).

[0022] Furthermore, the storage unit 12 stores control programs such as a voice control program (an example of a voice processing program in this disclosure) that causes the control unit 11 to execute the voice control processing described later (see Figures 5 to 9). For example, the voice control program may be non-temporarily recorded on a computer-readable recording medium such as a CD or DVD, read by a reading device (not shown) such as a CD drive or DVD drive provided by the voice processing device 1, and stored in the storage unit 12.

[0023] The control unit 11 includes control devices such as a CPU, ROM, and RAM. The CPU is a processor that performs various arithmetic operations. The ROM is a non-volatile memory unit that stores control programs such as a BIOS and OS in advance to cause the CPU to perform various arithmetic operations. The RAM is a volatile or non-volatile memory unit that stores various information and is used as a temporary memory (work area) for the various processes performed by the CPU. The control unit 11 controls the audio processing device 1 by executing various control programs that have been stored in advance in the ROM or memory unit 12 using the CPU.

[0024] Specifically, as shown in Figure 2, the control unit 11 includes various processing units such as an acquisition processing unit 111, an audio processing unit 112, an audio recognition processing unit 113, an audio output processing unit 114, a text output processing unit 115, and a setting processing unit 116. The control unit 11 functions as these various processing units by executing various processes according to the control program using the CPU. Some or all of these processing units may be composed of electronic circuits. The control program may also be a program that causes multiple processors to function as these processing units.

[0025] Figure 3 schematically shows an example of speech processing when the speech processing device 1 is applied to a conference. For example, when a conference starts and a user speaks, the acquisition processing unit 111 acquires the spoken voice (voice Va) input to the microphone of the user's voice device 2. The acquisition processing unit 111 performs routing processing to output voice Va to a predetermined output destination. Here, the acquisition processing unit 111 outputs the acquired voice Va (voice data) to the speech processing unit 112 for generating conference audio and to the speech recognition processing unit 113 for converting the voice to text.

[0026] The audio processing unit 112 performs audio processing on the audio Va acquired by the acquisition processing unit 111 in order to play the audio from the speaker. Specifically, the audio processing unit 112 performs at least one of the following on the audio Va: echo cancellation (EC processing), noise cancellation (NC processing), and gain adjustment (AGC processing). The audio processing unit 112 outputs the processed audio (audio Va1) to the audio output processing unit 114.

[0027] The speech recognition processing unit 113 performs speech recognition processing to convert speech to text based on the speech Va acquired by the acquisition processing unit 111. The speech recognition processing unit 113 converts speech to text using a predetermined speech recognition engine (trained model). The speech recognition engine is generated by learning from various speech data (training data), and this speech data has not undergone any speech processing such as echo cancellation, noise cancellation, or gain adjustment. The speech processing device 1 is equipped with the speech recognition engine.

[0028] Thus, the speech recognition processing unit 113 may perform the speech recognition processing based on the speech Va that has not undergone the aforementioned speech processing. The speech recognition processing unit 113 outputs the recognition result (text information Ta1) of the speech recognition processing on the speech Va to the text output processing unit 115.

[0029] The audio output processing unit 114 outputs the audio Va1, which has been processed by the audio processing unit 112, to the conference terminal 3 in conference room R1. The text output processing unit 115 outputs the recognition result (text information Ta1) of the speech recognition processing by the speech recognition processing unit 113 to the conference terminal 3 in conference room R1.

[0030] When conference terminal 3 in conference room R1 receives audio Va1 from the audio processing device 1, it outputs audio Va1 to conference server 4 (see Figure 1). When conference server 4 receives audio Va1 from conference terminal 3 in conference room R1, it outputs audio Va1 to conference terminal 3 in conference room R2. Conference terminal 3 in conference room R2 outputs audio Va1 to the users in conference room R2.

[0031] Furthermore, when the conference terminal 3 in conference room R1 receives text information Ta1 from the audio processing device 1, it displays the text information Ta1 on the display device 5 (see Figure 1). In another embodiment, the conference terminal 3 may store the text information Ta1 to create meeting minutes. Alternatively, the conference terminal 3 may display the text information Ta1 on each user's user terminal (not shown).

[0032] As described above, the voice processing device 1 acquires the user's spoken voice input to each voice device 2 and enables conversations between users (such as online meetings) by sending and receiving audio with the conference terminal 3 and each voice device 2. Here, the voice processing device 1 has a function that allows the settings of each voice device 2 to be changed individually. Specific examples of configurations for realizing this function (Examples 1 to 5) are described below.

[0033] [Example 1] In the audio processing device 1 according to Embodiment 1, the setting processing unit 116 sets the mute of a predetermined audio device 2. Specifically, the setting processing unit 116 sets whether or not to erase, block or stop (mute) the output to the outside of the audio input from the predetermined audio device 2 (whether to enable or disable the mute function). For example, when the acquisition processing unit 111 acquires a specific sound input to the microphone of audio device 2A from audio device 2A, the setting processing unit 116 sets the microphone of audio device 2A to mute (enables the mute function). As a result, the audio input from the microphone of audio device 2A (for example, the speech of user A) is erased. For example, if user A, who is wearing audio device 2A, taps the microphone of audio device 2A twice in a row with their finger, two consecutive tapping sounds are input to the microphone of audio device 2A. When the setting processing unit 116 determines that the sound acquired by the acquisition processing unit 111 from audio device 2A is two consecutive tapping sounds, it sets the microphone of audio device 2A to mute.

[0034] Figure 5 shows an example of the procedure for voice control processing performed by the control unit 11 of the voice processing device 1 according to Embodiment 1.

[0035] This disclosure can be understood as a voice control method (voice processing method of this disclosure) that performs one or more steps included in the voice control process. Furthermore, one or more steps included in the voice control process described herein may be omitted as appropriate. In addition, the execution order of each step in the voice control process may differ to the extent that similar effects are produced. Furthermore, although this description uses the case in which the control unit 11 executes each step in the voice control process as an example, in other embodiments, one or more processors may distribute and execute each step in the voice control process. The same applies to the voice control processes in Examples 2 to 5 described later.

[0036] First, in step S11, the control unit 11 (setting processing unit 116) determines whether or not sound has been input to the microphone of any of the audio devices 2. That is, the control unit 11 determines whether or not it has acquired input sound from any of the audio devices 2 to the microphone. If the control unit 11 has acquired input sound (S11: Yes), it proceeds to step S12. If the control unit 11 waits until it acquires input sound (S11: No).

[0037] In step S12, the control unit 11 (setting processing unit 116) detects the waveform of the input sound.

[0038] In step S13, the control unit 11 (setting processing unit 116) determines whether the detected waveform of the input sound matches a pre-registered waveform (a registered waveform corresponding to mute).

[0039] Here, the memory unit 12 stores a setting information registration list D1. Figure 4 is an example of the setting information registration list D1. The setting information registration list D1 registers a predetermined sound waveform (registered waveform) in association with the setting content of the sound device 2. For example, ID "0001" is registered in association with a waveform representing two consecutive taps and the setting content for mute. ID "0002" is registered in association with a waveform representing three consecutive taps and the setting content for unmute (see [Example 2] below). ID "0003" is registered in association with a waveform representing four consecutive taps and the setting content for speaker setting mode (see [Example 3] and [Example 5] below). ID "0004" is registered in association with a waveform representing five consecutive taps and the setting content for full mute (see [Example 4] below). For example, the administrator of voice device 2 pre-registers combinations of registered waveforms and settings. The registered contents of setting information registration list D1 may be notified to each user of voice device 2.

[0040] In step S13, the control unit 11 determines whether the waveform of the detected input sound matches the waveform corresponding to two consecutive taps registered in the setting information registration list D1 (the registered waveform corresponding to muting). If the control unit 11 determines that the waveform of the detected input sound matches the waveform corresponding to two consecutive taps (S13: Yes), it proceeds to step S14. On the other hand, if the control unit 11 determines that the waveform of the detected input sound does not match the waveform corresponding to two consecutive taps (S13: No), it proceeds to step S11.

[0041] In step S14, the control unit 11 (setting processing unit 116) sets the microphone of the audio device 2 that acquired the input sound to mute. For example, if user A of audio device 2A taps the microphone of audio device 2A twice in a row with their finger, the control unit 11 sets the microphone of audio device 2A to mute. In this case, the control unit 11 does not change the settings of the microphones of the other audio devices 2B to 2D. However, if user B of audio device 2B taps the microphone of audio device 2B twice in a row with their finger, the control unit 11 sets the microphone of audio device 2B to mute. In this way, the control unit 11 can individually set the mute setting for each audio device 2 according to the user's operation.

[0042] Furthermore, the control unit 11 (setting processing unit 116) notifies the user by sound when the setting of a predetermined setting item of the audio device 2 is changed. For example, when the control unit 11 sets the microphone of audio device 2A to mute, it causes the speaker of audio device 2A to output information (such as a voice message) indicating that the microphone has been muted. In this case, the control unit 11 does not output information indicating that the microphone of audio device 2A has been muted from the speakers of the other audio devices 2B to 2D.

[0043] Furthermore, the control unit 11 (audio output processing unit 114) does not output the input sound to an external device if the input sound is a registered sound registered in the setting information registration list D1. For example, the control unit 11 cancels two consecutive tapping sounds input to the microphone of the audio device 2A using a noise canceller and does not output them to the conference terminal 3 and each audio device 2. This prevents unwanted sounds (tapping sounds) from being played back to the other party in the conference (conference room R2).

[0044] [Example 2] In the audio processing device 1 according to Embodiment 2, the setting processing unit 116 unmutes a predetermined audio device 2. For example, when the acquisition processing unit 111 acquires a specific sound input to the microphone of audio device 2A from audio device 2A, the setting processing unit 116 unmutes the microphone of audio device 2A (sets the mute function to be disabled). For example, if user A, who is wearing audio device 2A which is set to mute, taps the microphone of audio device 2A three times in a row with their finger, three consecutive tapping sounds are input to the microphone of audio device 2A. When the setting processing unit 116 determines that the sound acquired by the acquisition processing unit 111 from audio device 2A is three consecutive tapping sounds, it unmutes the microphone of audio device 2A.

[0045] Figure 6 shows an example of the procedure for voice control processing performed by the control unit 11 of the voice processing device 1 according to Embodiment 2.

[0046] First, in step S21, the control unit 11 (setting processing unit 116) determines whether or not sound has been input to the microphone of any of the audio devices 2. That is, the control unit 11 determines whether or not it has acquired input sound from any of the audio devices 2 to the microphone. If the control unit 11 has acquired input sound (S21: Yes), it proceeds to step S22. If the control unit 11 waits until it acquires input sound (S21: No).

[0047] In step S22, the control unit 11 (setting processing unit 116) detects the waveform of the input sound.

[0048] In step S23, the control unit 11 (setting processing unit 116) determines whether the detected waveform of the input sound matches a pre-registered waveform (a registered waveform corresponding to unmuting).

[0049] Specifically, the control unit 11 determines whether the waveform of the detected input sound matches the waveform corresponding to three consecutive taps registered in the setting information registration list D1 (see Figure 4) (the registered waveform corresponding to unmuting). If the control unit 11 determines that the waveform of the detected input sound matches the waveform corresponding to three consecutive taps (S23: Yes), it moves the process to step S24. On the other hand, if the control unit 11 determines that the waveform of the detected input sound does not match the waveform corresponding to three consecutive taps (S23: No), it moves the process to step S21.

[0050] In step S24, the control unit 11 (setting processing unit 116) unmutes the microphone of the audio device 2 that acquired the input sound. For example, if user A of audio device 2A taps the microphone of audio device 2A three times in a row with their finger, the control unit 11 unmutes the microphone of audio device 2A. In this case, the control unit 11 does not change the settings of the microphones of the other audio devices 2B to 2D. However, if user B of audio device 2B, whose microphone is muted, taps the microphone of audio device 2B three times in a row with their finger, the control unit 11 unmutes the microphone of audio device 2B. In this way, the control unit 11 can individually unmute each audio device 2 in response to user operations.

[0051] Furthermore, similar to Embodiment 1, when the microphone of audio device 2A is unmuted, the control unit 11 outputs information (such as voice) from the speaker of audio device 2A indicating that the microphone has been unmuted, but does not output this information from the speakers of the other audio devices 2B to 2D. In addition, the control unit 11 cancels three consecutive tapping sounds input to the microphone of audio device 2A using a noise canceller and does not output them to the conference terminal 3 and each audio device 2.

[0052] [Example 3] In the voice processing device 1 according to Embodiment 3, the setting processing unit 116 switches a predetermined voice device 2 to speaker setting mode and assigns a user name to the voice device 2 (microphone) in speaker setting mode. For example, when the acquisition processing unit 111 acquires a specific sound input to the microphone of voice device 2A from voice device 2A, the setting processing unit 116 switches voice device 2A to speaker setting mode. For example, if user A wearing voice device 2A taps the microphone of voice device 2A four times in a row with their finger, four consecutive tapping sounds are input to the microphone of voice device 2A. When the setting processing unit 116 determines that the sound acquired by the acquisition processing unit 111 from voice device 2A is four consecutive tapping sounds, it switches voice device 2A to speaker setting mode.

[0053] Furthermore, in speaker setting mode, the setting processing unit 116 sets the username to be assigned to the microphone of the voice device 2. The assigned username is displayed in association with the text when, for example, the speech-recognized text is displayed. Specifically, after switching to speaker setting mode, the setting processing unit 116 sets the text of the spoken voice acquired from the voice device 2A by the acquisition processing unit 111 as the username of the voice device 2A.

[0054] Figure 7 shows an example of the procedure for voice control processing performed by the control unit 11 of the voice processing device 1 according to Embodiment 3.

[0055] First, in step S31, the control unit 11 (setting processing unit 116) determines whether or not sound has been input to the microphone of any of the audio devices 2. That is, the control unit 11 determines whether or not it has acquired input sound from any of the audio devices 2 to the microphone. If the control unit 11 has acquired input sound (S31: Yes), it proceeds to step S32. If the control unit 11 waits until it acquires input sound (S31: No).

[0056] In step S32, the control unit 11 (setting processing unit 116) detects the waveform of the input sound.

[0057] In step S33, the control unit 11 (setting processing unit 116) determines whether the detected waveform of the input sound matches a pre-registered waveform (a registered waveform corresponding to the speaker setting mode).

[0058] Specifically, the control unit 11 determines whether the waveform of the detected input sound matches the waveform corresponding to four consecutive taps registered in the setting information registration list D1 (see Figure 4) (registered waveform corresponding to the speaker setting mode). If the control unit 11 determines that the waveform of the detected input sound matches the waveform corresponding to four consecutive taps (S33: Yes), it moves the process to step S34. On the other hand, if the control unit 11 determines that the waveform of the detected input sound does not match the waveform corresponding to four consecutive taps (S33: No), it moves the process to step S31.

[0059] In step S34, the control unit 11 (setting processing unit 116) switches the voice device 2 that acquired the input sound to speaker setting mode. For example, if user A of voice device 2A taps the microphone of voice device 2A four times in a row with their finger, the control unit 11 switches voice device 2A to speaker setting mode. In this case, the control unit 11 does not change the settings of the other voice devices 2B to 2D. If user B of voice device 2B taps the microphone of voice device 2B four times in a row with their finger, the control unit 11 switches voice device 2B to speaker setting mode. In this way, the control unit 11 can individually switch each voice device 2 to speaker setting mode in response to user operation.

[0060] Furthermore, similar to Embodiment 1, when the audio device 2A is switched to speaker setting mode, the control unit 11 outputs information (such as voice) indicating that the audio device 2A has been switched to speaker setting mode, and does not output this information from the speakers of the other audio devices 2B to 2D. In addition, the control unit 11 cancels four consecutive tapping sounds input to the microphone of audio device 2A using a noise canceller and does not output them to the conference terminal 3 and each audio device 2.

[0061] In step S35, the control unit 11 (setting processing unit 116) determines whether or not it has acquired sound from the voice device 2A in speaker setting mode. Specifically, if the control unit 11 acquires the spoken voice input to the microphone of the voice device 2A within a predetermined time (S35: Yes, S37: No), it proceeds to step S36. On the other hand, if the control unit 11 does not acquire the spoken voice from the voice device 2A within a predetermined time (S35: No, S37: Yes), that is, if a predetermined time has elapsed without acquiring the spoken voice, it cancels the speaker setting mode and proceeds to step S31.

[0062] In step S36, the control unit 11 (setting processing unit 116) assigns the text of the spoken voice as the username to the voice device 2A, which has entered speaker setting mode. For example, if user A of the voice device 2A speaks "Tanaka" in speaker setting mode, the control unit 11 assigns "Tanaka" to the voice device 2A.

[0063] The speaker setting mode is executed sequentially in each audio device 2, for example, before the start of a meeting. For example, after user A registers the username of audio device 2A, user B taps the microphone of audio device 2B four times in a row to switch to speaker setting mode, and after switching to speaker setting mode, speaks "Suzuki", the control unit 11 assigns "Suzuki" to audio device 2B. Similarly, the control unit 11 assigns usernames to audio devices 2C and 2D.

[0064] Once the above username assignments are complete and the meeting begins, each user speaks, and their voice is transcribed into text, which is then associated with their username.

[0065] [Example 4] In the audio processing device 1 according to Embodiment 4, the setting processing unit 116 sets all audio devices 2 to mute. For example, when the acquisition processing unit 111 acquires a specific sound input to the microphone of any of the audio devices 2, the setting processing unit 116 sets all the microphones of the audio devices 2 to mute (enables the mute function). As a result, the sound input from the microphones of all audio devices 2 is silenced. For example, if user D wearing audio device 2D taps the microphone of audio device 2D five times in a row with their finger, five consecutive tapping sounds are input to the microphone of audio device 2D. When the setting processing unit 116 determines that the sound acquired by the acquisition processing unit 111 from audio device 2D is five consecutive tapping sounds, it sets the microphones of audio devices 2A to 2D to mute.

[0066] Figure 8 shows an example of the procedure for voice control processing performed by the control unit 11 of the voice processing device 1 according to Embodiment 4.

[0067] First, in step S41, the control unit 11 (setting processing unit 116) determines whether or not sound has been input to the microphone of any of the audio devices 2. That is, the control unit 11 determines whether or not it has acquired input sound from any of the audio devices 2 to the microphone. If the control unit 11 has acquired input sound (S41: Yes), it proceeds to step S42. If the control unit 11 waits until it acquires input sound (S41: No).

[0068] In step S42, the control unit 11 (setting processing unit 116) detects the waveform of the input sound.

[0069] In step S43, the control unit 11 (setting processing unit 116) determines whether the detected waveform of the input sound matches a pre-registered waveform (a registered waveform corresponding to full mute).

[0070] Specifically, the control unit 11 determines whether the waveform of the detected input sound matches the waveform corresponding to five consecutive taps registered in the setting information registration list D1 (see Figure 4) (the registered waveform corresponding to full mute). If the control unit 11 determines that the waveform of the detected input sound matches the waveform corresponding to five consecutive taps (S43: Yes), it moves the process to step S44. On the other hand, if the control unit 11 determines that the waveform of the detected input sound does not match the waveform corresponding to five consecutive taps (S43: No), it moves the process to step S41.

[0071] In step S44, the control unit 11 (setting processing unit 116) mutes the microphones of all audio devices 2. For example, if user D of audio device 2D taps the microphone of audio device 2D five times in a row with their finger, the control unit 11 mutes all microphones of audio devices 2A to 2D. In this way, the control unit 11 can mute all audio devices 2 at once.

[0072] For example, if a user of any of the audio devices 2 taps the microphone of an audio device 2 six times in a row, the control unit 11 may unmute all the microphones of the audio devices 2. In this way, the control unit 11 may unmute all the audio devices 2 at once.

[0073] [Example 5] In the voice processing device 1 according to Embodiment 5, the setting processing unit 116 switches a predetermined voice device 2 to speaker setting mode and assigns a user name to the voice device 2 (microphone) in speaker setting mode. Embodiment 5 is a modification of Embodiment 3. For example, when the acquisition processing unit 111 acquires a specific sound input to the microphone of voice device 2A from voice device 2A, the setting processing unit 116 switches voice device 2A to speaker setting mode. For example, if user A wearing voice device 2A taps the microphone of voice device 2A four times in a row with their finger, four consecutive tapping sounds are input to the microphone of voice device 2A. When the setting processing unit 116 determines that the sound acquired by the acquisition processing unit 111 from voice device 2A is four consecutive tapping sounds, it switches voice device 2A to speaker setting mode.

[0074] In speaker setting mode, the setting processing unit 116 sets the username to be assigned to the microphone of the voice device 2 based on the tapping sound to the voice device 2. Specifically, when the acquisition processing unit 111 acquires a predetermined tapping sound from the voice device 2A after switching to speaker setting mode, the setting processing unit 116 sets the username that has been previously associated with the tapping sound as the username for the voice device 2A.

[0075] Figure 9 shows an example of the procedure for voice control processing performed by the control unit 11 of the voice processing device 1 according to Embodiment 5.

[0076] First, in step S51, the control unit 11 (setting processing unit 116) determines whether or not sound has been input to the microphone of any of the audio devices 2. That is, the control unit 11 determines whether or not it has acquired input sound from any of the audio devices 2 to the microphone. If the control unit 11 has acquired input sound (S51: Yes), it proceeds to step S52. If the control unit 11 waits until it acquires input sound (S51: No).

[0077] In step S52, the control unit 11 (setting processing unit 116) detects the waveform of the input sound.

[0078] In step S53, the control unit 11 (setting processing unit 116) determines whether the detected waveform of the input sound matches a pre-registered waveform (a registered waveform corresponding to the speaker setting mode).

[0079] Specifically, the control unit 11 determines whether the waveform of the detected input sound matches the waveform corresponding to four consecutive taps registered in the setting information registration list D1 (see Figure 4) (registered waveform corresponding to the speaker setting mode). If the control unit 11 determines that the waveform of the detected input sound matches the waveform corresponding to four consecutive taps (S53: Yes), it moves the process to step S54. On the other hand, if the control unit 11 determines that the waveform of the detected input sound does not match the waveform corresponding to four consecutive taps (S53: No), it moves the process to step S51.

[0080] In step S54, the control unit 11 (setting processing unit 116) switches the voice device 2 that acquired the input sound to speaker setting mode. For example, if user A of voice device 2A taps the microphone of voice device 2A four times in a row with their finger, the control unit 11 switches voice device 2A to speaker setting mode. In this case, the control unit 11 does not change the settings of the other voice devices 2B to 2D. If user B of voice device 2B taps the microphone of voice device 2B four times in a row with their finger, the control unit 11 switches voice device 2B to speaker setting mode. In this way, the control unit 11 can individually switch each voice device 2 to speaker setting mode in response to user operations.

[0081] Furthermore, similar to Embodiment 1, when the audio device 2A is switched to speaker setting mode, the control unit 11 outputs information (such as voice) indicating that the audio device 2A has been switched to speaker setting mode, and does not output this information from the speakers of the other audio devices 2B to 2D. In addition, the control unit 11 cancels four consecutive tapping sounds input to the microphone of audio device 2A using a noise canceller and does not output them to the conference terminal 3 and each audio device 2.

[0082] In step S55, the control unit 11 (setting processing unit 116) determines whether or not sound has been input to the microphone of the voice device 2A in speaker setting mode. Specifically, if the control unit 11 acquires input sound from the microphone of the voice device 2A within a predetermined time (S55: Yes, S57: No), it proceeds to step S56. On the other hand, if the control unit 11 does not acquire input sound from the voice device 2A within a predetermined time (S55: No, S57: Yes), it cancels the speaker setting mode and proceeds to step S51.

[0083] In step S56, the control unit 11 (setting processing unit 116) detects the waveform of the input sound.

[0084] In step S58, the control unit 11 (setting processing unit 116) determines whether the detected waveform of the input sound matches a pre-registered registered waveform (a registered waveform corresponding to a user name).

[0085] Here, the memory unit 12 stores a user information registration list D2. Figure 10 shows an example of the user information registration list D2. The user information registration list D2 registers a combination of a predetermined sound waveform (registered waveform) and a username. For example, ID "5001" is registered with a waveform representing two consecutive taps and the username "Tanaka". ID "5002" is registered with a waveform representing three consecutive taps and the username "Suzuki". ID "5003" is registered with a waveform representing two consecutive taps and one tap with a predetermined interval and the username "Sato". ID "5004" is registered with a waveform representing two consecutive taps and two consecutive taps with a predetermined interval and the username "Yamada". For example, the administrator of the sound device 2 pre-registers combinations of registered waveforms and usernames. The contents of user information registration list D2 may be notified to each user of voice device 2.

[0086] In step S58, the control unit 11 determines whether the waveform of the detected input sound matches a waveform registered in the user information registration list D2 (registered waveform). If the control unit 11 determines that the waveform of the detected input sound matches a registered waveform (S58: Yes), it proceeds to step S59. On the other hand, if the control unit 11 determines that the waveform of the detected input sound does not match a registered waveform (S58: No), it proceeds to step S60.

[0087] In step S59, the control unit 11 (setting processing unit 116) assigns a username associated with a registered waveform that matches the waveform of the detected input sound to the voice device 2A in speaker setting mode. For example, in speaker setting mode, if user A of the voice device 2A taps the microphone of the voice device 2A twice in a row with their finger, the control unit 11 assigns "Tanaka" to the voice device 2A.

[0088] The speaker setting mode is executed sequentially in each audio device 2, for example, before the start of a meeting. For example, if user A registers the username for audio device 2A, and then user B taps the microphone of audio device 2B four times in a row to switch to speaker setting mode, and then taps the microphone of audio device 2B three times in a row after switching to speaker setting mode, the control unit 11 assigns "Suzuki" to audio device 2B. Similarly, the control unit 11 assigns usernames to audio devices 2C and 2D.

[0089] Once the above username assignments are complete and the meeting begins, each user speaks, and their voice is transcribed into text, which is then associated with their username.

[0090] The voice processing device 1 executes the voice control processing of Examples 1 to 5 as described above. The voice processing device 1 may have any one of the configurations of Examples 1 to 5, or it may have at least two of the configurations of Examples 1 to 5. By storing the setting information registration list D1 shown in Figure 4 and the user information registration list D2 shown in Figure 10, the voice processing device 1 can execute voice control processing that combines all of Examples 1 to 5.

[0091] As described above, the audio processing system 100 according to this disclosure is a system that acquires input sounds input to each microphone of a plurality of audio devices 2 and performs predetermined audio processing. In the audio processing system 100, the audio processing device 1 acquires the input sound input to the microphone of the first audio device 2 among the plurality of audio devices 2, and if the acquired input sound is a pre-registered registered sound (see Figure 4), it changes the setting content of a predetermined setting item of the first audio device 2 to a setting content (see Figure 4) that has been pre-registered in association with the said registered sound. Furthermore, the audio processing device 1 does not change the setting content of predetermined setting items of the other audio devices 2 among the plurality of audio devices 2, excluding the first audio device 2.

[0092] For example, the audio processing device 1 refers to a storage unit 12 (setting information registration list D1) that stores predetermined sound waveforms and setting contents in a pre-associated manner, and if a waveform matching the acquired input sound waveform is stored, it changes the setting contents of the predetermined setting item of the first audio device 2 to the setting contents associated with that waveform.

[0093] With the above configuration, the settings of the audio device 2 can be individually changed in the audio processing device 1 without performing any setting change operations on the audio device 2 itself. Therefore, even in cases where multiple audio devices 2 of different types, all of which are general-purpose products, are used together for a conference, the settings of each audio device 2 can be changed individually. Thus, it is possible to improve the convenience of the audio devices 2.

[0094] Furthermore, the audio processing device 1 may notify the user by sound when the setting of a predetermined setting item of the first audio device 2 is changed. For example, the audio processing device 1 may be configured to output information indicating that the setting has been changed from the speaker of the first audio device 2, but not from the speaker of the other audio device 2.

[0095] Furthermore, if the audio processing device 1 is configured to output the acquired input sound to an external device (such as a conference terminal 3), it may be configured not to output the input sound to the external device if the input sound is the registered sound. This prevents input sounds for setting changes (such as tapping sounds) from being played externally (to the other party in the conference).

[0096] The input sound may be, for example, a tapping sound against the microphone, but is not limited to this; it may also be the sound of breath blown into the microphone, or a scraping sound against the microphone. Furthermore, the input sound may also be the user's spoken voice (for example, spoken words such as "mute" or "unmute").

[0097] Furthermore, the predetermined setting items are, for example, microphone mute, unmute, and speaker setting mode, but are not limited to these. They may also include setting the volume (high / low) of the playback sound played (output) by the audio device 2, setting the frequency (high / low) of the playback sound played (output) by the audio device 2, or checking the battery level of the audio device 2.

[0098] In the above embodiment, an example was shown in which conference rooms R1 and R2 are connected via a network to conduct an online meeting. However, the audio processing system 100 of this disclosure may consist of only one conference room R1. In this case, for example, in conference room R1, the conference terminal 3 plays the audio input to the microphone of one audio device 2 through the speaker of the other audio device 2, and displays the converted text information of the audio on the display device 5. For example, the audio output processing unit 114 may output the audio processed by the audio processing unit 112 through the speaker of the audio device 2, and the text output processing unit 115 may display the text information, which is the recognition result of the speech recognition processing by the speech recognition processing unit 113, on the display device 5.

[0099] The control unit 11 of the audio processing device 1 controls the entire audio processing device 1. The control unit 11 realizes various functions by reading and executing various programs stored in the storage unit 12 (for example, storage or ROM). The control unit 11 may be realized by one or more control devices / arithmetic units (CPU (Central Processing Unit), SoC (System on a Chip)). The control unit 11 may also be composed of one or more control circuits (electronic circuits).

[0100] [Disclosure Note] The following is an overview of the disclosures extracted from the above-described embodiments. Note that each configuration and processing function described in the following notes can be selected and combined as desired.

[0101] <Note 1> A sound processing system that acquires input sounds input to the microphones of multiple sound devices and performs predetermined sound processing, An acquisition processing unit that acquires the input sound input to the microphone of the first audio device among the plurality of audio devices, If the input sound acquired by the acquisition processing unit is a pre-registered registered sound, the setting processing unit changes the setting content of a predetermined setting item of the first sound device to a setting content that has been pre-registered in association with the said registered sound, and does not change the setting content of the predetermined setting item of the other sound devices among the plurality of sound devices except for the first sound device. A voice processing system equipped with the following features.

[0102] <Note 2> The setting processing unit notifies the first audio device of the change in the setting content of a predetermined setting item by sound when the setting content of the first audio device is changed. The audio processing system described in Appendix 1.

[0103] <Note 3> The setting processing unit outputs information indicating that the setting has been changed from the speaker of the first audio device, but does not output it from the speaker of the other audio device. The audio processing system described in Appendix 2.

[0104] <Note 4> The system includes an output processing unit that outputs the input sound acquired by the acquisition processing unit to an external device. The output processing unit, when the input sound acquired by the acquisition processing unit is the registered sound, does not output the input sound to the external device. The voice processing system described in any of the appendices 1 to 3.

[0105] <Note 5> The input sound is a tapping sound against the microphone. The voice processing system described in any of the appendices 1 to 4.

[0106] <Note 6> The predetermined setting item is to mute or unmute the microphone. The voice processing system described in any of the appendices 1 to 5.

[0107] <Note 7> The predetermined setting item is the registration of the username of the first voice device. A voice processing system as described in any of the appendices 1 to 6.

[0108] <Note 8> The setting processing unit refers to a storage unit that stores predetermined sound waveforms and setting contents in a pre-associated manner, and if a waveform matching the input sound waveform acquired by the acquisition processing unit is stored in the storage unit, it changes the setting contents of the predetermined setting item of the first sound device to the setting contents associated with that waveform. A voice processing system as described in any of the appendices 1 to 7.

[0109] <Note 9> A sound processing method that acquires input sounds input to each microphone of multiple sound devices and performs predetermined sound processing, To acquire the input sound input to the microphone of the first audio device among the aforementioned plurality of audio devices, If the acquired input sound is a pre-registered registered sound, the setting of a predetermined setting item of the first sound device is changed to a setting registered in association with the said registered sound, and the setting of the predetermined setting item of the other sound devices among the plurality of sound devices, excluding the first sound device, is not changed. A method of audio processing performed by one or more processors.

[0110] <Note 10> A sound processing program that acquires input sounds from the microphones of multiple sound devices and performs predetermined sound processing, To acquire the input sound input to the microphone of the first audio device among the aforementioned plurality of audio devices, If the acquired input sound is a pre-registered registered sound, the setting of a predetermined setting item of the first sound device is changed to a setting registered in association with the said registered sound, and the setting of the predetermined setting item of the other sound devices among the plurality of sound devices, excluding the first sound device, is not changed. A speech processing program for causing one or more processors to execute, or a non-temporary computer-readable recording medium on which the speech processing program is recorded. [Explanation of Symbols]

[0111] 100: Voice processing system 1: Audio processing device 2: Audio device 3: Conference terminal 4: Conference Server 5:Display device 11: Control Unit 12: Storage section 13: Communications Department 111: Acquisition Processing Unit 112: Audio Processing Unit 113: Speech Recognition Processing Unit 114: Audio output processing unit 115: Text output processing unit 116: Configuration Processing Unit D1: Configuration Information Registration List D2: User Information Registration List

Claims

1. A sound processing system that acquires input sounds input to the microphones of multiple sound devices and performs predetermined sound processing, An acquisition processing unit that acquires the input sound input to the microphone of the first audio device among the plurality of audio devices, If the input sound acquired by the acquisition processing unit is a pre-registered registered sound, the setting processing unit changes the setting content of a predetermined setting item of the first sound device to a setting content that has been pre-registered in association with the registered sound, and does not change the setting content of the predetermined setting item of the other sound devices among the plurality of sound devices except the first sound device, A voice processing system equipped with the following features.

2. The setting processing unit notifies the first audio device of the change in the setting of a predetermined setting item by sound when the setting of that item is changed. The voice processing system according to claim 1.

3. The setting processing unit outputs information indicating that the setting has been changed from the speaker of the first audio device, but does not output it from the speaker of the other audio device. The voice processing system according to claim 2.

4. The system includes an output processing unit that outputs the input sound acquired by the acquisition processing unit to an external device. The output processing unit, when the input sound acquired by the acquisition processing unit is the registered sound, does not output the input sound to the external device. The voice processing system according to claim 1.

5. The input sound is a tapping sound against the microphone. The voice processing system according to claim 1.

6. The predetermined setting item is to mute or unmute the microphone. The voice processing system according to claim 1.

7. The predetermined setting item is the registration of the username of the first voice device. The voice processing system according to claim 1.

8. The setting processing unit refers to a storage unit that stores predetermined sound waveforms and setting contents in a pre-associated manner, and if a waveform matching the input sound waveform acquired by the acquisition processing unit is stored in the storage unit, it changes the setting contents of the predetermined setting item of the first sound device to the setting contents associated with that waveform. The voice processing system according to any one of claims 1 to 7.

9. A sound processing method that acquires input sounds input to each microphone of multiple sound devices and performs predetermined sound processing, To acquire the input sound input to the microphone of the first audio device among the aforementioned plurality of audio devices, If the acquired input sound is a pre-registered registered sound, the setting of a predetermined setting item of the first sound device is changed to a setting registered in association with the said registered sound, and the setting of the predetermined setting item of the other sound devices among the plurality of sound devices, excluding the first sound device, is not changed. A method of audio processing performed by one or more processors.

10. A sound processing program that acquires input sounds from the microphones of multiple sound devices and performs predetermined sound processing, To acquire the input sound input to the microphone of the first audio device among the aforementioned plurality of audio devices, If the acquired input sound is a pre-registered registered sound, the setting of a predetermined setting item of the first sound device is changed to a setting registered in association with the said registered sound, and the setting of the predetermined setting item of the other sound devices among the plurality of sound devices, excluding the first sound device, is not changed. A speech processing program that causes one or more processors to execute.

Citation Information

Patent Citations

  • Conference system

    JP2023037813A