Audio processing system, audio processing method, and audio processing program
The audio processing system addresses the challenge of reduced conversation quality in shared spaces by enabling or disabling audio output based on device type, ensuring clear and interference-free communication among multiple users.
Patent Information
- Application Number
- JP2023199606
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2025-06-06
AI Technical Summary
When multiple users use audio devices individually in the same space, issues such as reduced conversation quality arise due to problems like blocked ears making it difficult to hear others, and double audio from overlapping voice playback.
An audio processing system that includes an acquisition processing unit to acquire audio inputs from multiple audio devices and a setting processing unit to enable or disable the output of these audio inputs to speakers of other audio devices, based on the type of each audio device.
The system improves conversation quality by ensuring that users can properly hear each other's voices without interference, regardless of the type of audio device used, thereby enhancing communication in shared spaces.
Smart Images

Figure 2025085904000001_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to a technology for multiple users to converse in the same space, each using an audio device individually. [Background technology]
[0002] Conventionally, there is known a system in which a plurality of users can converse using audio devices each equipped with a microphone and a speaker. For example, there is known a system that includes a plurality of audio devices (personal call devices) and a hub device that is installed in a conference space and allows the plurality of audio devices to be simultaneously connected to a local network, and that constructs a group call network in which the hub device allows simultaneous mutual calls to the connected audio devices (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2023-037813 A Summary of the Invention [Problem to be solved by the invention]
[0004] When multiple users use audio devices individually in the same space (such as a conference room), the following problems arise. For example, when a first user uses an earphone-type audio device, the voice spoken by a second user in the same space is not played back from the speaker of the audio device, and the first user's ears are blocked by the audio device, making it difficult to hear the voice spoken by the second user. In addition, when a first user uses a tabletop audio device, even though the first user can directly hear the voice spoken by the second user in the same space, the voice of the second user is played back from the first user's audio device, resulting in double audio. In this way, when multiple users use audio devices individually in the same space, a problem of reduced conversation quality arises.
[0005] An object of the present disclosure is to provide a voice processing system, a voice processing method, and a voice processing program that can improve conversation quality when multiple users are conversing in the same space, each using an individual voice device. [Means for solving the problem]
[0006] An audio processing system according to one aspect of the present disclosure is a system for controlling audio from multiple audio devices arranged in the same space. The audio processing system includes an acquisition processing unit and a setting processing unit. The acquisition processing unit acquires a first audio input to a microphone of a first audio device. The setting processing unit enables or disables a function for outputting the first audio from a speaker of one or more other second audio devices for each of the second audio devices.
[0007] An audio processing method according to another aspect of the present disclosure is a method for controlling audio from multiple audio devices arranged in the same space, in which one or more processors acquire a first audio input into a microphone of a first audio device, and enable or disable, for each of the second audio devices, a function for outputting the first audio from a speaker of one or more other second audio devices.
[0008] An audio processing program according to another aspect of the present disclosure is a program for controlling audio from multiple audio devices arranged in the same space, and for causing one or more processors to acquire a first audio input into a microphone of a first audio device, and enable or disable a function for outputting the first audio from a speaker of one or more other second audio devices, for each of the second audio devices. Effect of the Invention
[0009] According to the present disclosure, it is possible to provide a voice processing system, a voice processing method, and a voice processing program that can improve the quality of a conversation when multiple users are conversing in the same space, each using an individual voice device. [Brief description of the drawings]
[0010] [Figure 1] FIG. 1 is a diagram illustrating an application example of a voice processing system according to an embodiment of the present disclosure. [Diagram 2] FIG. 2 is a diagram illustrating a configuration of a voice processing system according to an embodiment of the present disclosure. [Figure 3A] FIG. 3A is a diagram illustrating an example of loopback registration information used in the voice processing system according to an embodiment of the present disclosure. [Figure 3B] FIG. 3B is a diagram illustrating an example of loopback registration information used in the voice processing system according to an embodiment of the present disclosure. [Figure 4] FIG. 4 is a diagram showing an example of a loopback registration screen displayed on the voice processing system according to an embodiment of the present disclosure. [Figure 5A] FIG. 5A is a diagram illustrating an example of loopback registration information used in the voice processing system according to an embodiment of the present disclosure. [Figure 5B] FIG. 5B is a diagram illustrating an example of loopback registration information used in the voice processing system according to an embodiment of the present disclosure. [Figure 6A] FIG. 6A is a diagram illustrating an example of usage type registration information used in the voice processing system according to an embodiment of the present disclosure. [Figure 6B] FIG. 6B is a diagram illustrating an example of usage type registration information used in the voice processing system according to an embodiment of the present disclosure. [Figure 7] FIG. 7 is a diagram showing an example of a usage type registration screen displayed on the voice processing system according to an embodiment of the present disclosure. [Figure 8A] FIG. 8A is a diagram illustrating an example of usage type registration information used in the voice processing system according to an embodiment of the present disclosure. [Figure 8B] FIG. 8B is a diagram illustrating an example of usage type registration information used in the voice processing system according to an embodiment of the present disclosure. [Figure 9A]FIG. 9A is a diagram illustrating an example of device type registration information used in the voice processing system according to an embodiment of the present disclosure. [Figure 9B] FIG. 9B is a diagram illustrating an example of device type registration information used in the voice processing system according to an embodiment of the present disclosure. [Figure 10] FIG. 10 is a diagram illustrating an example of a device type registration screen displayed on the voice processing system according to an embodiment of the present disclosure. [Figure 11A] FIG. 11A is a diagram illustrating an example of device type registration information used in the voice processing system according to an embodiment of the present disclosure. [Figure 11B] FIG. 11B is a diagram illustrating an example of device type registration information used in the voice processing system according to an embodiment of the present disclosure. [Figure 12] FIG. 12 is a diagram illustrating an application example of the voice processing device according to the first embodiment of the present disclosure. [Figure 13] FIG. 13 is a flowchart illustrating an example of a procedure of a voice control process executed in the voice processing device according to the first embodiment of the present disclosure. [Figure 14] FIG. 14 is a diagram illustrating an application example of the voice processing device according to the second embodiment of the present disclosure. [Figure 15] FIG. 15 is a flowchart illustrating an example of a procedure of a voice control process executed in a voice processing device according to the second embodiment of the present disclosure. [Figure 16] FIG. 16 is a diagram illustrating an application example of a voice processing device according to a third embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0011] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. Note that the following embodiments are examples of the present disclosure and are not intended to limit the technical scope of the present disclosure.
[0012] The voice processing system according to the present disclosure can be applied to a case where, for example, multiple users in the same space (e.g., a conference room) use audio devices each equipped with a microphone and a speaker to have a conversation (conference) with a user in another space. The voice processing system can also be applied to a case where multiple users in one space use their own audio devices to have a conversation.
[0013] FIG. 1 shows an application example of the voice processing system 100 according to this embodiment. As shown in FIG. 1, users A to D participate in a conference in a conference room R1, and other users (not shown) participate in the conference in a conference room R2. User A uses a neckband-type audio device 2A that can be worn around the neck, user B uses an earphone-type (inner-type) audio device 2B that can be inserted into the ears, user C uses a tabletop-type audio device 2C that is placed on a desk, and user D uses a headset-type audio device 2D that can be worn to cover the entire ears, to carry out a conversation. The user in conference room R2 may use the audio device 2, or may use one microphone speaker device installed in conference room R2.
[0014] Each audio device 2 in the conference room R1 is wirelessly connected (connected by Bluetooth (registered trademark)) to the audio processing device 1, and audio input to the microphone of each audio device 2 is output (played) from the speaker of the audio device 2 (or microphone speaker device) of the user in the conference room R2 via the conference terminal 3 and the conference server 4. Also, audio input to the microphone of the audio device 2 (or microphone speaker device) in the conference room R2 is output from the speaker of the audio device 2 of each user in the conference room R1 via the conference server 4 and the conference terminal 3.
[0015] In this way, the voice processing system 100 is a system that enables multiple users to have a conversation in the same space (conference room R1 in FIG. 1 ) each individually using the voice device 2. The voice processing system 100 may also include a display device 5 that can be used in a conference. On the display device 5, a conference application in the conference server 4 displays conference information such as camera images of the conference participants and conference materials.
[0016] As shown in FIG. 1, the voice processing system 100 includes a voice processing device 1, a voice device 2, a conference terminal 3, and a conference server 4. The voice device 2 is a wirelessly connected audio device equipped with a microphone and a speaker. The microphone of the voice device 2 may be omitted. The voice device 2 may have functions such as an AI speaker or a smart speaker. The voice processing system 100 includes a plurality of voice devices 2, and is a system that transmits and receives voice data of a user's speech between the plurality of voice devices 2. The voice processing system 100 is an example of a voice processing system of the present disclosure.
[0017] The voice processing device 1 controls the voice (input voice, output voice, etc.) of the voice device 2, and executes a process of transmitting and receiving voice between the voice device 2 when a conference starts in a conference room, for example. For example, the voice processing device 1 controls multiple voice devices 2 arranged in the same space. Note that the voice processing device 1 alone may constitute the voice processing system of the present disclosure. When the voice processing system of the present disclosure is constituted by the voice processing device 1 alone, the voice processing device 1 may store the voice acquired from the voice device 2 as voice for recording, or execute a process of recognizing the acquired voice within the device itself (voice recognition process).
[0018] The voice processing system of the present disclosure may also include various servers that provide various services such as a conference service, a subtitling service using voice recognition, a translation service, and a meeting minutes service. In this embodiment, the system includes a conference server 4 that provides a conference service. The conference server 4 provides an online meeting service of a conference application, which is a type of general-purpose software. For example, the conference application is installed in a conference terminal 3. By starting up the conference terminal 3 and logging in, it becomes possible to hold an online meeting using the conference application (for example, an online conference in conference rooms R1 and R2).
[0019] [Speech processing device 1] As shown in FIG. 2, the voice processing device 1 is a device including a control unit 11, a storage unit 12, a communication unit 13, and the like.
[0020] For example, the audio processing device 1 is configured with a device (for example, a mixer box) having a function of transmitting and receiving audio and a function of mixing or splitting audio.
[0021] The communication unit 13 is a communication unit for connecting the audio processing device 1 to a communication network in a wired or wireless manner and for executing data communication in accordance with a predetermined communication protocol with external devices such as the audio device 2 and the conference terminal 3 (see FIG. 1) via the communication network. For example, the communication unit 13 executes pairing processing by Bluetooth to wirelessly connect to each audio device 2.
[0022] The storage unit 12 is a non-volatile storage unit such as a hard disk drive (HDD), a solid state drive (SSD), or a flash memory that stores various information. Specifically, the storage unit 12 stores data such as loopback registration information D1 regarding the setting registration of the loopback playback function for each audio device 2 connected to the audio processing device 1, usage type registration information D2 regarding the usage type for each audio device 2, and device type registration information D3 regarding the type for each audio device 2.
[0023] FIG. 3A shows an example of the loopback registration information D1. As shown in FIG. 3A, the loopback registration information D1 includes information such as a "connection ID", a "device name", a "loopback playback function", and a "registration date". The connection ID is identification information used when connecting an audio device 2, and is, for example, a Bluetooth address. The device name is the device name of the audio device 2. Instead of the device name, a model name or a type name may be registered. The loopback playback function is a function for outputting a sound input to a microphone of one audio device 2 from a speaker of another audio device 2 in a plurality of audio devices 2 arranged in the same space. In the loopback registration information D1, setting information of the loopback playback function ("disabled", "enabled") is registered for each audio device 2. In addition, the registration date of the setting information of the loopback playback function is registered in the loopback registration information D1.
[0024] Here, the loopback playback function is registered by, for example, a user's operation. For example, in a loopback registration screen P1 shown in FIG. 4, the user selects whether to enable ("Yes") or disable ("No") the loopback playback function for the audio device 2 connected to the connection port of the audio processing device 1. Note that the control unit 11 displays the registered information on the loopback registration screen P1 for the audio device 2 for which "enabled" or "disabled" has already been registered. As shown in FIG. 4, when an audio device 2 ("audio device E") is newly connected to the connection port "6", the control unit 11 accepts a selection operation of "Yes" or "No" from the user. The control unit 11 registers the setting information of the loopback playback function ("disabled" or "enabled") in the loopback registration information D1 based on the user's selection operation. FIG. 3A shows the loopback registration information D1 before the information of "audio device E" is registered, and FIG. 3B shows the loopback registration information D1 after the information of "audio device E" is registered.
[0025] In Figures 3A and 3B, the loopback registration information D1 is configured to manage the setting information of the loopback playback function by a connection ID, but in another embodiment, as shown in Figures 5A and 5B, the loopback registration information D1 may be configured to manage the setting information of the loopback playback function by a device name.
[0026] FIG. 6A shows an example of the usage type registration information D2. As shown in FIG. 6A, the usage type registration information D2 includes information such as "connection ID", "device name", "usage type", and "registration date". The usage type is information indicating a type of a usage method of the audio device 2, and includes, for example, a closed type that covers the ears when used and an open type that does not cover the ears when used. For example, the usage type of the neckband type audio device 2A and the tabletop type audio device 2C is "open type" because the user's ears are not covered when used. In contrast, the usage type of the earphone type (inner type) audio device 2B and the headset type audio device 2D is "closed type" because the user's ears are covered when used. In the usage type registration information D2, the usage type ("open type", "closed type") is registered for each audio device 2. In addition, the registration date of the usage type is registered in the usage type registration information D2.
[0027] Here, the usage type is registered, for example, by a user operation. For example, on the usage type registration screen P2 shown in FIG. 7, the user selects the usage type ("open" or "closed") for the audio device 2 connected to the connection port of the audio processing device 1. Note that the control unit 11 displays the registered information for the audio device 2 whose usage type has already been registered on the usage type registration screen P2. The control unit 11 may also automatically display setting information for the loopback playback function according to the usage type selected by the user. For example, the control unit 11 displays "no loopback" (disabled) when the usage type is "open type", and displays "with loopback" (enabled) when the usage type is "closed type".
[0028] As shown in FIG. 7, when a new audio device 2 ("audio device E") is connected to connection port "6", control unit 11 accepts a selection operation of "open" or "closed" from the user. Based on the user's selection operation, control unit 11 registers the usage type ("open type" or "closed type") in usage type registration information D2. FIG. 6A shows usage type registration information D2 before the information on "audio device E" is registered, and FIG. 6B shows usage type registration information D2 after the information on "audio device E" is registered.
[0029] In Figures 6A and 6B, usage type registration information D2 is configured to manage usage type registration information by connection ID, but in another embodiment, as shown in Figures 8A and 8B, usage type registration information D2 may be configured to manage usage type registration information by device name.
[0030] FIG. 9A shows an example of device type registration information D3. As shown in FIG. 9A, the device type registration information D3 includes information such as "connection ID", "equipment name", "device type", and "registration date". The device type is information indicating the type of audio device 2, and includes, for example, "neckband type", "earphone type", "desktop type", and "headset type". In the device type registration information D3, the device type ("neckband type", "earphone type", "desktop type", and "headset type") is registered for each audio device 2. In addition, the registration date of the device type is registered in the device type registration information D3.
[0031] Here, the device type is registered, for example, by a user operation. For example, on a device type registration screen P3 shown in FIG. 10, the user selects a device type ("neckband type", "earphone type", "desktop type", or "headset type") for the audio device 2 connected to the connection port of the audio processing device 1. Note that the control unit 11 displays registered information on the device type registration screen P3 for the audio device 2 whose device type has already been registered. In addition, the control unit 11 may automatically display setting information for the loopback playback function according to the device type selected by the user. For example, the control unit 11 displays "no loopback" (disabled) when the device type is "neckband type" or "desktop type", and displays "loopback" (enabled) when the device type is "earphone type" or "headset type".
[0032] As shown in Fig. 10, when a new audio device 2 ("audio device E") is connected to connection port "6", control unit 11 accepts a selection operation of "neckband type", "earphone type", "desktop type", or "headset type" from the user. Based on the selection operation by the user, control unit 11 registers the device type ("neckband type", "earphone type", "desktop type", or "headset type") in device type registration information D3. Fig. 9A shows device type registration information D3 before information on "audio device E" is registered, and Fig. 9B shows device type registration information D3 after information on "audio device E" is registered.
[0033] In Figures 9A and 9B, the device type registration information D3 is configured to manage the registration information of the device type by the connection ID, but in another embodiment, as shown in Figures 11A and 11B, the device type registration information D3 may be configured to manage the registration information of the device type by the device name.
[0034] The storage unit 12 also stores control programs such as a voice control program (an example of a voice processing program of the present disclosure) for causing the control unit 11 to execute a voice control process (see FIGS. 13 and 15) described below. For example, the voice control program may be non-temporarily recorded on a computer-readable recording medium such as a CD or DVD, read by a reading device (not shown) such as a CD drive or DVD drive provided in the voice processing device 1, and stored in the storage unit 12.
[0035] The control unit 11 has control devices such as a CPU, a ROM, and a RAM. The CPU is a processor that executes various arithmetic processes. The ROM is a non-volatile storage unit in which control programs such as a BIOS and an OS for causing the CPU to execute various arithmetic processes are stored in advance. The RAM is a volatile or non-volatile storage unit that stores various information, and is used as a temporary storage memory (work area) for various processes executed by the CPU. The control unit 11 controls the audio processing device 1 by executing various control programs stored in advance in the ROM or the storage unit 12 by the CPU.
[0036] Specifically, as shown in Fig. 2, the control unit 11 includes various processing units such as an acquisition processing unit 111, an output processing unit 112, a specific processing unit 113, and a setting processing unit 114. The control unit 11 functions as the various processing units by executing various processes according to the control program with the CPU. Some or all of the processing units may be configured with electronic circuits. The control program may be a program for causing a plurality of processors to function as the processing units.
[0037] Here, the voice processing device 1 may have configurations shown in the following Examples 1 to 3. Specific configurations of each of Examples 1 to 3 will be described below.
[0038] [Example 1] FIG. 12 shows a schematic application example of the voice processing device 1 according to the first embodiment. In FIG. 12, "Va1" indicates a voice that a user A speaks, is input to the microphone of the voice device 2A, and is output from the voice device 2A to the voice processing device 1. "Vb1" indicates a voice that a user B speaks, is input to the microphone of the voice device 2B, and is output from the voice device 2B to the voice processing device 1. "Vc1" indicates a voice that a user C speaks, is input to the microphone of the voice device 2C, and is output from the voice device 2C to the voice processing device 1. "Vd1" indicates a voice that a user D speaks, is input to the microphone of the voice device 2D, and is output from the voice device 2D to the voice processing device 1. "Vb2" indicates a voice (loopback voice) in which the voice Va1, Vc1, or Vd1 is output from the voice processing device 1 to the voice device 2B. "Vd2" indicates a voice (loopback voice) in which the voice Va1, Vb1, or Vc1 is output from the voice processing device 1 to the voice device 2D.
[0039] 12, "V1" indicates the audio output from the audio processing device 1 in conference room R1 to the conference server 4 via the conference terminal 3, and then output from the conference server 4 to the audio processing device 1 in conference room R2. "V2" indicates the audio output from the audio processing device 1 in conference room R2 to the conference server 4, and then output from the conference server 4 to the audio processing device 1 via the conference terminal 3 in conference room R1.
[0040] The acquisition processing unit 111 of the control unit 11 acquires the voice input to the microphone of the audio device 2. In the example shown in Fig. 12, the acquisition processing unit 111 acquires the speech voice (voice Va1) of user A from the audio device 2A, the speech voice (voice Vb1) of user B from the audio device 2B, the speech voice (voice Vc1) of user C from the audio device 2C, and the speech voice (voice Vd1) of user D from the audio device 2D.
[0041] Moreover, the acquisition processing unit 111 acquires, via the network, the voice output from the conference server 4 via the conference terminal 3. For example, the acquisition processing unit 111 acquires the voice V2 uttered by the user in the conference room R2 from the conference server 4 via the conference terminal 3.
[0042] The output processing unit 112 outputs the audio acquired by the acquisition processing unit 111. For example, the output processing unit 112 outputs the audio V1 (audio Va1 to Vd1) acquired from the audio device 2 in the conference room R1 to the conference server 4 via the conference terminal 3. The output processing unit 112 also outputs the audio V2 acquired from the conference server 4 to each audio device 2 in the conference room R1. The output processing unit 112 is set so as not to output the audio input to the microphone of the audio device 2 itself from the speaker of the audio device 2 itself. For example, the audio input from the microphone of the audio device 2A is not output from the speaker of the audio device 2A.
[0043] The identification processing unit 113 identifies the type of the audio device 2. For example, the identification processing unit 113 identifies the use type of the audio device 2 connected to the audio processing device 1 by referring to the information registered in the usage type registration information D2 (see FIG. 6A). Specifically, the identification processing unit 113 determines whether the audio device 2 is a closed type that blocks the user's ears when used, or an open type that does not block the user's ears when used. For example, the identification processing unit 113 identifies the device type of the audio device 2 connected to the audio processing device 1 by referring to the information registered in the device type registration information D3 (see FIG. 9A). Specifically, the identification processing unit 113 determines whether the audio device 2 is a "neckband type", "earphone type", "desktop type", or "headset type". For example, the identification processing unit 113 identifies the setting information ("disabled" or "enabled") of the loopback playback function of the audio device 2 connected to the audio processing device 1 by referring to the information registered in the loopback registration information D1 (see FIG. 3A). Specifically, the specific processing unit 113 determines whether the loopback playback function of the audio device 2 is set to "enabled" or "disabled."
[0044] If the usage type of the audio device 2 connected to the audio processing device 1 is not registered in the usage type registration information D2, the control unit 11 inquires the user about the usage type of the audio device 2 (see FIG. 7), and when the usage type of the audio device 2 is acquired, registers it in the usage type registration information D2. If the device type of the audio device 2 connected to the audio processing device 1 is not registered in the device type registration information D3, the control unit 11 inquires the user about the device type of the audio device 2 (see FIG. 10), and when the device type of the audio device 2 is acquired, registers it in the device type registration information D3. If the setting information ("disabled" or "enabled") of the loopback playback function of the audio device 2 connected to the audio processing device 1 is not registered in the loopback registration information D1, the control unit 11 inquires the user about the setting information of the loopback playback function of the audio device 2 (see FIG. 4), and when the setting information of the loopback playback function of the audio device 2 is acquired, registers it in the loopback registration information D1. The control unit 11 may perform a network search (Internet search) based on the identification information of the audio device 2 to acquire information of the audio device 2 (such as a usage type and a device type) and register the acquired information.
[0045] 12, the identification processing unit 113 identifies the usage type ("open type" or "closed type") of each of the audio devices 2A-2C, for example, by referring to the usage type registration information D2. Here, the identification processing unit 113 identifies the audio devices 2A and 2C as "open type" and the audio devices 2B and 2D as "closed type."
[0046] The setting processing unit 114 enables or disables, for each audio device 2, a function (loopback function) for outputting audio input to a microphone of one audio device 2 from a speaker of one or more other audio devices 2, among multiple audio devices 2 arranged in the same space. Specifically, the setting processing unit 114 enables or disables the loopback function of the audio device 2 based on the type of the audio device 2. That is, the identification processing unit 113 determines whether the audio device 2 is a closed type or an open type, and the setting processing unit 114 enables or disables the loopback function of the audio device 2 based on the result of the determination by the identification processing unit 113 of the type of the audio device 2.
[0047] Here, "enabling the loopback function of the audio device 2" means that the output processing unit 112 causes audio input to the microphone of one audio device 2 to be output from the speaker of one or more other audio devices 2. Also, "disabling the loopback function of the audio device 2" means that the output processing unit 112 does not cause audio input to the microphone of one audio device 2 to be output from the speaker of one or more other audio devices 2.
[0048] For example, the setting processing unit 114 enables the loopback function when the audio device 2 is a closed type, and disables the loopback function when the audio device 2 is an open type. When the setting processing unit 114 enables the loopback function of a specific audio device 2, the output processing unit 112 outputs the audio input to the microphone of another audio device 2 from the speaker of the specific audio device 2. On the other hand, when the setting processing unit 114 disables the loopback function of a specific audio device 2, the output processing unit 112 does not output the audio input to the microphone of another audio device 2 from the speaker of the specific audio device 2.
[0049] In the example shown in FIG. 12, since the audio devices 2B and 2D are closed type, the setting processing unit 114 enables the loopback function of the audio devices 2B and 2D. In contrast, since the audio devices 2A and 2C are open type, the setting processing unit 114 disables the loopback function of the audio devices 2A and 2C. In this case, for example, when the user A speaks, the output processing unit 112 outputs the voice Va1 of the user A to each of the audio devices 2B and 2D (corresponding to the voices Vb2 and Vd2) and causes them to be output (played) from the speakers of the audio devices 2B and 2D. In contrast, the output processing unit 112 does not output the voice Va1 of the user A to the audio device 2C.
[0050] In this way, when audio device 2 is connected, the setting processing unit 114 obtains the usage type that is pre-associated with the identification information of audio device 2, and enables or disables the loopback function of the audio device based on the obtained usage type.
[0051] As a result, users B and D, whose ears are blocked by audio devices 2B and 2D and who have difficulty directly hearing the voice of user A, can hear the voice of user A reproduced from audio devices 2B and 2D, respectively. On the other hand, user C, whose ears are not blocked by audio device 2C, can directly hear the voice of user A, and since the voice of user A (Va1) is not reproduced from audio device 2C, the voice of user A is not heard twice.
[0052] When the audio device 2 is connected, the setting processing unit 114 may obtain a device type previously associated with the identification information of the audio device 2, and enable or disable the loopback function of the audio device based on the obtained device type. When the audio device 2 is connected, the setting processing unit 114 may obtain setting information ("disabled" or "enabled") of the loopback playback function previously associated with the identification information of the audio device 2, and enable or disable the loopback function of the audio device based on the obtained setting information.
[0053] [Voice control processing] FIG. 13 illustrates an example of a procedure of a voice control process executed by the control unit 11 of the voice processing device 1 according to the first embodiment.
[0054] The present disclosure can be understood as a voice control method (voice processing method of the present disclosure) that executes one or more steps included in the voice control process. One or more steps included in the voice control process described here may be omitted as appropriate. The steps in the voice control process may be executed in a different order as long as the same action and effect is achieved. Furthermore, although an example is described here in which the control unit 11 executes each step in the voice control process, in other embodiments, one or more processors may execute each step in the voice control process in a distributed manner.
[0055] Here, an example will be described in which a user in a conference room R1 connects (pairs) the audio device 2 to the audio processing device 1 and participates in the conference.
[0056] First, in step S11, the control unit 11 determines whether or not multiple audio devices 2 are connected. For example, the control unit 11 determines whether or not multiple audio devices 2 are connected in the conference room R1 at the time of starting a conference. If the control unit 11 determines that multiple audio devices 2 are connected (S11: Yes), the control unit 11 shifts the process to step S12. On the other hand, if the control unit 11 determines that multiple audio devices 2 are not connected, that is, that only one audio device 2 is connected (S11: No), the control unit 11 ends the audio control process.
[0057] When multiple audio devices 2 are connected, the control unit 11 executes the following processing of steps S12 to S17 for each audio device 2. Specifically, in step S12, the control unit 11 determines whether or not the device information of the target audio device 2 has been registered in the usage type registration information D2 (see FIG. 6A). If the control unit 11 determines that the device information of the target audio device 2 has been registered in the usage type registration information D2 (S12: Yes), the control unit 11 proceeds to processing of step S15. On the other hand, if the control unit 11 determines that the device information of the target audio device 2 has not been registered in the usage type registration information D2 (S12: No), the control unit 11 proceeds to processing of step S13.
[0058] In step S13, the control unit 11 inquires about device information regarding the target audio device 2. For example, the control unit 11 may perform a network search (internet search) based on the identification information of the audio device 2 acquired when the audio device 2 is connected, or may request the user of the audio device 2 to input device information of the audio device 2. For example, the control unit 11 accepts an operation from the user to select a usage type ("open" or "closed") on the usage type registration screen P2 (see FIG. 7).
[0059] In step S14, the control unit 11 acquires and registers device information of the target audio device 2. For example, the control unit 11 acquires a usage type and registers it in usage type registration information D2 (see FIG. 6B). After registering the device information, the control unit 11 transitions to step S15.
[0060] In step S15, the control unit 11 determines whether the type (usage type) of the target audio device 2 is a closed type. The control unit 11 determines the usage type of the audio device 2 by referring to the usage type registration information D2. If the control unit 11 determines that the target audio device 2 is a closed type (S15: Yes), the control unit 11 proceeds to step S16. On the other hand, if the control unit 11 determines that the target audio device 2 is an open type (S15: No), the control unit 11 proceeds to step S17.
[0061] In step S16, the control unit 11 enables the loopback playback function of the target audio device 2. In the example shown in Fig. 12, since the audio devices 2B and 2D are closed types, the control unit 11 enables the loopback function of the audio devices 2B and 2D.
[0062] On the other hand, in step S17, the control unit 11 disables the loopback playback function of the target audio device 2. In the example shown in Fig. 12, since the audio devices 2A and 2C are of the open type, the control unit 11 disables the loopback function of the audio devices 2A and 2C.
[0063] In this way, the control unit 11 decides whether or not to enable the loopback function for each of the audio devices 2A to 2D. When the setting of the loopback function for each audio device 2 is completed, the control unit 11 controls the output destination of the audio in the conference. For example, the control unit 11 outputs the audio acquired through the audio device 2 by the user speaking in the conference room R1 to the audio devices 2B and 2D that have the loopback function enabled (audio Vb2, Vd2 in FIG. 12), but does not output it to the audio devices 2A and 2C that have the loopback function disabled.
[0064] In another embodiment, the control unit 11 may inquire about the setting of the loopback playback function ("disabled" or "enabled") when the device information of the target audio device 2 is not registered. For example, the control unit 11 accepts an operation from the user to select whether to enable ("Yes") or disable ("No") the loopback playback function on the loopback registration screen P1 (see FIG. 4). The control unit 11 sets the loopback playback function of the target audio device 2 based on the inquiry result (user's selection operation).
[0065] In another embodiment, the control unit 11 may inquire about the device type when the device information of the target audio device 2 is not registered. For example, the control unit 11 accepts an operation from the user to select a device type ("neckband type", "earphone type", "desktop type", or "headset type") on the device type registration screen P3 (see FIG. 10). The control unit 11 sets the loopback playback function of the target audio device 2 based on the inquiry result (user's selection operation).
[0066] The control unit 11 registers the information selected by the user (usage type, device type, and loopback playback function settings), and the next time an audio device 2 is connected, sets the loopback playback function based on the registered information corresponding to that audio device 2.
[0067] As described above, the voice processing device 1 according to the first embodiment acquires a first voice input to a microphone of a first voice device 2, and enables or disables a function (loop-back function) of outputting the first voice from a speaker of one or more other second voice devices 2 for each second voice device 2. In addition, the voice processing device 1 identifies the type of the voice device 2, and enables or disables the loop-back function of the second voice device 2 based on the type of the second voice device 2 (closed type or open type).
[0068] According to the above configuration, it is possible to properly hear the speech of users in the same space regardless of the type of audio device 2 used by the user, thereby improving the quality of conversation.
[0069] As another embodiment, the setting processor 114 may enable or disable the loopback function for each audio device 2 in response to a setting operation by the user.
[0070] As another embodiment, the loopback registration information D1, the usage type registration information D2, and the device type registration information D3 may be created outside the voice processing system 100 and downloaded to the voice processing device 1. Also, the loopback registration information D1, the usage type registration information D2, and the device type registration information D3 may be files created regularly, such as text files or csv files. This allows the information of the voice device 2 to be connected to be stored in advance in the voice processing device 1, thereby making it possible to omit settings at the time of initial connection.
[0071] [Example 2] 14 is a schematic diagram illustrating an application example of the voice processing device 1 according to the embodiment 2. The embodiment 2 is a configuration related to voice control in a case where a voice V2 uttered by a user in a conference room R2 is output to a voice device 2 of a user in a conference room R1.
[0072] For example, earphone-type audio device 2B is inserted into the user's ear, so if the speaker volume is too high, it puts a strain on the ear. Also, if the speaker volume of tabletop audio device 2C is too high, it will disturb other users. Thus, the appropriate volume varies depending on the type of audio device 2 used by the user.
[0073] Therefore, the setting processing unit 114 sets the signal level of the speaker of audio device 2 based on the type of audio device 2. For example, the setting processing unit 114 lowers the signal level of the speaker of audio device 2B by x decibels (xdb) and lowers the signal level of the speaker of audio device 2C by y decibels. Note that the setting processing unit 114 may set the signal level of the speaker of audio device 2B lower than the signal level of the speaker of audio device 2C (for example, x>y).
[0074] Furthermore, the setting processor 114 may change the signal levels of the audio devices 2A and 2D. As another embodiment, the setting processor 114 may set the signal level of each audio device 2 based on a combination of the types of multiple audio devices 2 used in the same space.
[0075] Furthermore, the setting processor 114 may set a signal level for the sound (loopback sound) by the loopback function corresponding to the embodiment 1. For example, the signal level of the sound output to the audio device 2B, which is more sealed out of the audio devices 2B and 2D, may be set to a level lower than the signal level of the sound output to the audio device 2D.
[0076] [Voice control processing] Fig. 15 shows an example of a procedure of a voice control process executed by the control unit 11 of the voice processing device 1 according to the embodiment 2. Steps S21 to S24 shown in Fig. 15 are the same as steps S11 to S14 shown in Fig. 13 of the embodiment 1, and therefore the description thereof will be omitted.
[0077] In step S25, the control unit 11 determines whether the target audio device 2 is an earphone type or a tabletop type. The control unit 11 determines the usage type of the audio device 2 by referring to the usage type registration information D2 (see FIG. 6A). If the control unit 11 determines that the target audio device 2 is an earphone type or a tabletop type (S25: Yes), the control unit 11 shifts the process to step S26. On the other hand, if the control unit 11 determines that the target audio device 2 is a neckband type or a headset type (S25: No), the control unit 11 ends the audio control process.
[0078] In step S26, the control unit 11 reduces the signal level of the speaker of the target audio device 2. In the example shown in Fig. 12, the control unit 11 reduces the signal level of the speaker of the earphone-type audio device 2B by x decibels (xdb) and reduces the signal level of the speaker of the tabletop-type audio device 2C by y decibels.
[0079] In this manner, the control unit 11 sets the speaker signal level for each of the audio devices 2A to 2D. After completing the setting of the speaker signal level for each audio device 2, the control unit 11 controls each signal level for each audio device 2 when outputting audio during a conference.
[0080] According to the configuration of the second embodiment, it is possible to prevent an unnecessarily loud sound from being output, which may cause strain on the ears or cause a nuisance to those around.
[0081] [Example 3] 16 is a schematic diagram illustrating an application example of the voice processing device 1 according to the embodiment 3. The embodiment 3 is configured to control the voice output by the loopback function according to the embodiment 1 and the voice V2 uttered by the user output from the conference room R2.
[0082] For example, if the signal level at which user A's spoken voice (voice Va1) output by the loopback function is input to user B's audio device 2B (voice Vb2) and output from the speaker differs from the signal level at which voice V2 spoken by a user in conference room R2 is input to user B's audio device 2B (voice V2) and output from the speaker, user B will feel uncomfortable due to the difference in volume.
[0083] Therefore, the setting processing unit 114 adjusts the signal level of the audio Vb2 (loopback audio) to be output from the speaker of the audio device 2B to the signal level of the audio V2 (remote audio) to be output from the speaker of the audio device 2B. For example, the setting processing unit 114 may adjust the signal level of the audio Vb2 to approximately match the signal level of the audio V2, or may adjust the signal level of the audio V2 to approximately match the signal level of the audio Vb2.
[0084] According to the above configuration, in the audio device 2, it is possible to improve the quality of audio when outputting audio acquired through different paths (loopback path, network path).
[0085] The above-mentioned first to third embodiments can be appropriately combined. The voice processing device 1 according to the present disclosure may have at least one of the configurations of the first to third embodiments.
[0086] In addition, in the voice processing system 100, the voice processing device 1 and the conference terminal 3 may be configured as a single device. For example, the voice processing device 1 may be configured to have the functions of the conference terminal 3 (such as a conference application) and to be capable of data communication with the conference server 4.
[0087] The voice processing system of the present disclosure may be configured by a single voice processing device 1, or may be configured by a combination of the voice processing device 1, a plurality of voice devices 2, a conference terminal 3, and a conference server 4.
[0088] [Disclosure Note] The following will provide an outline of the disclosure extracted from the above-described embodiment. Note that the configurations and processing functions described in the following supplementary notes can be selected and combined in any desired manner.
[0089] <Appendix 1> An audio processing system for controlling audio of multiple audio devices arranged in the same space, an acquisition processing unit that acquires a first voice input to a microphone of the first audio device; a setting processing unit that enables or disables a function of outputting the first sound from a speaker of one or more other second audio devices for each of the second audio devices; A voice processing system comprising:
[0090] <Appendix 2> A processor for identifying a type of the audio device; The setting processing unit enables or disables the function of the second audio device based on the type of the second audio device. 2. The speech processing system of claim 1.
[0091] <Appendix 3> the specific processing unit determines whether the audio device is a closed type that covers the user's ears when in use, or an open type that does not cover the user's ears when in use, the setting processing unit enables the function when the second audio device is the closed type, and disables the function when the second audio device is the open type. 3. The speech processing system of claim 2.
[0092] <Appendix 4> The setting processing unit sets a signal level of a speaker of the second audio device based on the type of the second audio device. 4. The speech processing system according to claim 2 or 3.
[0093] <Appendix 5> The setting processing unit enables or disables the function of the second audio device in response to a setting operation by a user. 5. A voice processing system according to any one of claims 1 to 4.
[0094] <Appendix 6> When the second audio device is connected, the setting processing unit acquires the type associated with the identification information of the second audio device from a storage unit that associates and stores the identification information of the audio device with the type of the audio device, and enables or disables the function of the second audio device based on the acquired type. 6. A voice processing system according to any one of claims 2 to 5.
[0095] <Appendix 7> When the second audio device is connected and the type of the second audio device is not stored in the storage unit, the type of the second audio device is inquired, and when the type of the second audio device is acquired, the type is stored in the storage unit. 7. The speech processing system of claim 6.
[0096] <Appendix 8> The acquisition processing unit acquires a second voice input to a microphone of an audio device disposed in a space different from the space, The setting processing unit adjusts a signal level of the first sound to be output from a speaker of the second audio device and a signal level of the second sound to be output from a speaker of the second audio device. 8. A voice processing system according to any one of claims 1 to 7. [Explanation of symbols]
[0097] 100: Audio processing system 1: Audio processing device 2: Audio equipment 3: Conference terminal 4: Conference server 5:Display device 11: Control section 12: Storage section 13: Operation display section 14: Communications Department 111: Acquisition processing unit 112: Output processing section 113: Specific processing unit 114: Setting processing section D1: Loopback registration information D2: Usage type registration information D3: Device type registration information
Claims
1. An audio processing system for controlling audio from multiple audio devices arranged in the same space, an acquisition processing unit that acquires a first voice input to a microphone of the first audio device; a setting processing unit that enables or disables a function of outputting the first sound from a speaker of one or more other second audio devices for each of the second audio devices; A voice processing system comprising:
2. A processor for identifying a type of the audio device; The setting processing unit enables or disables the function of the second audio device based on the type of the second audio device. The audio processing system of claim 1 .
3. the specific processing unit determines whether the audio device is a closed type that covers the user's ears when in use, or an open type that does not cover the user's ears when in use, the setting processing unit enables the function when the second audio device is the closed type, and disables the function when the second audio device is the open type. The audio processing system of claim 2 .
4. The setting processing unit sets a signal level of a speaker of the second audio device based on the type of the second audio device. The audio processing system of claim 2 .
5. The setting processing unit enables or disables the function of the second audio device in response to a setting operation by a user. The audio processing system of claim 1 .
6. When the second audio device is connected, the setting processing unit acquires the type associated with the identification information of the second audio device from a storage unit that associates and stores the identification information of the audio device with the type of the audio device, and enables or disables the function of the second audio device based on the acquired type. The audio processing system of claim 2 .
7. When the second audio device is connected and the type of the second audio device is not stored in the storage unit, the type of the second audio device is inquired, and when the type of the second audio device is acquired, the type is stored in the storage unit.
7. The audio processing system of claim 6.
8. The acquisition processing unit acquires a second voice input to a microphone of an audio device disposed in a space different from the space, The setting processing unit adjusts a signal level of the first sound to be output from a speaker of the second audio device and a signal level of the second sound to be output from a speaker of the second audio device. The audio processing system of claim 1 .
9. 1. An audio processing method for controlling audio of multiple audio devices arranged in the same space, comprising: acquiring a first voice input to a microphone of a first audio device; enabling or disabling a function of outputting the first sound from a speaker of one or more other second audio devices for each of the second audio devices; The audio processing method is executed by one or more processors.
10. A sound processing program for controlling sounds from a plurality of sound devices arranged in the same space, acquiring a first voice input to a microphone of a first audio device; enabling or disabling a function of outputting the first sound from a speaker of one or more other second audio devices for each of the second audio devices; An audio processing program for causing one or more processors to execute the above.
Citation Information
Patent Citations
Conference system
JP2023037813A