Audio processing system and audio processing method
The voice processing system uses wearable devices to deliver specific conference information to the facilitator without interrupting the conference, enhancing efficiency by maintaining focus among participants.
Patent Information
- Application Number
- JP2021038028
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-03-10
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2041-03-10
AI Technical Summary
In multi-user conferences, specific information intended for the facilitator is often inadvertently shared with other participants, disrupting their concentration and the conference flow.
A voice processing system with wearable microphone/speaker devices that transmit and receive voice data, including a determination unit to identify conference-affecting conditions and a notification unit to discreetly inform the facilitator of specific information without disturbing others.
Enables targeted information delivery to the facilitator, maintaining conference continuity and participant focus.
Smart Images

Figure 0007752949000001 
Figure 0007752949000002 
Figure 0007752949000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an audio processing system and an audio processing method for transmitting and receiving audio from a microphone speaker device. [Background technology]
[0002] Conventionally, there has been known a system that enables multiple users in different locations (such as conference rooms) to hold a conference (online conference) using terminals such as personal computers. For example, Patent Document 1 discloses a remote conference system in which multiple terminals are connected to each other, in which one of the multiple terminals acts as a moderator terminal and only the moderator terminal can control the speaking rights of the other terminals participating in the remote conference. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2012-217068 Summary of the Invention [Problem to be solved by the invention]
[0004] When multiple users participate in a conference in the same conference room, even if you want to notify only the facilitator (the user who moderates the conference) of specific information (such as the progress of the conference), the information will also be notified to other users in the same conference room. If the specific information is notified to other users in the same conference room in this way, for example, a user who is speaking may lose concentration or stop speaking, causing discomfort to the user.
[0005] An object of the present invention is to provide a voice processing system and a voice processing method that can notify specific information to specific users without interrupting the progress of a conference. [Means for solving the problem]
[0006] One aspect of the present invention provides a voice processing system that includes a plurality of wearable microphone speaker devices that are worn by a plurality of users, and transmits and receives voice data of the users' speech between the plurality of microphone speaker devices. The system comprises: a voice acquisition unit that acquires the voice data from the microphone speaker devices; a voice transmission unit that transmits the voice data acquired by the voice acquisition unit to other microphone speaker devices; a determination processing unit that determines whether or not a predetermined condition is met with respect to an element that affects the progress of the conference; and a notification processing unit that, if the predetermined condition is met, causes a microphone speaker device selected from the plurality of microphone speaker devices in accordance with the element that affects the progress of the conference to notify specific information regarding the predetermined condition.
[0007] Another aspect of the present invention is an audio processing method that includes a plurality of wearable microphone / speaker devices that are worn by a plurality of users, and transmits and receives audio data of the users' speech between the plurality of microphone / speaker devices, in which one or more processors execute an audio acquisition step of acquiring the audio data from one microphone / speaker device, an audio transmission step of transmitting the audio data acquired by the audio acquisition step to another microphone / speaker device, a determination step of determining whether or not a predetermined condition is met with respect to an element that affects the progress of the conference, and a notification step of, if the predetermined condition is met, causing a microphone / speaker device selected from the plurality of microphone / speaker devices in accordance with the element that affects the progress of the conference to notify specific information regarding the predetermined condition. [Effects of the Invention]
[0008] According to the present invention, it is possible to provide a voice processing system and a voice processing method that can notify specific information to specific users without interrupting the progress of a conference. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a diagram showing the configuration of a conference system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a diagram showing an application example of a conference system according to an embodiment of the present invention. [Figure 3] FIG. 3 is an external view showing the configuration of the microphone speaker device according to the embodiment of the present invention. [Figure 4] FIG. 4 is a diagram showing an example of conference information used in the conference system according to the embodiment of the present invention. [Figure 5] FIG. 5 is a diagram showing an example of setting information used in the conference system according to the embodiment of the present invention. [Figure 6] FIG. 6 is a flowchart illustrating an example of a procedure of a conference support process executed in the conference system according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, an embodiment of the present invention will be described with reference to the accompanying drawings. Note that the following embodiment is an example of the present invention, and does not limit the technical scope of the present invention.
[0011] The voice processing system according to the present invention can be applied to a case where, for example, multiple users at two locations (e.g., conference rooms R1 and R2) each use a microphone / speaker device to hold a conference (such as an online conference). The microphone / speaker device has, for example, a neckband shape, and each user wears the microphone / speaker device around their neck to participate in the conference. Each user can hear the sound output from the speaker of the microphone / speaker device, and can also have their own spoken voice collected by the microphone of the microphone / speaker device. The voice processing system according to the present invention can also be applied to a case where multiple users at one location each use a microphone / speaker device to hold a conference.
[0012] [Conference System 100] FIG. 1 is a diagram showing the configuration of a conference system according to an embodiment of the present invention. The conference system 100 includes a voice processing device 1, multiple microphone / speaker devices 2, and a conference server 3. The microphone / speaker device 2 is an audio device equipped with a microphone 24 and a speaker 25. The microphone / speaker device 2 may also have functions such as an AI speaker or a smart speaker. The conference system 100 includes multiple wearable microphone / speaker devices 2 that are worn by multiple users, and is a system that transmits and receives audio data of the users' speech between the multiple microphone / speaker devices 2. The conference system 100 is an example of a voice processing system according to the present invention.
[0013] The conference server 3 executes a conference application that realizes the online conference. The conference server 3 also manages conference information. The audio processing device 1 controls each microphone / speaker device 2, and executes processing to transmit and receive audio to and from each microphone / speaker device 2 when the conference starts. Note that the audio processing device 1 alone may constitute the audio processing system of the present invention.
[0014] In this embodiment, an online conference shown in FIG. 2 will be described as an example. Of users A to H who are participants in the online conference, users A, B, C, and D are located in conference room R1, and users E, F, G, and H are located in conference room R2. Users A to H participate in the conference wearing microphone / speaker devices 2A to 2H around their necks, respectively. An audio processing device 1a and a display DP1 are installed in conference room R1, and an audio processing device 1b and a display DP2 are installed in conference room R2. The displays DP1 and DP2 share their respective screens and display, for example, conference materials. The audio processing device 1a and the display DP1, and the audio processing device 1b and the display DP2 are configured to be able to communicate data via a communication network N1 (for example, the Internet). The audio processing devices 1a and 1b are information processing devices (for example, personal computers) having the same functions. When the audio processing devices 1a and 1b are described in common, they will be referred to as "audio processing device 1."
[0015] Specifically, the conference server 3 uses the Internet communication network N1 to transmit and receive audio data from conference rooms R1 and R2 via the microphone / speaker device 2 and audio processing devices 1a and 1b. For example, when audio processing device 1a acquires data of user A's speech from microphone / speaker device 2A, it transmits the audio data to the conference server 3. The conference server 3 transmits the audio data acquired from audio processing device 1a to audio processing devices 1a and 1b. Audio processing device 1a transmits the audio data acquired from the conference server 3 to microphone / speaker devices 2B to 2D of users B to D, respectively, to output (emit) the speech. Similarly, audio processing device 1b transmits the audio data acquired from the conference server 3 to microphone / speaker devices 2E to 2H of users E to H, respectively, to output (emit) the speech. Furthermore, the conference server 3 accepts user operations and displays conference materials and the like on displays DP1 and DP2. In this way, the conference server 3 realizes an online conference.
[0016] The conference server 3 also stores data such as conference information D1 related to online conferences. FIG. 4 shows an example of the conference information D1. As shown in FIG. 4, the conference information D1 includes, for each conference, information on the conference identification information (conference ID), the location of the conference, the start and end dates and times of the conference, the conference participants, and the materials to be used in the conference. Information corresponding to the online conference shown in FIG. 2 is registered in the conference ID "M001." For example, the organizer of the online conference registers the conference information D1 in advance using his or her own terminal (personal computer). The conference server 3 may be configured as a cloud server.
[0017] [Microphone speaker device 2] FIG. 3 shows an example of the appearance of the microphone speaker device 2. As shown in FIG. 3, the microphone speaker device 2 includes a power supply 22, a connection button 23, a microphone 24, a speaker 25, a communication unit (not shown), and the like. The microphone speaker device 2 is, for example, a neckband-type wearable device that can be worn around the user's neck. The microphone speaker device 2 acquires the user's voice via the microphone 24 and outputs voice to the user from the speaker 25. The microphone speaker device 2 may also include a display unit that displays various types of information.
[0018] As shown in FIG. 3, the main body 21 of the microphone speaker device 2 has left and right arms when viewed from the user wearing the microphone speaker device 2, and is formed in a U-shape.
[0019] The microphone 24 is disposed at the tip of the microphone speaker device 2 so as to easily collect the user's voice. The microphone 24 is connected to a microphone board (not shown) built into the microphone speaker device 2.
[0020] The speakers 25 include a speaker 25L arranged on the left arm and a speaker 25R arranged on the right arm when viewed from the perspective of a user wearing the microphone speaker device 2. The speakers 25L and 25R are arranged near the center of the arms of the microphone speaker device 2 so that the user can easily hear the output sound. The speakers 25L and 25R are connected to a speaker board (not shown) built into the microphone speaker device 2.
[0021] The microphone board is a transmitter board for transmitting audio data to the audio processing device 1 and is included in the communication unit. The speaker board is a receiver board for receiving audio data from the audio processing device 1 and is included in the communication unit.
[0022] The communication unit is a communication interface for wirelessly communicating data between the microphone speaker device 2 and the audio processing device 1 in accordance with a predetermined communication protocol. Specifically, the communication unit connects to and communicates with the microphone speaker device 2 using, for example, Bluetooth (Bluetooth; registered trademark). For example, when a user turns on the power supply 22 and then presses the connection button 23, the communication unit executes pairing processing to connect the microphone speaker device 2 to the audio processing device 1. Note that a transmitter may be disposed between the microphone speaker device 2 and the audio processing device 1, and the transmitter may be paired with the microphone speaker device 2 (Bluetooth connection), and the transmitter and the audio processing device 1 may be connected via the Internet.
[0023] [Speech processing device 1] 1, the voice processing device 1 is a server including a control unit 11, a storage unit 12, an operation display unit 13, a communication unit 14, etc. The voice processing device 1 is not limited to a single computer, but may be a computer system in which multiple computers operate in cooperation. Furthermore, various processes executed by the voice processing device 1 may be executed in a distributed manner by one or multiple processors.
[0024] The communication unit 14 is a communication unit that connects the audio processing device 1 to a communication network N2 by wire or wirelessly and executes data communication in accordance with a predetermined communication protocol with external devices such as the microphone speaker device 2 and displays DP1 and DP2 via the communication network N2. For example, the communication unit 14 executes pairing processing using Bluetooth to connect to the microphone speaker device 2. When an online conference is held, the communication unit 14 connects to the communication network N1 (for example, the Internet) and executes data communication between multiple locations (conference rooms R1 and R2).
[0025] The operation display unit 13 is a user interface that includes a display unit such as a liquid crystal display or an organic EL display that displays various information, and an operation unit such as a mouse, keyboard, or touch panel that accepts operations.
[0026] The storage unit 12 is a non-volatile storage unit such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive) that stores various types of information. Specifically, the storage unit 12 stores data such as setting information D2 of the microphone speaker device 2.
[0027] FIG. 5 shows an example of the setting information D2. As shown in FIG. 5, the setting information D2 includes information such as "device ID," "facilitator," "notification sound," "volume," and "microphone gain." The device ID is identification information of the microphone / speaker device 2, and, for example, a device number is registered. Here, "MS001" to "MS008" correspond to the microphone / speaker devices 2A to 2H, respectively. The facilitator is information indicating whether or not a user is the facilitator (moderator) of the online conference. In the example shown in FIG. 5, it is indicated that user A, who uses the microphone / speaker device 2A corresponding to "MS001," is the facilitator. The notification sound is information such as the volume and type of the notification sound output from the microphone / speaker device 2 of the facilitator. The volume is the volume of each microphone / speaker device 2, and the microphone gain is the microphone gain of each microphone / speaker device 2.
[0028] The user can operate (touch) a setting screen (not shown) displayed on the displays DP1 and DP2 to select and register a facilitator, select and register a notification sound, adjust the volume and microphone gain, etc. The control unit 11 stores setting information D2 in response to user operations.
[0029] The storage unit 12 also stores control programs such as a conference support program for causing the control unit 11 to execute a conference support process (see FIG. 6 ) described below. For example, the conference support program may be non-temporarily recorded on a computer-readable recording medium such as a CD or a DVD, read by a reading device (not shown) such as a CD drive or a DVD drive provided in the audio processing device 1, and stored in the storage unit 12.
[0030] The control unit 11 has control devices such as a CPU, a ROM, and a RAM. The CPU is a processor that executes various arithmetic processes. The ROM is a non-volatile storage unit that pre-stores control programs such as a BIOS and an OS that cause the CPU to execute various arithmetic processes. The RAM is a volatile or non-volatile storage unit that stores various information and is used as a temporary storage memory (work area) for various processes executed by the CPU. The control unit 11 controls the audio processing device 1 by having the CPU execute various control programs pre-stored in the ROM or the storage unit 12.
[0031] However, in conventional technology, when multiple users participate in a conference in the same conference room, even if it is desired to notify only the facilitator (the user who is moderating the conference) of specific information (such as the progress of the conference), the information is also notified to other users in the same conference room. If the specific information is notified to other users in the same conference room in this way, for example, a user who is speaking may lose concentration or stop speaking, causing discomfort to the user. In contrast, the voice processing device 1 according to this embodiment makes it possible to notify specific users of specific information without disrupting the progress of the conference.
[0032] Specifically, as shown in Fig. 1, the control unit 11 includes various processing units such as a setting processing unit 111, an audio acquisition unit 112, an audio transmission unit 113, a determination processing unit 114, and a notification processing unit 115. The control unit 11 functions as the various processing units by executing various processes in accordance with the control program using the CPU. Some or all of the processing units may be configured with electronic circuits. The control program may be a program for causing multiple processors to function as the processing units.
[0033] The setting processing unit 111 performs settings related to the microphone speaker device 2. Specifically, when the microphone speaker device 2 is connected (paired) to the audio processing device 1, the setting processing unit 111 acquires identification information (e.g., a device number) of the microphone speaker device 2 and registers it in the "device ID" of the setting information D2. Furthermore, when the setting processing unit 111 acquires information on the facilitator, notification sound, volume, and microphone gain in response to a user operation, the setting processing unit 111 registers the information in the "facilitator," "notification sound," "volume," and "microphone gain" of the setting information D2. In other words, when a conference is held with multiple users, the setting processing unit 111 sets the microphone speaker device 2 of the facilitator user among the multiple users.
[0034] The voice acquisition unit 112 acquires voice data of the speaker's voice collected by the microphone 24 of the microphone / speaker device 2 from the microphone / speaker device 2. For example, when an online conference starts and facilitator user A speaks, the microphone 24 of the microphone / speaker device 2A collects user A's voice, and the microphone / speaker device 2A transmits the voice data of the voice to the voice processing device 1. The voice acquisition unit 112 acquires the voice data of user A's voice from the microphone / speaker device 2A. The voice acquisition unit 112 acquires voice data from each microphone / speaker device 2. The voice acquisition unit 112 is an example of a voice acquisition unit of the present invention.
[0035] The voice transmitting unit 113 transmits the voice data acquired by the voice acquiring unit 112 to each microphone / speaker device 2. For example, when the voice acquiring unit 112 acquires voice data of the speech of user A, who is the facilitator, from the microphone / speaker device 2A, the voice transmitting unit 113 transmits the voice data to the microphone / speaker devices 2B to 2H other than the microphone / speaker device 2A. The voice data is transmitted to the microphone / speaker device 2 via the conference server 3.
[0036] Note that the voice transmitting unit 113 may transmit the voice data only to microphone / speaker devices 2 in a conference room different from the conference room from which the voice data was transmitted. For example, if the voice acquiring unit 112 acquires voice data of the voice uttered by user A in conference room R1 from microphone / speaker device 2A, the voice transmitting unit 113 may transmit the voice data only to microphone / speaker devices 2E-2H in conference room R2 via the conference server 3. Users B-D in the same conference room R1 as user A can directly hear the voice uttered by user A without going through microphone / speaker devices 2B-2D. The voice transmitting unit 113 is an example of the voice transmitting unit of the present invention.
[0037] The determination processing unit 114 determines whether a predetermined condition for notifying (announcing) specific information from the microphone / speaker device 2 is satisfied. Specifically, the determination processing unit 114 determines whether a predetermined condition is satisfied with respect to factors affecting the progress of the conference. The factors include at least one of the progress status of the conference held by the multiple users, the communication status of sending and receiving the audio data, and the user's speech status. For example, the determination processing unit 114 determines whether a predetermined condition is satisfied with respect to at least one of the progress status of the conference held by the multiple users, the communication status of sending and receiving the audio data, and the user's speech status. For example, the determination processing unit 114 determines whether a predetermined time (e.g., one hour) has elapsed since the conference started. Furthermore, for example, the determination processing unit 114 determines whether the current time is a predetermined time (15 minutes) before the conference end time, i.e., whether the remaining time until the conference end time is within a predetermined time. The predetermined condition is not limited to the above example. Other examples of the predetermined condition will be described later. The determination processing unit 114 is an example of a determination processing unit of the present invention.
[0038] When the predetermined condition is satisfied, notification processing unit 115 causes a predetermined microphone speaker device 2 to notify (announce) specific information related to the predetermined condition. Specifically, when the predetermined condition is satisfied, notification processing unit 115 causes a microphone speaker device 2 selected from multiple microphone speaker devices 2 according to the factors that affect the progress of the conference to notify the user of the microphone speaker device 2 of the specific information related to the predetermined condition.
[0039] For example, when the predetermined condition is satisfied, the notification processing unit 115 outputs a predetermined notification sound from a predetermined microphone / speaker device 2. For example, when a predetermined time (e.g., one hour) has passed since the start of the conference, or when the current time reaches a predetermined time (15 minutes) before the end of the conference, the notification processing unit 115 transmits audio data of the notification sound registered in the setting information D2 to the microphone / speaker device 2A of the facilitator (user A) and outputs the notification sound from the microphone / speaker device 2A. The notification processing unit 115 may output, for example, a buzzer sound (chime sound), or may output a sound indicating that one hour has passed since the start of the conference or that 15 minutes remain until the end of the conference. The notification sound is an example of the specific information of the present invention. The facilitator (user A) is an example of the first user of the present invention. In the above example, the control unit 11 selects the microphone / speaker device 2A of the facilitator, who is the moderator and has an influence on the progress of the conference, as the microphone / speaker device 2 to which the notification sound is to be sent.
[0040] Furthermore, the notification processing unit 115 may output the notification sound from the predetermined microphone speaker device 2 at a volume lower than the volume at which the user's speech is output from the microphone speaker device 2. This prevents the notification sound from being heard by other users.
[0041] According to the above configuration, the speakers 25L and 25R of the microphone speaker device 2A are placed near the ears of user A, so the notification sound can be heard only by user A, without being heard by other users. This allows user A to grasp the progress of the conference (elapsed time, remaining time, etc.) without being noticed by other users. Furthermore, the notification sound does not disturb other users' concentration on the agenda of the conference. This improves the efficiency of the conference.
[0042] If the microphone speaker device 2 has a display unit, the notification processing unit 115 may cause the display unit to display text information corresponding to the voice.
[0043] [Meeting support processing] An example of the procedure of the conference support processing executed by the control unit 11 of the audio processing device 1 will be described below with reference to FIG. 6. Note that the present invention can be understood as a conference support method (audio processing method of the present invention) that executes one or more steps included in the conference support processing. One or more steps included in the conference support processing described here may be omitted as appropriate. The steps in the conference support processing may be executed in a different order as long as the same operational effect is achieved. Furthermore, although an example in which the control unit 11 executes each step in the conference support processing will be described here, in other embodiments, one or more processors may execute each step in the conference support processing in a distributed manner.
[0044] Here, an explanation will be given taking as an example the online conference shown in Fig. 2. The control unit 11 starts the conference support process when the power supply 22 of the microphone speaker device 2 is turned on.
[0045] First, in step S11, the control unit 11 connects each microphone speaker device 2 to the audio processing device 1. For example, when each user participating in a conference presses the connection button 23 of their own microphone speaker device 2, the control unit 11 executes pairing processing with each microphone speaker device 2 using the Bluetooth system, and connects each microphone speaker device 2 to the audio processing device 1.
[0046] Next, in step S12, the control unit 11 registers setting information D2 (see FIG. 4). Specifically, the control unit 11 acquires identification information (e.g., device number) of each microphone / speaker device 2 and registers it in the "device ID" of the setting information D2. Furthermore, the control unit 11 acquires information on the facilitator, notification sound, volume, and microphone gain in response to a user operation, and registers the information in the "facilitator," "notification sound," "volume," and "microphone gain" of the setting information D2. Here, the setting information D2 corresponding to the microphone / speaker devices 2A-2H worn by users A-H participating in the online conference is stored in the storage unit 12. Furthermore, when user A is selected as the facilitator of the online conference, the identification information "1" of the "facilitator" is registered in the device ID "MS001" of user A's microphone / speaker device 2A.
[0047] Next, in step S13, the control unit 11 determines whether the conference has started. If the conference has started (S13: Yes), the process proceeds to step S14. The control unit 11 waits until the conference has started (S13: No). For example, the online conference starts when user A performs an operation to start the online conference. When the online conference starts, the control unit 11 starts measuring time (conference time).
[0048] In step S14, the control unit 11 acquires audio data of the user's speech from the microphone / speaker device 2 and starts a process of transmitting the audio data to the microphone / speaker devices 2 of the other users. For example, when the control unit 11 acquires audio data of the speech of user A, the facilitator, from the microphone / speaker device 2A, it starts a process of transmitting the audio data to the microphone / speaker devices 2B to 2H of users B to H. The audio data is transmitted to the microphone / speaker devices 2E to 2H of users E to H in conference room R2 via the conference server 3. Step S14 is an example of a voice acquisition step and a voice transmission step of the present invention.
[0049] Next, in step S15, the control unit 11 determines whether the conference time (measured time) has exceeded a predetermined time. For example, the control unit 11 determines whether one hour has passed since the conference started. If the conference time has exceeded the predetermined time (S15: Yes), the process proceeds to step S16. The control unit 11 continues the determination process until the conference time has exceeded the predetermined time (S15: No). Step S15 is an example of a determination step of the present invention.
[0050] In another embodiment, the control unit 11 may determine whether the remaining time until the end of the conference is within a predetermined time.
[0051] In step S16, the control unit 11 outputs a predetermined notification sound from a predetermined microphone / speaker device 2. For example, when a predetermined time (e.g., one hour) has passed since the start of the conference, or when the remaining time until the conference end time is within a predetermined time (15 minutes), the control unit 11 transmits the audio data of the notification sound registered in the setting information D2 to the microphone / speaker device 2A of the facilitator (user A), and outputs the notification sound from the microphone / speaker device 2A. The control unit 11 may output a predetermined buzzer sound (chime sound), or may output a predetermined sound (e.g., sound indicating that one hour has passed since the start of the conference or that there are 15 minutes left until the conference end). Step S16 is an example of a notification step of the present invention.
[0052] Next, in step S17, the control unit 11 determines whether the conference has ended. For example, the online conference ends when user A performs an operation to end the online conference. If the online conference has ended (S17: Yes), the control unit 11 ends the conference support process. On the other hand, if the online conference has not ended (S17: No), the process proceeds to step S15.
[0053] Returning to step S15, the control unit 11 determines, for example, whether two hours have passed since the start of the conference, or whether the remaining time until the conference end time is within five minutes. If the conference time has passed two hours, or if the remaining time is within five minutes (S15: Yes), in step S16, the control unit 11 again transmits the audio data of the notification sound to the microphone speaker device 2A of user A, and causes the microphone speaker device 2A to output the notification sound. Note that the control unit 11 may make the volume or type of the second notification sound different from that of the first notification sound. For example, the control unit 11 may increase the volume each time the notification sound is played.
[0054] As described above, the conference system 100 includes a plurality of wearable microphone / speaker devices 2 that are worn by a plurality of users, and transmits and receives audio data of the users' speech between the plurality of microphone / speaker devices 2. The conference system 100 also acquires the audio data from one microphone / speaker device 2 and transmits the acquired audio data to another microphone / speaker device 2. The conference system 100 also determines whether or not predetermined conditions are met regarding at least one of the progress of the conference held by the plurality of users, the communication status for transmitting and receiving the audio data, and the speech status of the user, and if the predetermined conditions are met, causes a predetermined microphone / speaker device 2 to notify specific information regarding the predetermined conditions.
[0055] This allows specific information to be notified only to the user of a specific microphone / speaker device 2. For example, it is possible to notify only the facilitator user of the progress of the conference, the communication status, the speech status, etc., without the participants of the conference noticing. This makes it possible to notify specific information to specific users without interrupting the progress of the conference. This allows the conference to be conducted efficiently.
[0056] The present invention is not limited to the above-described embodiment, and other embodiments of the present invention will be described below.
[0057] In the above-described embodiment, the notification processor 115 outputs the notification sound from a predetermined microphone / speaker device 2 (the facilitator's microphone / speaker device 2) when a predetermined time (e.g., one hour) has elapsed since the start of the conference, or when the remaining time until the conference end time is within a predetermined time (15 minutes). In another embodiment, the predetermined condition may be a condition related to the communication quality of the online conference, such as an error condition, a communication bandwidth condition, a communication speed condition, or a noise condition. Specifically, the determination processor 114 determines whether the communication quality of each user's voice data in an online conference held by multiple users at remote locations has fallen below a predetermined quality level. For example, the determination processor 114 determines whether the communication bandwidth of the Internet communication (communication network N1) for transmitting and receiving voice data between the conference rooms R1 and R2 has fallen below a predetermined bandwidth. The communication bandwidth is, for example, the communication bandwidth between the conference server 3 and the voice processing device 1. When the communication bandwidth falls below the predetermined bandwidth, the notification processor 115 outputs the notification sound from a predetermined microphone / speaker device 2 (the facilitator's microphone / speaker device 2). This allows the facilitator user to recognize that the communication quality has deteriorated, and thus allows the facilitator user to quickly take measures to improve the communication quality without being aware of this deterioration in the communication quality. The notification processing unit 115 may, for example, output a sound indicating that the communication bandwidth has deteriorated. In the above example, when the communication bandwidth between the conference server 3 and the audio processing device 1a falls below the predetermined bandwidth, the control unit 11 selects the microphone / speaker device 2 connected to the audio processing device 1a because the communication bandwidth between the conference server 3 and the audio processing device 1a affects the progress of the conference.
[0058] Furthermore, when the notification processing unit 115 detects a transmission error in audio data, for example, it outputs the notification sound from a predetermined microphone speaker device 2. When the communication speed of the audio data falls below a predetermined speed, for example, it outputs the notification sound from a predetermined microphone speaker device 2. When the noise component of the audio data exceeds a predetermined value, for example, it outputs the notification sound from a predetermined microphone speaker device 2.
[0059] In another embodiment, the predetermined condition may be a condition related to the input time of the microphone input of each microphone / speaker device 2. For example, the determination processing unit 114 determines whether the continuous input time of speech input to the microphone 24 of each microphone / speaker device 2 is equal to or longer than a predetermined time. Then, the notification processing unit 115 outputs the notification sound from the corresponding microphone / speaker device 2 when the continuous input time is equal to or longer than the predetermined time. For example, when user E continuously speaks for equal to or longer than a predetermined time (10 minutes), the notification processing unit 115 outputs the notification sound from the microphone / speaker device 2E of user E. In the above example, the control unit 11 selects the microphone / speaker device 2E of user E as the microphone / speaker device 2 to which the notification sound is to be sent because user E's speech affects the progress of the conference. This allows user E to recognize that he or she has been speaking for a long time without being noticed by other users. The predetermined time may be set according to the attributes of the user. For example, the predetermined time can be set to a long time (15 minutes) for user A, the facilitator in charge of explaining the agenda, and a short time (5 minutes) for other users B to H, who ask questions, etc. This allows for lively and efficient discussions.
[0060] In another embodiment, the notification processing unit 115 may output a notification sound from at least one of the speakers 25L and 25R depending on the importance of the notification content. For example, the control unit 11 sets the importance of the notification sound indicating that the communication quality of an online conference has deteriorated and the notification sound indicating that there are five minutes or less remaining until the end of the conference to "high." The control unit 11 sets the importance of the notification sound indicating that there are 15 minutes or less remaining until the end of the conference and the notification sound indicating that the microphone input time is greater than or equal to a predetermined time to "medium." The control unit 11 sets the importance of the notification sound indicating that one hour has passed since the start of the conference and the notification sound indicating that background noise is greater than or equal to a predetermined value to "low." For example, the notification processing unit 115 outputs the notification sound from both the speakers 25L and 25R when the condition for the importance level is "high." When the condition for the importance level is "medium," the notification sound is output only from the speaker 25L. When the condition for the importance level is "low," the notification sound is output only from the speaker 25R. This allows the target user to recognize the importance of the notification content.
[0061] In this way, notification processing unit 115 may notify specific information corresponding to the importance of the predetermined condition. That is, when the importance of the predetermined condition is a first importance, notification processing unit 115 causes the notification sound to be announced from both speakers 25L and 25R mounted on microphone speaker device 2, and when the importance of the predetermined condition is a second importance lower than the first importance, notification processing unit 115 causes the notification sound to be announced from one of speakers 25L and 25R.
[0062] In another embodiment, the notification processing unit 115 may notify the user by vibrating the microphone speaker device 2 instead of the notification sound. For example, the notification processing unit 115 vibrates a specific microphone speaker device 2 when the predetermined condition is met. The notification processing unit 115 may also change the vibration pattern depending on the importance. For example, the notification processing unit 115 vibrates the microphone speaker device 2 three times at short intervals when the condition for the importance is "high" is met, vibrates the microphone speaker device 2 two times at long intervals when the condition for the importance is "medium" is met, and vibrates the microphone speaker device 2 once when the condition for the importance is "low" is met. The notification processing unit 115 may also continuously vibrate the microphone speaker device 2 until the predetermined condition is no longer met. Note that the microphone speaker device 2 may vibrate a vibrator built into the microphone speaker device 2 based on a signal received from the notification processing unit 115, or may vibrate the speaker unit while outputting a low-pitched notification sound from the speakers 25L and 25R. The vibration is an example of specific information of the present invention.
[0063] In this way, the notification processing unit 115 may output a predetermined notification sound from a predetermined microphone speaker device 2, or may vibrate a predetermined microphone speaker device 2 in a predetermined vibration pattern.
[0064] In another embodiment, the notification processing unit 115 may notify all conference participants of the notification sound according to the predetermined condition. Specifically, the notification processing unit 115 causes all microphone / speaker devices 2 to notify the notification sound when the time remaining until the conference end time is within a predetermined time. For example, when there are five minutes left until the conference end time or when the conference end time has arrived, the notification processing unit 115 causes the microphone / speaker devices 2A to 2H of users A to H to output the notification sound. When notifying all conference participants of the notification sound, the notification processing unit 115 may cause the microphone / speaker device 2A of user A, the facilitator, to output a loud sound like a warning, and the microphone / speaker devices 2B to 2H of the other participants, users B to H, to output soft, gentle sounds.
[0065] In another embodiment, the conference server 3 may have the functions of the audio processing device 1. That is, the conference server 3 acquires audio data from the microphone / speaker devices 2 and transmits the acquired audio data to other microphone / speaker devices 2. The conference server 3 also determines whether or not a predetermined condition is met for an element that affects the progress of the conference, and, if the predetermined condition is met, causes a microphone / speaker device 2 selected from the multiple microphone / speaker devices 2 according to the element that affects the progress of the conference to notify specific information related to the predetermined condition.
[0066] The voice processing system of the present invention may be configured with the voice processing device 1 alone, the conference server 3 alone, or a combination of the voice processing device 1 and the conference server 3.
[0067] Furthermore, the voice processing system of the present invention can be constructed by freely combining the above-described embodiments within the scope of the invention described in each claim, or by appropriately modifying or omitting parts of each embodiment. [Explanation of symbols]
[0068] 1: Audio processing device 2: Microphone speaker device 3: Conference Server 11: Control section 12: Storage section 13: Operation display section 14: Communications Department 21: Main body 22: Power supply 23: Connect button 24:Mike 25: Speaker 25L: Speaker 25R: Speaker 100: Conference system 111: Setting processing section 112: Voice acquisition unit 113: Audio transmission unit 114: Judgment processing unit 115: Notification processing unit
Claims
1. A voice processing system including a plurality of wearable microphone / speaker devices attached to the body of each of a plurality of users participating in a conference in the same room, and transmitting and receiving voice data of the users' speech between the plurality of microphone / speaker devices, a voice acquisition unit that acquires the voice data from a first microphone / speaker device; a voice transmission unit that transmits the voice data acquired by the voice acquisition unit to another microphone / speaker device; a determination processing unit that determines whether or not a continuous input time of a speech voice input to a microphone of the first microphone / speaker device has reached a predetermined time or more; a notification processing unit that, when the continuous input time reaches or exceeds the predetermined time, causes the first microphone / speaker device to notify specific information regarding the continuous input time; Equipped with A voice processing system, wherein the predetermined time is set according to attributes of a user who speaks the speech input to the microphone of the first microphone / speaker device.
2. a setting processing unit that sets the microphone / speaker device of a first user who is a facilitator among the plurality of users when the conference is held; the notification processing unit causes the microphone / speaker device of the first user to notify the specific information when the continuous input time reaches or exceeds the predetermined time. The audio processing system of claim 1 .
3. The notification processing unit outputs a predetermined notification sound from the first microphone / speaker device.
3. The speech processing system according to claim 1 or 2.
4. the notification processing unit outputs the notification sound from the first microphone / speaker device at a volume lower than a volume at which the speech sound is output from the other microphone / speaker device; The audio processing system of claim 3 .
5. the notification processing unit vibrates the first microphone / speaker device in a predetermined vibration pattern; The speech processing system according to any one of claims 1 to 4.
6. The microphone speaker device has a neckband type shape. The speech processing system according to any one of claims 1 to 5.
7. A voice processing method includes a plurality of wearable microphone / speaker devices that are worn by a plurality of users participating in a conference in the same room, and transmits and receives voice data of the users' speech between the plurality of microphone / speaker devices, one or more processors, a voice acquisition step of acquiring the voice data from a first microphone / speaker device; a voice transmission step of transmitting the voice data acquired in the voice acquisition step to another microphone / speaker device; a determination step of determining whether or not a continuous input time of a speech voice input to a microphone of the first microphone / speaker device has reached a predetermined time or more; a notification step of causing the first microphone / speaker device to notify specific information regarding the continuous input time when the continuous input time reaches or exceeds the predetermined time; Run A voice processing method, wherein the predetermined time is set according to attributes of a user who speaks the speech input to the microphone of the first microphone / speaker device.
Citation Information
Patent Citations
Announcement meeting support system, sub-terminal, method for supporting announcement meeting, and method and program for controlling sub-terminal
JP2006134094A
Teleconference system
JP2012217068A
Information processing device used for web conference system, control method thereof, and program
JP2017147670A
Conference system, conference server and program
JP2019220067A
Processing device, program, and processing method
JP2020135556A