Remote conferencing system, control method, and program

JP2026126817APending Publication Date: 2026-08-05KONICA MINOLTA INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
KONICA MINOLTA INC
Filing Date
2025-01-24
Publication Date
2026-08-05

Smart Images

  • Figure 2026126817000001_ABST
    Figure 2026126817000001_ABST
Patent Text Reader

Abstract

This feature eliminates the inconvenience of users having to repeat what they said while muted during remote meetings, thereby improving meeting efficiency. [Solution] The remote conferencing system 1, which allows multiple users to connect and make voice calls simultaneously from different locations, comprises: an audio acquisition unit 32 that acquires the voices emitted by each user; an audio output unit 35 that provides the voices acquired by the audio acquisition unit 32 to other users; a setting unit 37 that sets a mute mode for each user so that the voices emitted by the user are not provided to other users; a speech recording unit 38 that records the content of speeches made by users set to mute mode; and a speech sharing unit 39 that transmits and shares the content of speeches recorded by the speech recording unit 38 to other users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a remote conference system, a control method, and a program.

Background Art

[0002] In recent years, in meetings such as in offices, remote meetings are increasingly being held instead of face-to-face meetings. In a remote meeting, multiple users can make a voice call by connecting to a server simultaneously from different locations. Each user can talk to other users through the speaker and microphone of their PC by connecting their PC to the server via the network. In such a remote meeting, in order to prevent ambient noise and the like from being heard by other users, each user generally sets the mute mode when not speaking themselves and only解除 the mute mode when speaking themselves.

[0003] However, a user may forget to解除 the mute mode and start speaking. Even if the user continues to speak while forgetting to解除 the mute mode, the user's speech will not be transmitted to other users. Therefore, the user解除 the mute mode at the timing when they notice that they have forgotten to解除 it. After解除 the mute mode, the user needs to speak the same content as the previous speech again. When such a case occurs, the time spent on the speech during the mute mode is wasted, and since the user has to speak the same content twice, there is a problem that the efficiency of the meeting is poor.

[0004] On the other hand, conventionally, a conference system that automatically解除 the mute mode has been proposed (for example, Patent Document 1). This conventional conference system takes a face image of a user set in the mute mode with a camera and determines whether the user is speaking based on the face image of the user. When it is determined that the user is speaking, the conference system automatically解除 the mute mode.

[0005] However, while conventional conferencing systems automatically unmute a user when they detect that the user is speaking, they cannot provide other users with the content of the user's speech before the mute was removed. Therefore, even when using conventional conferencing systems, users must respeak the same content they spoke while muted after the mute is removed, which leads to inefficient meetings. [Prior art documents] [Patent Documents]

[0006] [Patent Document 1] Japanese Patent Publication No. 2023-160002 [Overview of the Initiative] [Problems that the invention aims to solve]

[0007] Therefore, the present invention was made to solve the above-mentioned conventional problems. In other words, the present invention aims to provide a remote conferencing system, control method, and program that eliminate the inconvenience of having to repeat what was said while in mute mode, and improve the efficiency of meetings. [Means for solving the problem]

[0008] To achieve the above objective, the invention according to claim 1 is a remote conferencing system that enables multiple users to connect simultaneously from different locations and conduct voice calls, comprising: an audio acquisition unit that acquires the voice emitted by each user; an audio output unit that provides the voice acquired by the audio acquisition unit to other users; a setting unit that sets a mute mode for each user so that the voice emitted by the user is not provided to other users; a speech recording unit that records the content of speeches made by users set to mute mode; and a speech sharing unit that transmits and shares the content of speeches recorded by the speech recording unit to other users.

[0009] The invention according to claim 2 is a remote conferencing system according to claim 1, characterized in that the speech recording unit records the speech acquired by the speech acquisition unit as the content of speech by a user set to mute mode.

[0010] The invention according to claim 3 is a remote conferencing system according to claim 2, characterized in that the speech sharing unit transmits and shares the audio recorded by the speech recording unit to other users.

[0011] The invention according to claim 4 is a remote conferencing system according to claim 2, characterized in that the speech recording unit divides and records the audio file each time the audio acquired by the audio acquisition unit is interrupted.

[0012] The invention according to claim 5 is a remote conferencing system according to claim 4, wherein the speech sharing unit, when multiple audio files are recorded by the speech recording unit, transmits and shares the audio recorded in an audio file selected from among the multiple audio files to other users.

[0013] The invention according to claim 6 is a remote conferencing system according to claim 3 or 5, characterized in that when the speech sharing unit transmits and shares audio with other users, it increases the playback speed of the audio to a speed equal to or greater than the normal speed.

[0014] The invention according to claim 7 is a remote conferencing system according to claim 1, characterized in that the speech recording unit records text information obtained by converting the audio acquired by the audio acquisition unit into text data as the content of speech by a user set to mute mode.

[0015] The invention according to claim 8 is a remote conferencing system according to claim 7, characterized in that the speech sharing unit transmits and shares the text information recorded by the speech recording unit to other users.

[0016] The invention according to claim 9 is a remote conferencing system according to claim 8, characterized in that the speech recording unit divides and records a text file each time the audio acquired by the audio acquisition unit is interrupted.

[0017] The invention according to claim 10 is a remote conferencing system according to claim 9, wherein the speech sharing unit, when multiple text files are recorded by the speech recording unit, transmits and shares a selected text file from among the multiple text files to other users.

[0018] The invention according to claim 11 is a remote conferencing system according to claim 1, wherein the speech recording unit converts the audio acquired by the audio acquisition unit into text data as the content of speech by a user set to mute mode, and records summary information which is a summary of the text contained in the text data.

[0019] The invention according to claim 12 is a remote conferencing system according to claim 1, characterized in that the speech sharing unit transmits and shares the content of speech recorded by the speech recording unit to other users based on an operation by a user set to mute mode.

[0020] The invention according to claim 13 is a remote conferencing system according to claim 12, characterized in that the speech sharing unit transmits and shares the content of speech recorded by the speech recording unit to other users based on the operation of unmuting the mute mode by a user who has been set to mute mode.

[0021] The invention according to claim 14 is a remote conferencing system according to claim 12, wherein the speech sharing unit transmits and shares the content of speech recorded by the speech recording unit to other users based on a sharing instruction operation by a user set to mute mode.

[0022] The invention according to claim 15 is a remote conferencing system according to claim 12, wherein when the speech sharing unit detects an operation by a user set in the mute mode and the voice acquisition unit is acquiring the voice of another user, at the timing when the voice acquired by the voice acquisition unit is interrupted, the speech content recorded by the speech recording unit is transmitted and shared with other users.

[0023] The invention according to claim 16 is a remote conferencing system according to claim 15, wherein at the timing when the speech sharing unit detects an operation by a user set in the mute mode, it notifies other users that the speech content recorded by the speech recording unit is to be shared.

[0024] The invention according to claim 17 is a method for controlling a remote conferencing system in which a plurality of users can simultaneously connect from different locations and conduct a voice call, comprising: a voice acquisition step of acquiring the voice uttered by each user; a voice output step of providing the voice acquired by the voice acquisition step to other users; a setting step of setting a mute mode for each user in which the voice uttered by the user is not provided to other users; a speech recording step of recording the speech content of the user set in the mute mode; and a speech sharing step of transmitting and sharing the speech content recorded by the speech recording step with other users.

[0025] The invention according to claim 18 is a program executed on a server of a remote conference system that enables a plurality of users to simultaneously connect from different locations and conduct a voice call. The server is caused to execute: a voice acquisition step of acquiring the voice uttered by each user; a voice output step of providing the voice acquired in the voice acquisition step to other users; a setting step of setting, for each user, a mute mode in which the voice uttered by the user is not provided to other users; a speech recording step of recording the speech content of the user set in the mute mode; and a speech sharing step of transmitting and sharing the speech content recorded in the speech recording step with other users.

Effect of the Invention

[0026] According to the present invention, it is possible to eliminate the annoyance of having to repeat the content of a user's speech during the mute mode in a remote conference, and improve the efficiency of the conference.

Brief Description of the Drawings

[0027] [Figure 1] It is a diagram showing a configuration example of a remote conference system. <00001​​​​​​​​​​​​​​​​​​​​​​This diagram shows the sequence of actions for sharing user comments. [Figure 10] This diagram shows the message sharing screen when the content of a message is shared as text information or summary information. [Figure 11] This is an example of meeting minutes displayed on a remote meeting screen. [Modes for carrying out the invention]

[0028] Preferred embodiments of the present invention will be described in detail below with reference to the drawings. In the embodiments described below, elements common to all are denoted by the same reference numerals, and redundant explanations of these elements will be omitted.

[0029] (Embodiment of the invention) Figure 1 shows an example configuration of a remote conferencing system 1 in one embodiment of the present invention. This remote conferencing system 1 has a configuration in which a server 2 and a plurality of information processing devices 3a, 3b, and 3c are connected to each other via a network 4 so that they can communicate with one another. The network 4 is a communication network including a LAN (Local Area Network) or the Internet.

[0030] Server 2 is a device that supports remote meetings for multiple users A, B, and C, who are located at different locations A, B, and C. Server 2 is installed on a cloud 5 such as the internet and provides services to support remote meetings by communicating with information processing devices 3a, 3b, and 3c used by each user A, B, and C. For example, Server 2 has the functionality of a web server. Server 2 uses its web server functionality to provide a user interface for remote meetings to information processing devices 3a, 3b, and 3c.

[0031] The information processing devices 3a, 3b, and 3c are composed of, for example, personal computers (PCs), tablet terminals, and smartphones. Figure 2 shows an example configuration of the information processing devices 3a, 3b, and 3c. The information processing devices 3a, 3b, and 3c include a control unit 10, a display unit 11, an operation unit 12, a communication interface 13, a microphone 14, a speaker 15, and a camera 16.

[0032] The display unit 11 is composed of, for example, a color liquid crystal display and displays various screens to the user. The operation unit 12 is equipped with, for example, touchscreen keys located on the screen of the display unit 11 and accepts user input. The operation unit 12 may also be configured to include a keyboard or mouse. The communication interface 13 is an interface that connects the information processing devices 3a, 3b, and 3c to the network 4. The microphone 14 inputs the voice spoken by the user and converts it into audio information. The speaker 15 outputs audio based on the input audio information. The camera 16 captures images of the user's face and generates image data. Note that the information processing devices 3a, 3b, and 3c may not be configured to include the camera 16.

[0033] The control unit 10 consists of a control board on which a processor such as a CPU and memory are mounted, and controls the operation of each part. The processor executes a predetermined program to make the control unit 10 function as a browser 17. The browser 17 is a web browser and accesses the server 2 via the network 4. The browser 17 then retrieves the screen for the remote meeting from the server 2 and displays it on the display unit 11. The browser 17 also drives the microphone 14, speaker 15, and camera 16 based on instructions from the server 2.

[0034] For example, browser 17 acquires audio information based on the user's voice from microphone 14 and sends that audio information to server 2. When browser 17 acquires audio information from server 2, it outputs audio based on that audio information via speaker 15. Furthermore, if camera 16 is capturing an image of the user, browser 17 sends the image data of the user to server 2.

[0035] Figure 3 is a block diagram showing the hardware and functional configuration of Server 2. Server 2 comprises a processor 20, a communication interface 21, and a storage unit 22. The processor 20 is composed of a CPU and other components, and reads and executes computer-readable programs. The communication interface 21 connects Server 2 to the network 4 and communicates with the information processing devices 3a, 3b, and 3c. The storage unit 22 is a non-volatile storage device, such as a hard disk drive (HDD) or solid-state drive (SSD). The storage unit 22 pre-stores programs 23 executed by the processor 20. The storage unit 22 also stores speech record information 24.

[0036] The processor 20 reads program 23 from the memory unit 22 and executes it. As a result, the processor 20 functions as a screen generation unit 30, a screen provision unit 31, an image acquisition unit 32, an audio acquisition unit 33, an audio processing unit 34, an audio output unit 35, an operation detection unit 36, a setting unit 37, a speech recording unit 38, and a speech sharing unit 39.

[0037] The screen generation unit 30 generates a screen for remote conferencing to be provided to each information processing device 3a, 3b, and 3c. The image acquisition unit 32 acquires image data of the user from each information processing device 3a, 3b, and 3c. When the image data is acquired by the image acquisition unit 32, the screen generation unit 30 places the user's image on the screen for remote conferencing. However, if the camera 16's shooting function is turned off in each information processing device 3a, 3b, and 3c, the image acquisition unit 32 does not acquire image data of the user. In that case, the screen generation unit 30 generates a screen that does not include the user's image. The screen provision unit 31 transmits the screen for remote conferencing generated by the screen generation unit 30 to each information processing device 3a, 3b, and 3c.

[0038] The voice acquisition unit 33 acquires the voice of each user detected by the microphone 14 of each information processing device 3a, 3b, and 3c. When the microphone 14 of each information processing device 3a, 3b, and 3c detects a user's voice, it sends voice information based on that voice to the server 2. The voice acquisition unit 33 identifies the information processing devices 3a, 3b, and 3c that are the sources of the voice information and determines which user's voice it is. The voice acquisition unit 33 then outputs the voice information acquired from each information processing device 3a, 3b, and 3c to the voice processing unit 34.

[0039] The voice processing unit 34 distributes the output destination of the voice information acquired by the voice acquisition unit 33 to each user. That is, the voice processing unit 34 distributes the output destination of the voice information acquired by the voice acquisition unit 33 to information processing devices 3a, 3b, and 3c used by users other than the user who made the statement. For example, if the voice information is from a statement made by user A, the voice processing unit 34 decides to output the voice information to information processing devices 3b and 3c used by users B and C. Similarly, if the voice information is from a statement made by user B, the voice processing unit 34 decides to output the voice information to information processing devices 3a and 3c used by users A and C. Furthermore, if the voice information is from a statement made by user C, the voice processing unit 34 decides to output the voice information to information processing devices 3a and 3b used by users A and B.

[0040] When the audio processing unit 34 determines the output destination for the audio information, it instructs the audio output unit 35 to output the audio information. At this time, the audio processing unit 34 specifies the output destination for the audio information to the audio output unit 35. The audio output unit 35 then transmits the audio information to the output destination specified by the audio processing unit 34. In other words, the audio output unit 35 provides the audio spoken by one user so that other users can hear it. For example, if the audio information was spoken by user A, the audio output unit 35 outputs the audio information to the information processing devices 3b and 3c, which were determined as output destinations by the audio processing unit 34. As a result, the information processing devices 3b and 3c output the audio spoken by user A from the speaker 15. Therefore, users B and C can hear the audio spoken by user A.

[0041] Furthermore, the audio processing unit 34 manages the mute mode setting for each user. Mute mode is an operating mode that prevents other users from hearing the voice spoken by a user. Mute mode can be set for each user. For example, if mute mode is set for user A, the audio processing unit 34 does not determine the output destination of the audio information acquired from user A's information processing unit 3a. Therefore, audio information based on the voice spoken by user A is not output to information processing units 3b and 3c. Consequently, the voice spoken by user A while mute mode is set cannot be heard by other users B and C.

[0042] Figure 4 shows an example of screen G1 for remote conferencing displayed in information processing devices 3a, 3b, and 3c. Figure 4(a) shows screen G1 when mute mode is set for user A. This screen G1 includes image display areas Ra, Rb, and Rc that display images of multiple users A, B, and C participating in the remote conferencing. For example, if each user A, B, and C has turned off the camera 16's capture function, their respective usernames will be displayed in image display areas Ra, Rb, and Rc. Also, when server 2 detects a user speaking, the frame of the image display area Rc of the speaking user is highlighted.

[0043] Additionally, icons 51, 52, and 53 that the user can operate are displayed at the bottom of the remote meeting screen G1. Icon 51 is an operation button for setting or disabling mute mode. In the example in Figure 4(a), mute mode is set, so icon 51 is accompanied by an image 55 indicating that mute mode is active. Icon 52 is an operation button for turning the camera 16's shooting function on or off. Figure 4(a) illustrates the case where the camera 16's shooting function is off. Therefore, icon 52 is accompanied by an image 55 indicating that the camera 16's shooting function is off. Icon 53 is an operation button for instructing other users to share the content of one's own statements made while in mute mode. This icon 53 is an operation button that can only be operated by users who are set to mute mode. Therefore, on the information processing device of a user who is not set to mute mode, icon 53 may be displayed in a grayed-out state to indicate that it cannot be operated. Each user can operate these icons 51, 52, and 53 by operating the mouse on the operation unit 12 and moving the mouse cursor 54.

[0044] For example, if mute mode is set for user A, an image 56 indicating that user A is in mute mode will be displayed near the image display area Ra that displays user A, as shown in Figure 4(a). This allows other users B and C to understand that user A is currently in mute mode.

[0045] When information processing devices 3a, 3b, and 3c detect an operation on icons 51, 52, and 53 by users A, B, and C, they send operation information to server 2. This operation information includes information indicating which icons 51, 52, and 53 were operated on by users A, B, and C.

[0046] All users can set a mute mode for a given user. For example, not only user A, but also users B and C can set mute mode for user A. In contrast, only the user who has been set to mute mode can unmute another user. For example, if user A is set to mute mode, only user A can unmute that user.

[0047] When the screen G1 shown in Figure 4(a) is displayed on the information processing device 3a, user A can disable mute mode by operating icon 51. When such operation is performed, user A's mute mode is disabled, and screen G1 transitions to the state shown in Figure 4(b). In screen G1 shown in Figure 4(b), the image 55 attached to icon 51 disappears, and the image 56 that was displayed near the image display area Ra that displays user A also disappears. Therefore, users A, B, and C can each understand that user A's mute mode has been disabled. When mute mode is disabled, the voice emitted by user A becomes audible to users B and C.

[0048] Returning to Figure 3, the operation detection unit 36 ​​detects the operations of users A, B, and C on screen G1 as described above. That is, the operation detection unit 36 ​​detects the operations of each user A, B, and C based on the operation information transmitted from the information processing devices 3a, 3b, and 3c, respectively. For example, if a user's operation is to set or de-mute mode, the operation detection unit 36 ​​activates the setting unit 37. Also, if the icon 53 is operated by a user who has set mute mode, the operation detection unit 36 ​​detects that an instruction has been given to share the content of what the user said while in mute mode with other users.

[0049] The settings unit 37 sets or deactivates mute mode for each user. When a mute mode setting operation is detected, the settings unit 37 identifies the user to whom the mute mode is to be set and sets the mute mode for the identified user in the audio processing unit 34. Also, when a mute mode deactivation operation is detected, the settings unit 37 identifies the user to whom the mute mode is to be deactivated and deactivates the mute mode set in the audio processing unit 34.

[0050] When the setting unit 37 sets the mute mode, it activates the speech recording unit 38. The speech recording unit 38 records the content of speech spoken by a user who is in mute mode. The speech recording unit 38 acquires audio information spoken by a user in mute mode from the audio acquisition unit 33. Based on the audio information acquired from the audio acquisition unit 33, the speech recording unit 38 records the content of speech spoken in mute mode. The speech recording unit 38 records the content of speech spoken in mute mode as audio information, text information, or summary information. For example, the speech recording unit 38 is pre-configured to record the content of speech spoken in mute mode as audio information, text information, or summary information. The speech recording unit 38 records the content of speech spoken in mute mode based on that setting. The speech recording unit 38 can also record all of the following: audio information, text information, and summary information. This speech recording unit 38 comprises a recording unit 41, a text generation unit 42, and a text summarization unit 43.

[0051] The recording unit 41 records the content of statements made by users in mute mode in the statement recording information 24, for each user. If it is set to record the content of statements as audio information, the recording unit 41 records the audio information of users in mute mode, acquired by the audio acquisition unit 33, in the statement recording information 24. At this time, the recording unit 41 splits the audio file and records it in the statement recording information 24 each time the audio of a user in mute mode is interrupted. For example, each time the audio of a user in mute mode is interrupted for a predetermined time, the recording unit 41 generates one audio file and repeats the process of recording that audio file in the statement recording information 24. Therefore, the statement recording information 24 may contain multiple audio files as the content of statements made by users in mute mode.

[0052] Furthermore, if it is set to record the content of speech as text information, the recording unit 41 records the text information generated by the text generation unit 42 in the speech recording information 24. The text generation unit 42 performs speech recognition on the audio information acquired by the audio acquisition unit 33 and converts the speech spoken by the user into text data. The recording unit 41 records the text information converted into text data by the text generation unit 42 in the speech recording information 24.

[0053] The text generation unit 42 splits the text file each time the user's voice is interrupted while in mute mode. For example, each time the user's voice is interrupted for a predetermined period of time while in mute mode, the text generation unit 42 repeats the process of generating one text file. The text generation unit 42 may also determine the content of the user's speech using natural language processing with artificial intelligence (AI) and split the text file each time the content of the speech changes. If multiple text files are generated sequentially while in mute mode, the recording unit 41 sequentially records these text files in the speech recording information 24.

[0054] Furthermore, if it is configured to record the content of the speech as summary information, the recording unit 41 activates the text summarization unit 43. The text summarization unit 43 generates summary information by summarizing the text information generated by the text generation unit 42. The text summarization unit 43 generates summary information by summarizing the text information, for example, using natural language processing with artificial intelligence (AI). If the text generation unit 42 generates multiple text files, the text summarization unit 43 generates summary information from each of those multiple text files. The recording unit 41 then records the summary information generated by the text summarization unit 43 in the speech recording information 24.

[0055] Therefore, the content of statements made by a user who is set to mute mode is recorded in the statement log information 24 for each statement. In addition, the statement log information 24 records the content of statements for each user. For example, if mute mode is set for multiple users, the statement log information 24 records the content of statements made by each of those multiple users while they are in mute mode, associated with the user who made the statement.

[0056] The recording process of spoken content by the speech recording unit 38 continues for the duration that mute mode is set. Therefore, the speech recording information 24 accumulates and records the content of spoken content made by the user while mute mode is enabled.

[0057] Furthermore, the message sharing unit 39 starts operating when the mute mode is deactivated by the setting unit 37, or when it detects that the icon 53 has been operated by a user who is in mute mode. The message sharing unit 39 performs the process of sharing the content of messages made by a user who is in mute mode with other users.

[0058] The message sharing unit 39 identifies users whose mute mode has been deactivated, or users who have operated the icon 53 while in mute mode. The message sharing unit 39 then reads the content of messages made by the identified user while in mute mode from the message log information 24. As described above, the message log information 24 may contain multiple messages from the identified user. Some of these messages may not need to be shared with other users. Therefore, the message sharing unit 39 prompts the identified user to select the messages to be shared with other users from the messages read from the message log information 24.

[0059] For example, if the identified user is user A, the message sharing unit 39, via the screen generation unit 30, displays a selection screen on user A's information processing device 3a to select the message content to share with other users B and C. This selection screen is displayed only on user A's information processing device 3a and not on the information processing devices 3b and 3c of users B and C. The message sharing unit 39 then identifies the message content selected by user A as the message content to be shared with other users B and C.

[0060] When the speech sharing unit 39 identifies speech content to be shared with other users, it outputs the identified speech content to the audio processing unit 34 or the screen generation unit 30, and transmits and shares it with other users. For example, when user A shares audio information with other users B and C, the speech sharing unit 39 outputs the audio information to the information processing devices 3b and 3c of other users B and C via the audio processing unit 34 and the audio output unit 35. At this time, the speech sharing unit 39 instructs the audio output unit 35 to play the audio based on the audio information at a speed of equal speed (1x speed) or faster. For example, the speech sharing unit 39 instructs to increase the audio playback speed to a speed faster than equal speed.

[0061] The information processing devices 3b and 3c output the audio provided from the server 2 by the speech sharing unit 39 through the speaker 15. As a result, users B and C listen to the audio output from the information processing devices 3b and 3c and understand what user A said while in mute mode. Therefore, user A does not need to repeat what they said while in mute mode. In addition, because the playback speed of the audio output from the information processing devices 3b and 3c is increased, users B and C can hear what user A said in a shorter amount of time, thereby improving the efficiency of the meeting.

[0062] Furthermore, when sharing audio information, the speech sharing unit 39 determines whether immediate sharing is possible. For example, if any user participating in a remote meeting is speaking, outputting audio based on the content of their speech recorded while in mute mode would result in two audio streams being output simultaneously. Therefore, the speech sharing unit 39 determines that immediate sharing is not possible when a user who is not in mute mode is speaking. If immediate sharing is not possible, the speech sharing unit 39 waits until the current speaker's speech ends. Then, the speech sharing unit 39 begins sharing the audio information at the moment the user's speech ends. This ensures that the content of user A's speech while in mute mode is accurately transmitted to other users B and C.

[0063] Furthermore, if the message sharing unit 39 cannot share immediately, it is preferable to inform other users B and C that the message recorded by the message recording unit 38 will be shared. For example, the message sharing unit 39 sends an information processing unit 3b and 3c a notification screen via the screen generation unit 30. This notification screen is displayed on the information processing units 3b and 3c, informing other users B and C that user A's message will be shared. As a result, users who are currently speaking will stop speaking. Consequently, it is possible to prevent the meeting from proceeding without user A's message being shared with other users B and C while user A is in mute mode.

[0064] Furthermore, when user A shares text information or summary information based on their statements with other users B and C, the statement sharing unit 39 transmits a statement sharing screen to the information processing devices 3b and 3c of users B and C via the screen generation unit 30. This statement sharing screen includes text information or summary information based on user A's statements. The statement sharing unit 39 may also provide the statement sharing screen to user A's information processing device 3a via the screen generation unit 30. When the server 2 provides the statement sharing screen, the information processing devices 3a, 3b, and 3c display it. Therefore, users B and C can view the statement sharing screen to confirm the content of user A's statements while in mute mode.

[0065] Figure 5 is a flowchart showing an example of a processing procedure performed by Server 2. This process is performed by Processor 20 executing Program 23, and is repeated during the progress of a remote meeting.

[0066] When Server 2 starts processing based on the flowchart in Figure 5, it determines whether or not mute mode is set for any of the users A, B, or C participating in the remote meeting (step S10). If mute mode is not set (NO in step S10), Server 2's processing ends. If mute mode is set for any of the users (YES in step S10), Server 2 executes the processes in steps S11 to S22. In the following explanation, the case where mute mode is set for user A will be used as an example.

[0067] If mute mode is set for user A (step S10), server 2 determines whether or not it has detected a statement from user A (step S11). For example, if server 2 obtains audio information from user A's information processing device 3a, it detects a statement from user A. Conversely, if it does not obtain audio information from user A's information processing device 3a, server 2 does not detect a statement from user A. If server 2 detects a statement from user A (YES in step S11), it records the content of user A's statement (step S12). In other words, the content of statements made by user A when mute mode is set is recorded in the statement recording information 24. At this time, the information recorded may be audio information, text information, or summary information.

[0068] Next, Server 2 determines whether or not it has detected User A's operation to disable mute mode (step S13). If it has detected an operation to disable mute mode (YES in step S13), Server 2 disables mute mode for User A (step S14). If it has not detected an operation to disable mute mode (NO in step S13), Server 2 determines whether or not it has detected User A's instruction to share the content of their speech (step S15). In other words, Server 2 determines whether or not User A has operated icon 53. If it has not detected an instruction to share the content of User A's speech (NO in step S15), Server 2 returns to step S11. In this case, Server 2 repeatedly executes the processes in steps S11 and S12. Therefore, each time Server 2 acquires audio information of User A, who is set to mute mode, it records the content of their speech based on that audio information in the speech record information 24.

[0069] Figure 6 shows the operation of the information processing device 3a and the server 2 during mute mode. When the information processing device 3a detects voice uttered by user A during mute mode (process P1), it generates voice information 48 representing that voice and sends it to the server 2 (process P2). When the server 2 receives the voice information 48 from the information processing device 3a of user A, who is in mute mode, it records the content of the utterance based on that voice information 48 (process P3). The information processing device 3a and the server 2 repeatedly execute processes P1 to P3 until user A's mute mode is released.

[0070] Returning to the flowchart in Figure 5, when the server detects that user A has disabled mute mode or has instructed user A to share their message, the server 2 displays a selection screen on user A's information processing device 3a (step S16). This selection screen allows user A to select the message to share with other users B and C.

[0071] Figure 7 shows the selection screen G2 displayed in the information processing device 3a. For example, the selection screen G2 is displayed as a pop-up in front of the remote meeting screen G1. The selection screen G2 displays multiple statements made by user A while in mute mode in a list format. If user A made only one statement while in mute mode, the selection screen G2 displays the content of that single statement. The selection screen G2 displays each statement with a checkbox 61 and a play button 62. The checkbox 61 is an operation unit for specifying whether to share with other users B and C. In the example in Figure 7, four statements are displayed, and it shows that user A has selected the second statement from the top. The play button 62 is a button for user A to confirm each statement. By operating this play button 62, user A plays back the statements in the information processing device 3a. In this case, playback is not performed by the other information processing devices 3b and 3c.

[0072] Furthermore, the selection screen G2 displays a share button 63 and a cancel button 64. The share button 63 is an operation button that instructs user A to share the content of their message with other users B and C. The cancel button 64 is an operation button that instructs user A not to share the content of their message with other users B and C.

[0073] Returning to the flowchart in Figure 5, when Server 2 displays the selection screen G2 on User A's information processing device 3a, it determines whether or not User A has instructed it to share the content of its statement (step S17). In other words, Server 2 determines whether or not User A has operated the share button 63 on the selection screen G2. If there is no instruction to share (NO in step S17), Server 2 proceeds to step S22. On the other hand, if there is an instruction to share (YES in step S17), Server 2 displays a guidance screen on the information processing devices 3b and 3c of the other users B and C (step S18).

[0074] Figure 8 shows the guidance screen G3 displayed in the information processing devices 3b and 3c. Guidance screen G3 is a screen that guides users to share the content of user A's statements while in mute mode. For example, guidance screen G3 is displayed as a pop-up in front of the remote meeting screen G1. When such guidance screen G3 is displayed in the information processing devices 3b and 3c of users B and C, users B and C know that the content of user A's statements will be shared. Therefore, users B and C will stop speaking if they are in the middle of speaking.

[0075] Returning to the flowchart in Figure 5, Server 2 displays the guidance screen G3 and then determines whether User A's statements can be immediately shared (step S19). If immediate sharing is not possible (NO in step S19), Server 2 waits until the statements of other users B and C are interrupted (step S20). If immediate sharing is possible (YES in step S19), or if the statements of other users B and C are interrupted (YES in step S20), Server 2 shares the statements made by User A while in mute mode with other users B and C (step S21). For example, if it is audio information, Server 2 sends audio information indicating the content of User A's statements to the information processing devices 3b and 3c of Users B and C. This allows Users B and C to hear the content of User A's statements while in mute mode. If it is text information or summary information, Server 2 sends a statement sharing screen containing text information or summary information indicating the content of User A's statements to the information processing devices 3b and 3c. Users B and C can understand the content of User A's statements by viewing the shared statements screen displayed on the information processing devices 3b and 3c.

[0076] Next, Server 2 determines whether User A's mute mode is still active (step S22). If the mute mode is active (YES in step S22), Server 2 returns to step S11 and repeats the process described above. If the mute mode is not active (NO in step S22), Server 2 terminates its process.

[0077] Figure 9 shows an example of how Server 2 shares the content of User A's statements. For example, a remote meeting is started by Users A, B, and C at timing T0. At timing T0, User A is already in mute mode. Once the remote meeting starts, Users B and C speak alternately, and the meeting progresses (processes P10, P11, P12, P13).

[0078] On the other hand, user A, who has muted mode enabled, forgets to disable it and speaks while users C and B are speaking (process P14). However, because user A is muted, user A's statements in process P14 are not heard by users B and C. At this time, server 2 records and retains the content of user A's statements in process P14.

[0079] User A makes a statement in process P14, but Users B and C do not respond, so User A realizes that they forgot to unmute. Therefore, after making a statement in process P14, User A unmutes at timing T1.

[0080] At timing T1, user C is making a statement (process P13). Therefore, server 2 waits until user C's statement ends. For example, if server 2 detects a period of silence exceeding a predetermined time in the audio information acquired from user C's information processing device 3c, it determines that user C's statement has ended. When user C's statement ends, at timing T2, server 2 outputs the content of the statement made by user A while in mute mode and transmits it to other users B and C (process P15). If the output content is audio information, server 2 increases the playback speed of the audio to a speed faster than normal. Therefore, the playback time Tb of the audio becomes shorter than the time Ta required for user A to speak while in mute mode. As a result, users B and C can hear user A's statement in less time than it would take for user A to repeat the same content themselves after the mute mode is released.

[0081] Figure 10 shows the message sharing screen G4 when the content of a message is shared as text information or summary information. The message sharing screen G4 is a screen that displays the content of user A's message as text information or summary information while in mute mode. For example, the message sharing screen G4 is displayed as a pop-up on the front side of the remote meeting screen G1. When such a message sharing screen G4 is displayed on the information processing devices 3b and 3c of users B and C, users B and C can understand the content of user A's message while in mute mode. Therefore, user A does not need to resend the same content themselves.

[0082] As described above, this remote conferencing system 1 includes a speech recording unit 38 that records the content of speech spoken by a user who is set to mute mode, and a speech sharing unit 39 that transmits and shares the content of speech recorded by the speech recording unit 38 with other users. Therefore, even if a user who is set to mute mode forgets to unmute themselves and speaks during a remote conference, the remote conferencing system 1 can transmit and share the content of speech spoken by the user while they were muted with other users. Thus, this remote conferencing system 1 eliminates the inconvenience for users of having to repeat what they said while muted, and improves the efficiency of remote conferences.

[0083] (modified version) Preferred embodiments of the present invention have been described above. However, the present invention is not limited to those described in the above embodiments, and various modifications are applicable.

[0084] For example, in the above embodiment, an example was described in which the speech recording unit 38 records speech made by a user who is set to mute mode while in mute mode. However, the present invention is not limited to this. For example, the speech recording unit 38 may record the content of speech made by a user who is not set to mute mode separately from the content of speech made by a user who is set to mute mode. In this case, the speech recording unit 38 can record the content of speech made by a user who is not set to mute mode as meeting minutes by activating the text generation unit 42. Server 2 may display the meeting minutes on screen G1 for remote meetings. When the meeting minutes are displayed on screen G1 for remote meetings, the displayed meeting minutes are updated each time a user who is not set to mute mode speaks. In such a case, when user A, who is set to mute mode, shares the content of speech made while in mute mode with other users B and C, Server 2 may reflect user A's speech in the meeting minutes displayed on screen G1 for remote meetings.

[0085] Figure 11 shows an example of meeting minutes displayed on screen G1 for remote meetings. Figure 11(a) shows meeting minutes 70 when users B and C speak during a remote meeting. If user A is set to mute mode, even if user A speaks while users B and C are conducting the meeting, users B and C will not hear what user A says. Therefore, user A's comments will not be reflected in meeting minutes 70.

[0086] If user A realizes after making a statement while muted that they forgot to unmute, they will either unmute or share the content of their statement made while muted. When such an operation is performed, server 2 adds user A's statement made while muted to the meeting minutes 70 in Figure 11(a), and generates the meeting minutes 71 shown in Figure 11(b). Server 2 then displays the meeting minutes 71 on screen G1 for the remote conference. In the meeting minutes 71 shown in Figure 11(b), user A's statement 72 made while muted has been added. Furthermore, it is indicated that this statement 72 was made by user A while muted. It is preferable that server 2 displays the statement 72 in a different manner from the other statements to draw the attention of users B and C to user A's statement 72 made while muted.

[0087] Furthermore, the above embodiment illustrates a case where server 2 is installed on a cloud 5 such as the internet. However, the location of server 2 is not necessarily limited to cloud 5. For example, server 2 may be installed in a local environment such as an internal company network.

[0088] Furthermore, the above embodiment illustrates a case where the program 23 executed on server 2 is pre-installed on server 2. However, program 23 is not limited to being pre-installed on server 2. That is, program 23 executed on server 2 can be the subject of a transaction on its own. Therefore, program 23 may be provided in a manner that allows it to be downloaded to server 2 via a network such as the Internet. Also, program 23 may be provided in a manner that is recorded on a computer-readable recording medium such as a CD-ROM or USB memory. [Explanation of Symbols]

[0089] 1. Remote conferencing system 2 servers 3a,3b,3c Information processing equipment 4 Network 20 processors 22 Memory section 24. Record of statements 30 Screen generation section 31 Screen provision department 32 Image acquisition unit 33. Voice acquisition unit 35 Audio output section 36 Operation detection unit 37 Settings Section 38. Recording Section 39. Discussion Sharing Section 41 Records Section 42 Text generation unit 43 Text Summary Section

Claims

1. A remote conferencing system that allows multiple users to connect simultaneously from different locations and conduct voice calls, A voice acquisition unit that acquires the voice spoken by each user, An audio output unit provides audio acquired by the audio acquisition unit to other users, A settings section that allows users to set a mute mode that prevents their voice from being shared with other users, A speech recording unit that records the content of speech spoken by a user who has set the device to mute mode, A speech sharing unit transmits and shares the content of speech recorded by the speech recording unit to other users, A remote conferencing system characterized by having the following features.

2. The remote conferencing system according to claim 1, characterized in that the speech recording unit records the audio acquired by the audio acquisition unit as the content of speech by a user who is set to mute mode.

3. The remote conferencing system according to claim 2, characterized in that the speech sharing unit transmits and shares the audio recorded by the speech recording unit with other users.

4. The remote conferencing system according to claim 2, characterized in that the speech recording unit divides and records the audio file each time the audio acquired by the audio acquisition unit is interrupted.

5. The remote conferencing system according to claim 4, characterized in that, when multiple audio files are recorded by the speech recording unit, the speech sharing unit transmits and shares the audio recorded in an audio file selected from among the multiple audio files to other users.

6. The remote conferencing system according to claim 3 or 5, characterized in that when the speech sharing unit transmits and shares audio with other users, it increases the playback speed of the audio to a speed equal to or greater than the normal speed.

7. The remote conferencing system according to claim 1, characterized in that the speech recording unit records text information obtained by converting the audio acquired by the audio acquisition unit into text data as the content of speech by a user set to mute mode.

8. The remote conferencing system according to claim 7, characterized in that the speech sharing unit transmits and shares the text information recorded by the speech recording unit to other users.

9. The remote conferencing system according to claim 8, characterized in that the speech recording unit divides and records the text file each time the audio acquired by the audio acquisition unit is interrupted.

10. The remote conferencing system according to claim 9, characterized in that, when the speech sharing unit has recorded multiple text files, it transmits and shares a selected text file from among the multiple text files to other users.

11. The remote conferencing system according to claim 1, characterized in that the speech recording unit converts the audio acquired by the audio acquisition unit into text data as the content of speech by a user set to mute mode, and records summary information which is a summary of the text contained in the text data.

12. The remote conferencing system according to claim 1, characterized in that the speech sharing unit transmits and shares the content of speech recorded by the speech recording unit to other users based on an operation by a user set to mute mode.

13. The remote conferencing system according to claim 12, characterized in that the speech sharing unit transmits and shares the content of speech recorded by the speech recording unit to other users based on the operation of a user who has set the mute mode to unmute mode.

14. The remote conferencing system according to claim 12, characterized in that the speech sharing unit transmits and shares the content of speech recorded by the speech recording unit to other users based on a sharing instruction operation by a user set to mute mode.

15. The remote conferencing system according to claim 12, characterized in that when the speech sharing unit detects an operation by a user set to mute mode and the audio acquisition unit is acquiring audio from another user, the speech recording unit transmits and shares the content of the speech recorded by the speech recording unit to the other user at the moment the audio acquired by the audio acquisition unit is interrupted.

16. The remote conferencing system according to claim 15, characterized in that the speech sharing unit, upon detecting an operation by a user set to mute mode, notifies other users that the content of the speech recorded by the speech recording unit will be shared.

17. A method for controlling a remote conferencing system that allows multiple users to connect simultaneously from different locations and conduct voice calls, A voice acquisition step that acquires the voice spoken by each user, A voice output step provides the voice acquired in the voice acquisition step to another user, A setting step to enable a mute mode for each user, which prevents the voice spoken by that user from being shared with other users, A speech recording step that records the content of speech spoken by a user who is set to mute mode, A speech sharing step involves transmitting and sharing the content of the speech recorded in the speech recording step with other users. A control method characterized by having the following features.

18. A program that runs on a server of a remote conferencing system that allows multiple users to connect simultaneously from different locations and conduct voice calls, wherein the server has: A voice acquisition step that acquires the voice spoken by each user, A voice output step provides the voice acquired in the voice acquisition step to another user, A setting step to enable a mute mode for each user, which prevents the voice spoken by that user from being shared with other users, A speech recording step that records the content of speech spoken by a user who is set to mute mode, A speech sharing step involves transmitting and sharing the content of the speech recorded in the speech recording step with other users. A program characterized by causing the execution of a specific action.